Senior Data Center Performance Engineer, Benchmarking and Optimization

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$184,000–$287,500 / yr
Posted
15 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $201k
This role $236k
$134k most similar roles pay here $304k

This role pays more than 80% of similar roles. Most pay $165,750–$235,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $201k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Data Center Performance Engineer, Benchmarking and Optimization

Senior Data Center Performance Engineer - Benchmarking and Optimization will lead performance benchmarking and optimization efforts for data center products to ensure solutions deliver industry-leading performance for accelerated computing workloads. The role involves designing comprehensive benchmarking strategies, characterizing real-world AI training, inference, and HPC workloads at scale, and developing automation tools for monitoring and analysis. You will identify bottlenecks across compute, memory, network, and storage subsystems while providing architectural recommendations for future systems. Required skills include proficiency in Linux perf and NVIDIA Nsight Systems, along with experience in GPU computing, CUDA, InfiniBand, RoCE, and NVLink. The candidate must possess programming skills in Python, C++, and shell scripting. Experience with PyTorch, TensorFlow, JAX, NCCL, Docker, Kubernetes, and SLURM is highly valued to solve complex performance issues within the data center platform ecosystem.

What you'll do

  • Design and execute comprehensive performance benchmarking strategies for data center platforms and products.
  • Characterize real-world AI training, inference, and HPC workloads at scale.
  • Define, track, and report key performance indicators including throughput, latency, efficiency, and scaling.
  • Build automation tools and frameworks for performance monitoring and analysis.
  • Identify and analyze performance bottlenecks across compute, memory, network, and storage subsystems.
  • Drive performance improvements through system tuning, configuration optimization, and architectural recommendations.

What we're looking for

  • M.S. or Ph.D. in Computer Science, Electrical Engineering, or a related field (or equivalent experience).
  • 8+ years of experience in performance engineering or system architecture.
  • Deep understanding of computer architecture, hardware-software interaction, and computing at scale.
  • Proficiency in performance profiling tools such as Linux perf and NVIDIA Nsight Systems.
  • Familiarity with GPU computing, parallel programming (CUDA), and HPC networking technologies like InfiniBand, RoCE, and NVLink.
  • Programming skills in Python, C++, and shell scripting.
  • Experience with AI/ML frameworks including PyTorch, TensorFlow, or JAX.
  • Knowledge of MPI, NCCL, distributed training, and container/orchestration tools like Docker and Kubernetes.

More like this

Similar roles

Senior Performance Engineer

Nvidia

Remote (Santa Clara, CA) +2 45 days ago $224,000$356,500
C++ Python CUDA PyTorch JAX XLA GPU Computing Distributed Systems Performance Engineering Benchmarking Profiling Observability High-Performance Computing Data Analysis Automation Workflows
Remote

Senior System Software Engineer, GPU Performance

Nvidia

Remote 6 days ago $152,000$241,500
HPC C++ Python CUDA NCCL UCX NVSHMEM MPI Infiniband Ethernet RDMA PyTorch TensorFlow Kubernetes Docker SLURM Ansible Performance Engineering Parallel Programming
Remote

Senior Data Center Facilities Engineer

Oracle

Abilene, TX 38 days ago $91,400$187,000
Data Center Infrastructure Critical Facility Operations CMMS Preventative Maintenance Fire and Life Safety Systems Incident Management Risk Management Contract Administration Electrical Engineering Industrial Engineering Process Engineering
8+ yrs exp

Chief Engineer, Data Centers

JLL (Jones Lang LaSalle)

Claude, TX 43 days ago
HVAC Electrical Distribution UPS BESS ATS CRAC CRAH BMS EPMS BACNet ModBus Corrigo CMMS Excel Google Suite root cause analysis (RCA
4+ yrs exp

Chief Engineer, Data Centers

JLL (Jones Lang LaSalle)

Charlotte, NC 7 days ago
Diesel Generator Systems UPS ATS Electrical Switchgear BMS DDC CMMS Pumps Cooling Towers Water Treatment CCTV MS Office Pneumatics
10+ yrs exp