Senior Systems Software Engineer, GPU Performance at Scale

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$184,000–$287,500 / yr
Posted
2 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $198k
This role $236k
$144k most similar roles pay here $303k

This role pays more than 82% of similar roles. Most pay $159,750–$235,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $198k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Systems Software Engineer, GPU Performance at Scale

As a Senior Systems Software Engineer - GPU Performance at Scale, you will join the team to drive innovation in AI and GPU computing. You will be responsible for implementing performance practices in large-scale GPU infrastructure, developing tools and methodologies to validate and improve multiple datacenter products concurrently. Your daily work involves aligning next-generation AI workloads with hardware, decomposing complex stability issues into reproduction cases, and resolving critical firmware and software issues. To succeed, you must possess expertise in CUDA, Linux-based operating systems, and container technologies like Docker. You will utilize C, C++, Python, and Bash to optimize performance platforms. The role focuses on the technical challenge of scaling high-performance compute runs, managing large-scale HPC environments, and optimizing system architectures for advanced deep learning workloads within complex datacenter infrastructures.

What you'll do

  • Implement performance practices and tools to validate and improve multiple datacenter products concurrently.
  • Align next-generation AI workloads with datacenter hardware including GPUs, CPUs, and networking components.
  • Develop engineering solutions to provide continuous insights into the performance of AI workloads in evolving environments.
  • Decompose complex performance or stability issues into minimal reproduction cases to identify root causes.
  • Analyze, debug, and resolve critical firmware and software issues to optimize high-performance AI workloads at scale.
  • Architect and optimize large-scale performance platforms by collaborating with HPC, OS, CPU, and GPU specialists.
  • Create improved workflows and new solutions for researchers, developers, and customers.

What we're looking for

  • Proven understanding of accelerated computing software stacks including CUDA.
  • Experience with modern cloud and container-based enterprise computing architectures, preferably with Slurm.
  • Strong programming and scripting skills in C, C++, Python, and Bash.
  • Deep expertise in systems architecture and the impact of various components on performance.
  • Experience with Linux-based operating systems and container technology, preferably with Docker.
  • Experience supporting high-performance computing or deep learning in engineering or academic research communities.
  • A BS in Engineering, Mathematics, Physics, or Computer Science is required; an MS or PhD is desirable.
  • At least 8 years of applicable experience are desired for advanced candidates.

More like this

Similar roles

Senior System Software Engineer, Performance - CUDA Driver

Nvidia

Santa Clara, CA 9 days ago $184,000$287,500
CUDA C++ C GPU Kernel Software performance-engineering Computer Architecture Operating Systems Concurrency Memory Management Python HPC Deep Learning Profiling Firmware Hardware/Software Co-design Pre-silicon Analysis
7+ yrs exp Hybrid

Senior System Software Engineer, GPU Performance

Nvidia

Remote 6 days ago $152,000$241,500
HPC C++ Python CUDA NCCL UCX NVSHMEM MPI Infiniband Ethernet RDMA PyTorch TensorFlow Kubernetes Docker SLURM Ansible Performance Engineering Parallel Programming
Remote

Senior Systems Software Engineer, CUDA Driver

Nvidia

Remote (Santa Clara, CA) +1 11 days ago $184,000$287,500
C C++ CUDA Kernel Mode Development Parallel Computing Linux Windows System Software Operating Systems Virtual Memory memory-mapped IO Interconnects
7+ yrs exp Remote