Senior System Software Engineer, GPU Performance

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$152,000–$241,500 / yr
Posted
6 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $198k
This role $197k
$141k most similar roles pay here $252k

This role pays more than 53% of similar roles. Most pay $159,562–$235,750 — the shaded band above. At the midpoint, this role pays about $197k versus about $198k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior System Software Engineer, GPU Performance

As a Senior System Software Engineer - GPU Performance on the GPU Communications Libraries and Networking team, you will influence the roadmap of communication libraries like NCCL, NVSHMEM, and UCX. You will conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters to optimize end-to-end application performance for Deep Learning and HPC workloads. Your daily responsibilities include studying hardware and software interactions, evaluating proof-of-concepts, triaging customer issues, and building tools to visualize performance data. You will utilize C/C++, Python, and experience with MPI or UCX to solve bottlenecks across high-speed interconnects like NVLink and PCIe. The role requires expertise in system architecture, operating systems, and container tools such as Kubernetes and SLURM to address the massive compute demands of scaling applications across tens of thousands of GPUs.

What does a System Software Engineer earn?

Median $235750 from 43 postings across 5 companies.

See salary data

What you'll do

  • Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters.
  • Analyze the interaction between communication libraries and hardware/software components across the stack.
  • Evaluate proof-of-concepts and perform trade-off analyses for multiple technical solutions.
  • Triage and root-cause performance issues reported by customers.
  • Build tools and infrastructure to collect, visualize, and analyze large volumes of performance data.
  • Implement micro-benchmarks in C/C++ and modify the existing codebase as required.
  • Debug performance issues across the entire hardware and software stack.

What we're looking for

  • Master's degree (or equivalent experience) or PhD in Computer Science or a related field.
  • 3+ years of experience with parallel programming and at least one communication runtime like MPI, NCCL, UCX, or NVSHMEM.
  • Experience conducting performance benchmarking and triage on large-scale HPC clusters.
  • Strong understanding of computer system architecture, operating systems principles, and hardware-software interactions.
  • Ability to implement micro-benchmarks in C/C++ and modify the existing codebase.
  • Proficiency in a scripting language, preferably Python.
  • Familiarity with containers and cloud provisioning tools such as Kubernetes, SLURM, Ansible, or Docker.
  • Experience with RDMA, GPU programming (CUDA), or Deep Learning frameworks like PyTorch and TensorFlow.

More like this

Similar roles

Senior Software Engineer, NCCL

Nvidia

Santa Clara, CA 3 days ago $152,000$241,500
C++ C CUDA NVIDIA GPUs PyTorch TensorFlow NCCL UCX MPI OpenSHMEM Linux High-Performance Computing InfiniBand iWARP Parallel Programming Deep Learning Frameworks
5+ yrs exp

Principal Deep Learning Communication Architect

Nvidia

Remote (Santa Clara, CA) +1 150 days ago $272,000$431,250
NCCL UCX UCC NVSHMEM CUDA InfiniBand RDMA RoCE TensorRT-LLM vLLM SGLang Megatron-Core DeepSpeed JAX XLA PyTorch Distributed Ray HPC High-Performance Computing
10+ yrs exp Remote