Senior Software Engineer, DGX Cloud AI Infrastructure

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CAAustin, TXRedmond, WA
Salary
$184,000–$287,500 / yr
Posted
99 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $187k
This role $236k
$126k most similar roles pay here $305k

This role pays more than 83% of similar roles. Most pay $149,350–$224,312 — the shaded band above. At the midpoint, this role pays about $236k versus about $187k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, DGX Cloud AI Infrastructure

Senior Software Engineer, DGX Cloud AI Infrastructure joins the team to lead the bring-up, triage, benchmarking, and optimization of distributed training and inference workloads across large-scale GPU platforms. This individual contributor role involves setting technical direction for communication libraries, model frameworks, and inference/training stacks to ensure efficient LLM performance. The engineer will perform deep performance investigations on multi-GPU and multi-node deployments, build resilience and failure-attribution capabilities, and develop repeatable benchmark suites. Key responsibilities include profiling workloads across compute, memory, networking, and communication layers while analyzing scaling efficiency for distributed models using data, tensor, pipeline, and expert parallelism. Required skills include expert-level Python and C/C++ programming, experience with NCCL, CUDA-aware execution, and the RDMA stack including IB verbs and UCX. The role addresses the technical challenge of maintaining productive large clusters for high-scale AI workloads.

What does a Software Engineer earn in California?

Median $214000 from 775 postings across 63 companies.

See salary data

What you'll do

  • Lead the bring-up, validation, and debugging of large-scale AI clusters and end-to-end workloads.
  • Tune and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo, and TensorRT-LLM.
  • Profile and optimize performance across compute, memory, networking, and communication layers using Nsight Systems and NCCL.
  • Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism.
  • Perform root-cause analysis of complex failures including hangs, regressions, and topology sensitivity in large environments.
  • Build resilience and failure-attribution systems to detect and triage node and fabric failures at scale.
  • Develop repeatable benchmark suites, automation, and qualification workflows for new hardware platforms.
  • Provide data-driven recommendations based on profiling results and cluster characterization to influence technical standards.

What we're looking for

  • Bachelor’s or Master’s degree in Computer Science or a related technical field (or equivalent experience).
  • 8+ years of experience developing software infrastructure for large-scale AI or HPC systems.
  • Proven track record of technical leadership and mentoring other engineers.
  • Expertise debugging and triaging AI applications across the full stack from application layer to hardware.
  • Deep hands-on experience with NCCL, CUDA-aware distributed execution, and multi-GPU/multi-node workloads.
  • Expert-level programming skills in Python and C/C++.
  • Experience operating workloads in scheduled, containerized cluster environments.
  • Strong analytical, debugging, and communication skills to influence across teams.

More like this

Similar roles

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp

Principal Developer, AI Networking

Nvidia

Remote (Santa Clara, CA) +3 91 days ago $272,000$431,250
NCCL RDMA CUDA PyTorch TensorFlow C++ Python Bash RoCE MPI SHARP Distributed Systems LLM Training High-Performance Networking Profiling Benchmarking
10+ yrs exp Remote

NCX Senior Engineer

Nvidia

Remote (Santa Clara, CA) +1 8 days ago $184,000$287,500
Kubernetes Python Go PyTorch TensorFlow CUDA MLOps Terraform Ansible Prometheus Grafana InfiniBand RoCE Triton NeMo CI/CD Linux Distributed Systems NVIDIA Reference Architectures
8+ yrs exp Remote

Senior Solutions Architect, Generative AI

Nvidia

Remote (Santa Clara, CA) +1 43 days ago $184,000$287,500
GPU InfiniBand RoCE RDMA NCCL NVLink NVSwitch Kubernetes Slurm Python Linux High-Performance Computing Distributed Systems Shell Scripting DCGM Nsight Systems
6+ yrs exp Remote

Senior Cloud Software Engineer

Nvidia

Remote 6 days ago $152,000$241,500
Kubernetes AWS GCP Azure Go Python Rust C++ Java Distributed Systems Cloud-Native Data Management Storage Systems Performance Engineering Observability
5+ yrs exp Remote