Software Engineer, DGX Cloud AI Infrastructure

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CAAustin, TXORWARedmond, WA
Salary
$116,000–$189,750 / yr
Employment
Full-time
Posted
4 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $199k
This role $153k
$102k most similar roles pay here $250k

This role pays less than 81% of similar roles. Most pay $161,300–$235,750 — the shaded band above. At the midpoint, this role pays about $153k versus about $199k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1150 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 912 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Software Engineer, DGX Cloud AI Infrastructure

As a Software Engineer, DGX Cloud AI Infrastructure, you will join the team responsible for bringing up, benchmarking, and optimizing distributed training and inference workloads across large-scale GPU platforms. You will perform root-cause analysis of failures in distributed environments while designing and implementing benchmarking tooling, automation, and debugging workflows. Your daily work involves validating end-to-end workloads, tuning runtime settings, and providing data-driven recommendations based on profiling and cluster characterization. The role requires proficiency in Python and C/C++, along with experience in PyTorch, NeMo, Megatron, and TensorRT-LLM. You will navigate complex technical challenges involving CUDA-aware distributed execution, the RDMA software stack including NCCL and InfiniBand, and containerized cluster environments. This position focuses on solving critical performance and reliability issues within large-scale AI infrastructure to support massive language model workloads across multi-node deployments.

What does a Software Engineer earn in California?

Median $214000 from 830 postings across 69 companies.

See salary data

What you'll do

  • Bring up, validate, and debug large-scale AI clusters and end-to-end workloads.
  • Tune and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo, and TensorRT-LLM.
  • Perform root-cause analysis of failures in large distributed environments.
  • Develop tools for resilience, failure detection, and fault attribution across the cluster.
  • Build and maintain repeatable benchmark suites, automation, and qualification workflows on new platforms.
  • Optimize runtime settings, communication parameters, and deployment configurations in partnership with internal teams.
  • Provide data-driven recommendations based on profiling, benchmarking results, and cluster characterization.

What we're looking for

  • Bachelor’s or Master’s degree in Computer Science or a related technical field (or equivalent experience).
  • 3+ years of experience developing software for AI, HPC, or systems-level applications.
  • Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.
  • Experience debugging and scaling distributed systems.
  • Experience debugging and triaging AI applications across the full stack from application to hardware.
  • Experience operating workloads in scheduled, containerized cluster environments.
  • Strong Python and C/C++ programming skills.
  • Familiarity with RDMA software stacks, InfiniBand/RoCE congestion debugging, or building resilience systems (preferred).

More like this

Similar roles

Software Engineer, DGX Cloud AI Infrastructure

Nvidia

Remote (Santa Clara, CA) +4 4 days ago $108,000–$178,250
PyTorch TensorRT-LLM NeMo Megatron Python C++ CUDA NCCL RDMA InfiniBand RoCE UCX libfabric MLPerf HPC Distributed Computing
Remote

Principal Developer, AI Networking

Nvidia

Remote (Santa Clara, CA) +3 113 days ago $272,000–$431,250
NCCL RDMA CUDA PyTorch TensorFlow C++ Python Bash RoCE MPI SHARP Distributed Systems LLM Training High-Performance Networking Profiling Benchmarking
10+ yrs exp Remote

Senior Deep Learning Frameworks CUDA Software Engineer

Nvidia

Remote (Santa Clara, CA) +1 33 days ago $184,000–$287,500
CUDA PyTorch JAX C++ Python TRT-LLM vLLM SGLang TensorRT Triton NCCL MPI UCX XLA HPC Kernel Authoring NVIDIA Nsight Systems Compiler Technologies
8+ yrs exp Remote