Software Engineer, DGX Cloud AI Infrastructure

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CAAustin, TXORWARedmond, WA
Salary
$108,000–$178,250 / yr
Employment
Full-time
Posted
4 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $199k
This role $143k
$92k most similar roles pay here $258k

This role pays less than 87% of similar roles. Most pay $161,300–$236,503 — the shaded band above. At the midpoint, this role pays about $143k versus about $199k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1150 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 912 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Software Engineer, DGX Cloud AI Infrastructure

Software Engineer, DGX Cloud AI Infrastructure - New College Grad 2026 will join the team to focus on the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across multi-GPU and multi-node deployments. The role involves debugging large-scale AI clusters, tuning pre-training and post-training workloads, and developing automated benchmarking tools and failure-attribution workflows. You will perform root-cause analysis for failures in distributed environments while managing runtime settings and communication parameters. Key technical requirements include proficiency in Python and C/C++, along with experience in PyTorch, NeMo, Megatron, and TensorRT-LLM. Candidates should possess skills in CUDA-aware distributed execution, the RDMA software stack including NCCL and InfiniBand, and containerized cluster environments. The work centers on solving complex problems in deep learning systems, GPU performance, and large-scale infrastructure for generative AI workloads.

What does a Software Engineer earn in California?

Median $214000 from 830 postings across 69 companies.

See salary data

What you'll do

  • Bring up, validate, and debug large-scale AI clusters and end-to-end workloads.
  • Tune and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo, and TensorRT-LLM.
  • Perform root-cause analysis of failures in large distributed environments.
  • Develop tools for resilience, failure detection, and fault attribution across the cluster.
  • Build and maintain repeatable benchmark suites, automation, and qualification workflows on new platforms.
  • Optimize runtime settings, communication parameters, and deployment configurations.
  • Provide data-driven recommendations based on profiling, benchmarking, and cluster characterization.

What we're looking for

  • Bachelor’s or Master’s degree in Computer Science or a related technical field (or equivalent experience).
  • Experience developing software for AI, HPC, or systems-level applications.
  • Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.
  • Background in debugging and scaling distributed systems.
  • Experience debugging and triaging AI applications across the full stack from application to hardware.
  • Experience operating workloads in scheduled, containerized cluster environments.
  • Strong Python and C/C++ programming skills.
  • Familiarity with RDMA software stacks, InfiniBand/RoCE congestion debugging, or building resilience systems (preferred).

More like this

Similar roles

Software Engineer, DGX Cloud AI Infrastructure

Nvidia

Remote (Santa Clara, CA) +4 4 days ago $116,000–$189,750
PyTorch NeMo Megatron TensorRT-LLM Python C/C++ CUDA NCCL RDMA InfiniBand RoCE UCX libfabric MLPerf HPC Distributed Computing
3+ yrs exp Remote

Principal Developer, AI Networking

Nvidia

Remote (Santa Clara, CA) +3 113 days ago $272,000–$431,250
NCCL RDMA CUDA PyTorch TensorFlow C++ Python Bash RoCE MPI SHARP Distributed Systems LLM Training High-Performance Networking Profiling Benchmarking
10+ yrs exp Remote

Senior Deep Learning Frameworks CUDA Software Engineer

Nvidia

Remote (Santa Clara, CA) +1 33 days ago $184,000–$287,500
CUDA PyTorch JAX C++ Python TRT-LLM vLLM SGLang TensorRT Triton NCCL MPI UCX XLA HPC Kernel Authoring NVIDIA Nsight Systems Compiler Technologies
8+ yrs exp Remote