Senior DGX Cloud AI Infrastructure Software Engineer

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CAAustin, TXRedmond, WA
Salary
$184,000–$287,500 / yr
Posted
162 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $193k
This role $236k
$133k most similar roles pay here $304k

This role pays more than 78% of similar roles. Most pay $149,462–$235,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $193k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior DGX Cloud AI Infrastructure Software Engineer

As a Senior DGX Cloud AI Infrastructure Software Engineer on the DGX Cloud AI Efficiency Team, you will develop infrastructure software and tools to support large-scale pre-training, post-training, and inference. You will be responsible for building and maintaining systems that ensure high efficiency and availability while co-designing APIs for integration with resiliency stacks. Your daily work involves optimizing libraries, defining reliability metrics, and performing root cause analysis on failures from the application level down to the hardware level. The role requires proficiency in Python, C/C++, and various scripting languages, alongside experience with observability platforms like ELK, Prometheus, or Loki. You will also utilize technologies such as RDMA stacks including NCCL, IB verbs, ucx, and libfabrics while working with deep learning frameworks like PyTorch, TensorFlow, JAX, and Ray to solve complex infrastructure challenges.

What you'll do

  • Develop infrastructure software and tools for large-scale pre-training, post-training, and inference.
  • Optimize tools and libraries to improve the efficiency and resiliency of AI systems.
  • Co-design and implement APIs for integration with NVIDIA's resiliency stacks.
  • Enhance products and infrastructure underpinning NVIDIA's AI platforms.
  • Define actionable reliability metrics to track and improve system and service performance.
  • Perform root cause analysis and triage failures from the application level down to the hardware level.

What we're looking for

  • Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems.
  • Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).
  • Proficiency in programming languages such as Python, C/C++, and various scripting languages.
  • Experience building and scaling large-scale distributed systems.
  • Experience with AI training and inferencing infrastructure services.
  • Strong debugging skills to analyze and triage failures from the application level down to the hardware level.
  • Experience with observability platforms for monitoring and logging such as ELK, Prometheus, or Loki.
  • Knowledge of quality software engineering practices including test development, defensive programming, version control, and CI.

More like this

Similar roles

Senior Cloud Software Engineer

Nvidia

Remote 6 days ago $152,000$241,500
Kubernetes AWS GCP Azure Go Python Rust C++ Java Distributed Systems Cloud-Native Data Management Storage Systems Performance Engineering Observability
5+ yrs exp Remote

Senior Full-Stack Lead Engineer

Nvidia

Santa Clara, CA +1 64 days ago $224,000$356,500
React Next.js Vue Nuxt TypeScript JavaScript Python Node.js Kubernetes Docker AWS GCP Azure CI/CD Prometheus Grafana OpenSearch Loki PyTorch TensorFlow JAX Ray
10+ yrs exp

Senior Performance Engineer

Nvidia

Remote (Santa Clara, CA) +2 45 days ago $224,000$356,500
C++ Python CUDA PyTorch JAX XLA GPU Computing Distributed Systems Performance Engineering Benchmarking Profiling Observability High-Performance Computing Data Analysis Automation Workflows
Remote