Senior Software Engineer, DGX Cloud Production Engineering

Nvidia

Confirmed live today High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CA
Salary
$184,000–$287,500 / yr
Employment
Full-time
Posted
11 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $191k
This role $236k
$122k most similar roles pay here $305k

This role pays more than 87% of similar roles. Most pay $155,050–$226,200 — the shaded band above. At the midpoint, this role pays about $236k versus about $191k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1391 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 1116 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, DGX Cloud Production Engineering

As a Senior Software Engineer, DGX Cloud Production Engineering, you will join the production engineering team to build and operate automation, tooling, and operational systems for large-scale GPU infrastructure. You will develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations across NVIDIA Cloud Partners and on-prem environments. Your daily work involves improving Day 0 through Day 2 workflows, reducing manual production touches via APIs, GitOps, and agent-assisted workflows while participating in incident response and debugging. The role requires proficiency in Python, Go, Linux, Kubernetes, containers, and infrastructure automation. You will solve technical challenges regarding the reliability and scalability of GPU clusters for AI research and production workloads. Key technologies include Terraform, ArgoCD, and fleet automation to ensure infrastructure is production-ready across complex distributed systems.

What does a Software Engineer earn in California?

Median $214000 from 925 postings across 71 companies.

See salary data

What you'll do

  • Build and operate automation for large-scale GPU clusters across cloud and on-prem environments.
  • Develop tools and services for provisioning, validation, upgrades, monitoring, and repair of cluster operations.
  • Improve Day 0, Day 1, and Day 2 workflows for cluster bringup and production handoff.
  • Reduce manual production touches through APIs, GitOps, and agent-assisted workflows.
  • Participate in on-call rotations, incident response, and debugging of distributed systems.
  • Partner with platform, storage, networking, and security teams to make infrastructure production-ready.

What we're looking for

  • 8+ years of experience building or operating production infrastructure.
  • Strong programming skills in Python, Go, or similar languages.
  • Experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
  • Ability to troubleshoot distributed systems in production environments.
  • Clear communication and ability to work across teams.
  • BS/MS in Computer Science or equivalent experience.
  • Experience with GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, or fleet automation (preferred).
  • Experience with SLOs, on-call, incident response, observability, reliability practices, BMaaS, VMaaS, managed Kubernetes, or multi-cloud infrastructure (preferred).

More like this

Similar roles

Principal Software Engineer, DGX Cloud Production Engineering

Nvidia

Remote (Santa Clara, CA) 139 days ago $272,000–$431,250
Kubernetes Go Python GitOps Linux GPU Clusters AI/ML Infrastructure Distributed Systems Infrastructure Automation APIs Observability SLOs Multi-cloud BMaaS VMaaS High-Performance Computing
10+ yrs exp Remote

Senior Production Engineer

Nvidia

Remote (Santa Clara, CA) 2 days ago $184,000–$287,500
Kubernetes Python Go Terraform Argo CD vLLM SGLang PyTorch TensorRT-LLM NVIDIA Dynamo CUDA NCCL GitOps CI/CD Linux Distributed Systems Infrastructure as Code
8+ yrs exp Remote

Senior Software Engineer, Distributed Systems Engineering

Nvidia

Remote (Santa Clara, CA) 2 days ago $184,000–$287,500
Kubernetes GPU Go Python Slurm Bright Cluster Manager Distributed Systems Cluster Management Data Structures Algorithms Systems Programming Monitoring Telemetry
8+ yrs exp Remote

Principal Software Engineer

Nvidia

Remote 58 days ago $272,000–$431,250
Kubernetes Go C Linux CI/CD GitLab Argo Flux Container Orchestration Distributed Systems Cloud Computing GPU DPU Confidential Computing
10+ yrs exp Remote