Software Engineer, SRE and Production Engineering

Nvidia

Confirmed live today High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CATXVASDUTSC
Salary
$184,000–$287,500 / yr
Employment
Full-time
Posted
2 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $180k
This role $236k
$125k $305k
below market most similar roles pay here above market

This role pays more than 85% of similar roles. Most pay $142,437–$217,006 — the blue band above. At the midpoint, this role pays about $236k versus about $180k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1472 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 1109 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Software Engineer, SRE and Production Engineering

JOB TITLE: Software Engineer, SRE and Production Engineering - DGX Cloud The Software Engineer, SRE and Production Engineering - DGX Cloud joins the team responsible for building and operating large-scale GPU infrastructure for AI workloads. This role involves building software and operational tooling to move GPU capacity from installed hardware to production services for BMaaS and VMaaS environments. Key responsibilities include building automation for bare-metal provisioning, hardware validation, firmware upgrades, and cluster lifecycle management. You will build tools using BMC and Redfish interfaces to assess hardware health and facilitate recovery workflows. Required skills include Go or Python, Linux, and experience with NVIDIA NVL72 systems and BlueField-3 DPUs. You will diagnose failures across servers, networking, and Kubernetes while managing production reliability through on-call duties, incident response, and root-cause analysis for large-scale AI infrastructure.

What does a Software Engineer earn in California?

Median $214500 from 1223 postings across 78 companies.

See salary data

What you'll do

  • Build automation for bare-metal provisioning, hardware validation, firmware upgrades, and cluster lifecycle management.
  • Develop tools using BMC and Redfish interfaces to assess hardware health and facilitate recovery workflows.
  • Manage and enhance NVIDIA NVL72 systems and BlueField-3 DPUs in cloud and on-premises environments.
  • Diagnose failures across servers, DPUs, GPUs, networking, Linux, and Kubernetes systems.
  • Convert recurring infrastructure issues into automated detection and repair systems.
  • Define validation and handoff criteria to ensure new capacity enters production safely and consistently.
  • Perform on-call duties, incident response, and root-cause analysis to implement permanent resolutions.

What we're looking for

  • 8+ years of experience building software for or operating production infrastructure with substantial hands-on bare-metal experience.
  • Strong Go or Python skills with a record of delivering production automation and services.
  • Direct experience working with BMC and Redfish for server provisioning, health inspection, power control, or fault diagnosis.
  • Practical experience working directly with NVIDIA GPU hardware, including NVL72 systems, and BlueField-3 or later DPUs.
  • Experience with Linux, firmware and driver management, network boot, and the server lifecycle from initial provisioning through repair.
  • Experience managing production reliability via on-call duties, incident handling, observability, and durable solutions.
  • Ability to debug failures across hardware, host operating systems, networking, and distributed services.
  • BS/MS in Computer Science or equivalent experience in a related field.

More like this

Similar roles

Senior Software Engineer, DGX Cloud Production Engineering

Nvidia

Santa Clara, CA 18 days ago
Kubernetes Python Go GitOps Terraform ArgoCD GPU Infrastructure Linux Containers Cloud Infrastructure Infrastructure Automation Distributed Systems Observability BMaaS VMaaS Managed Kubernetes Multi-cloud APIs Incident Response
8+ yrs exp

Senior Software Engineer, DGX Cloud Production Engineering

Nvidia

Remote (CA) +4 21 days ago
Kubernetes Python Go GitOps Terraform ArgoCD GPU Infrastructure Linux Containers Cloud Infrastructure Infrastructure Automation Distributed Systems Observability BMaaS VMaaS Managed Kubernetes Multi-cloud APIs Incident Response
8+ yrs exp Remote

Principal Software Engineer, DGX Cloud Production Engineering

Nvidia

Remote (Santa Clara, CA) 146 days ago $272,000–$431,250
Kubernetes Go Python Linux GitOps GPU Clusters Infrastructure Automation Distributed Systems Cloud Infrastructure APIs Observability SLOs Incident Response BMaaS VMaaS Multi-cloud High-Performance Computing
10+ yrs exp Remote

Senior Production Engineer

Nvidia

Remote (Santa Clara, CA) 9 days ago $184,000–$287,500
Kubernetes Python Go AWS Azure Google Cloud Terraform GitOps CI/CD vLLM SGLang PyTorch TensorRT-LLM NVIDIA Dynamo CUDA NCCL Linux Distributed Systems Argo CD SRE
8+ yrs exp Remote

Principal Software Engineer

Nvidia

Remote (Seattle, WA) +1 65 days ago $272,000–$431,250
Kubernetes Go C Linux Cloud Computing CI/CD Gitlab Argo Flux Container Orchestration Distributed Systems Unix Data Structures and Algorithms GPU DPU Confidential Computing
10+ yrs exp Remote