Senior Software Engineer, SRE and Production Engineering

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CA
Salary
$152,000–$241,500 / yr
Employment
Full-time
Posted
4 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $178k
This role $197k
$105k most similar roles pay here $256k

This role pays more than 71% of similar roles. Most pay $149,500–$206,450 — the shaded band above. At the midpoint, this role pays about $197k versus about $178k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1150 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 912 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, SRE and Production Engineering

Senior Software Engineer, SRE and Production Engineering - DGX Cloud joins the team building software and operational tooling to move GPU capacity from installed hardware into a production IaaS environment of BMaaS and VMaaS. You will build automation for bare-metal provisioning, hardware validation, firmware upgrades, and cluster lifecycle management while developing tools that interact with BMC and Redfish interfaces to monitor health and manage server states. The role involves managing NVL72 systems and BlueField-3 DPUs across cloud partner and on-premises environments. You will diagnose failures across servers, GPUs, networking, Linux, and Kubernetes to create automated detection and repair workflows. Required skills include Go or Python, experience with firmware management, network boot, and incident response. The work focuses on the technical challenge of making large-scale GPU infrastructure for AI workloads production-ready through robust automation and reliability engineering.

What does a Software Engineer earn in California?

Median $214000 from 830 postings across 69 companies.

See salary data

What you'll do

  • Build automation for bare-metal provisioning, hardware validation, firmware upgrades, and cluster lifecycle management.
  • Develop tools using BMC and Redfish interfaces to monitor hardware health and manage server states.
  • Manage and advance NVIDIA NVL72 systems and BlueField-3 DPUs across cloud and on-premises environments.
  • Diagnose failures across servers, GPUs, networking, Linux, and Kubernetes to create automated detection and repair workflows.
  • Define validation and handoff criteria to ensure new capacity enters production safely and consistently.
  • Perform on-call duties, incident response, and root-cause analysis to implement permanent infrastructure solutions.
  • Develop production automation and services using Go or Python for large-scale GPU infrastructure.

What we're looking for

  • 5+ years of experience building software for or operating production infrastructure with substantial hands-on bare-metal experience.
  • Strong Go or Python skills with a record of delivering production automation and services.
  • Direct experience with BMC and Redfish in server provisioning, health inspection, power management, or fault diagnosis.
  • Practical experience working directly with NVIDIA GPU hardware (NVL72 systems) and BlueField-3 or newer DPUs.
  • Experience with Linux, firmware/driver management, network boot, and the full server lifecycle from provisioning to repair.
  • Experience managing production reliability through on-call duties, incident response, observability, and durable solutions.
  • Ability to debug failures across hardware, host operating systems, networking, and distributed services.
  • BS/MS in Computer Science or equivalent experience in a practical setting.

More like this

Similar roles

Principal Software Engineer

Nvidia

Remote 57 days ago $272,000–$431,250
Kubernetes Go C Linux CI/CD GitLab Argo Flux Container Orchestration Distributed Systems Cloud Computing GPU DPU Confidential Computing
10+ yrs exp Remote

Senior Cloud Software Engineer

Nvidia

Remote 2 days ago $152,000–$241,500
Kubernetes AWS GCP Azure Go Python Rust C++ Java Distributed Systems Cloud-Native Data Management Storage Systems Performance Engineering Observability
5+ yrs exp Remote

Senior Site Reliability Engineer

Nvidia

Remote 15 days ago
Kubernetes Python Go Linux GitOps Terraform ArgoCD Prometheus Grafana OpenTelemetry ELK Stack Splunk PyTorch TensorRT-LLM CUDA NCCL vLLM SGLang Distributed Systems
8+ yrs exp Remote