Senior Production Engineer

Nvidia

Confirmed live today High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CA
Salary
$184,000–$287,500 / yr
Employment
Full-time
Posted
2 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $188k
This role $236k
$128k most similar roles pay here $305k

This role pays more than 84% of similar roles. Most pay $151,900–$223,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $188k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1391 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 1116 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Production Engineer

Senior Production Engineer - DGX Cloud As a Senior Production Engineer on the Production Engineering team, you will build and operate production software, automation, and tooling to ensure reliable, scalable, and safe services for inference and agentic workloads. You will manage control plane services, model deployments, and infrastructure across Kubernetes clusters in multi-cloud and on-premises environments. Your daily work involves improving endpoint availability, capacity management, and service health through GitOps, infrastructure as code, and automated workflows to replace manual tasks. You will define SLIs and SLOs while participating in incident response to troubleshoot failures across routing and model runtimes. Required skills include Python or Go programming, Linux, Kubernetes, and networking fundamentals. Experience with vLLM, SGLang, PyTorch, TensorRT-LLM, CUDA, Terraform, and Argo CD is highly valued for managing large-scale distributed systems and complex AI inference platforms.

What you'll do

  • Build and operate production software, automation, and tooling for control plane services and model deployments.
  • Improve the reliability of inference platforms through health validation, safer rollouts, and observability.
  • Manage endpoint availability, inference routing, and capacity to maintain performance during fluctuating demand.
  • Deploy and configure services across multi-cloud environments using infrastructure as code and GitOps.
  • Automate manual workflows for service enablement, model releases, and deprecation processes.
  • Define and instrument SLIs and SLOs to monitor system health and guide reliability improvements.
  • Participate in on-call rotations and troubleshoot failures across routing, runtimes, and hardware.
  • Develop durable fixes and automation to resolve recurring production issues.

What we're looking for

  • 8+ years of experience building or operating production services and large-scale distributed systems with hands-on automation.
  • Strong programming skills in Python, Go, or a comparable language for developing production tools.
  • Experience with infrastructure as code, configuration management, or GitOps to automate service deployments.
  • Strong knowledge of Linux, Kubernetes, containers, cloud infrastructure, distributed systems, and networking fundamentals.
  • Understanding of SRE principles including SLIs, SLOs, error budgets, incident response, and reducing operational toil.
  • Experience instrumenting services using metrics, logs, and traces to improve system reliability.
  • BS/MS in Computer Science or equivalent experience.
  • Familiarity with vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, or NCCL (preferred).

More like this

Similar roles

Senior Software Engineer, DGX Cloud Production Engineering

Nvidia

Remote (Santa Clara, CA) 11 days ago $184,000–$287,500
Kubernetes Python Go GitOps Terraform ArgoCD Linux Cloud Infrastructure GPU Infrastructure Distributed Systems Observability Infrastructure Automation BMaaS VMaaS
8+ yrs exp Remote

Senior Site Reliability Engineer

Nvidia

Remote 16 days ago
Kubernetes Python Go Linux GitOps Terraform ArgoCD Prometheus Grafana OpenTelemetry ELK Stack Splunk PyTorch TensorRT-LLM CUDA NCCL vLLM SGLang Distributed Systems
8+ yrs exp Remote

Principal Software Engineer, DGX Cloud Production Engineering

Nvidia

Remote (Santa Clara, CA) 139 days ago $272,000–$431,250
Kubernetes Go Python GitOps Linux GPU Clusters AI/ML Infrastructure Distributed Systems Infrastructure Automation APIs Observability SLOs Multi-cloud BMaaS VMaaS High-Performance Computing
10+ yrs exp Remote

DGX Cloud Automation Engineer

Nvidia

Santa Clara, CA 56 days ago
Golang Kubernetes KubeVirt Docker AWS CI/CD Infrastructure as Code Bare Metal PXE Boot DHCP DNS Firecracker KVM OpenStack Nutanix AHV Redhat OpenShift PaaS IaaS
8+ yrs exp

NCX Senior Engineer

Nvidia

Remote (Santa Clara, CA) +1 31 days ago $184,000–$287,500
Kubernetes Python Go PyTorch TensorFlow CUDA MLOps Terraform Ansible Prometheus Grafana InfiniBand RoCE Triton NeMo CI/CD Linux Distributed Systems NVIDIA Reference Architectures
8+ yrs exp Remote