Distinguished Engineer, Production Engineering, Cluster Management

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CAOR
Salary
$320,000–$488,750 / yr
Posted
8 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $241k
This role $404k
$151k most similar roles pay here $525k

This role pays more than 99% of similar roles. Most pay $186,825–$295,500 — the shaded band above. At the midpoint, this role pays about $404k versus about $241k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Distinguished Engineer, Production Engineering, Cluster Management

Distinguished Engineer, Production Engineering, Cluster Management serves as a senior technical leader within the Production Engineering organization focusing on DGX Cloud GPU capacity. This role involves defining long-range technical strategies and architectural directions for cluster lifecycle management, runtime delivery, and steady-state operability across local data centers, hyperscalers, and NeoCloud environments. The individual will build durable workflows, APIs, and automation frameworks to ensure infrastructure health, scalability, and reliability while reducing manual intervention. Key responsibilities include establishing operating standards for Kubernetes production services, hardware readiness, and service-layer reliability. Required expertise includes extensive experience in distributed systems, Linux, networking, and containers, alongside proficiency in programming languages like Python or Go. The role addresses the complex challenge of integrating diverse infrastructure components into a unified, automated production system to ensure high-performance availability for large-scale GPU resources.

What you'll do

  • Define the long-range technical strategy for operating DGX Cloud clusters across local data centers and hyperscaler environments.
  • Establish architectural directions and operating standards for cluster lifecycle, runtime delivery, and steady-state operability.
  • Develop durable workflows, interfaces, and engineering handshakes between Kubernetes services, hardware providers, and bare-metal infrastructure.
  • Build and evolve automation, APIs, and readiness gates to move new capacity into stable production environments.
  • Implement operating models that reduce manual input while increasing consistency, traceability, and release safety.
  • Identify recurring operational friction points and convert them into durable improvements in software and process interfaces.
  • Lead cross-organizational investments to improve production readiness, performance, and reliability across the DGX Cloud estate.
  • Raise engineering standards for scalability and resilience through architecture reviews and technical leadership.

What we're looking for

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience.
  • 18+ years of experience building and operating large-scale distributed systems, infrastructure platforms, or production environments.
  • Confirmed company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms.
  • Consistent track record of defining operating models, architectural direction, and engineering standards across multiple technical domains.
  • Experience leading large, cross-team technical efforts from concept through production while aligning collaborators and delivering measurable outcomes.
  • Deep experience with Kubernetes-based production systems, infrastructure automation, or distributed systems operations.
  • Strong software engineering skills in languages such as Python, Go, or similar low-level programming languages.
  • Deep understanding of distributed systems, Linux, networking, containers, and production reliability concerns.

More like this

Similar roles

HPC Infrastructure & Cluster Engineer

General Dynamics

Springfield, VA 9 days ago $119,850$162,150
Linux Run:AI OpenShift Kubernetes InfiniBand SLURM Python Bash SAN HPC Bare-metal Parallel File Systems Cluster Administration Infrastructure Optimization Network Management Storage Area Network Container Orchestration
5+ yrs exp

Distinguished Engineer

Capital One Financial

McLean, VA +2 121 days ago $269,100$307,200
AWS Microsoft Azure Google Cloud Python Java Go JavaScript TypeScript Swift Machine Learning Cloud Computing Platform Engineering Observability Scalability
7+ yrs exp

Distinguished Engineer

Elevance Health

Atlanta, GA +2 52 days ago
AI LLM Agentic Systems Distributed Systems Python Java Microservices Event-Driven Architecture APIs Cloud Platforms CI/CD DevOps SDLC
10+ yrs exp Hybrid