Distinguished Engineer, Production Engineering, Data Center Automation

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
CASanta Clara, CANYSDWY
Salary
$320,000–$488,750 / yr
Posted
8 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $210k
This role $404k
$107k most similar roles pay here $530k

This role pays more than 99% of similar roles. Most pay $153,720–$266,659 — the shaded band above. At the midpoint, this role pays about $404k versus about $210k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Distinguished Engineer, Production Engineering, Data Center Automation

Distinguished Engineer, Production Engineering, Data Center Automation serves as a senior technical leader within the Production Engineering group focused on cluster operations for DGX Cloud GPU capacity. This role involves defining long-range technical strategies and architectural visions for cluster lifecycles, runtime delivery, and steady-state operability across on-premises, hyperscaler, and NeoCloud environments. The individual will build robust workflows, interfaces, and engineering collaborations to ensure infrastructure health, scalability, and availability. Key responsibilities include managing Kubernetes services, coordinating release readiness, and overseeing bare-metal infrastructure operations. The role requires extensive expertise in distributed systems, infrastructure platforms, and production environments. By integrating software engineering and system knowledge, the candidate will establish operating standards and technical directions to solve complex problems regarding the reliability and performance of large-scale GPU infrastructure for researchers and customers across diverse cloud provider locations.

What you'll do

  • Define the long-range technical strategy for operating DGX Cloud clusters across on-premise and multi-cloud environments.
  • Establish the architectural vision and operational guidelines for cluster lifecycle, runtime delivery, and steady-state operability.
  • Guide cross-organizational investments to improve production readiness, operational safety, and performance of GPU infrastructure.
  • Make high-impact technical decisions to coordinate platform, hardware, provider, and service teams in production environments.
  • Develop robust workflows and interfaces for Kubernetes services, bare-metal operations, and service-layer reliability.
  • Set operating standards and engineering guidelines to ensure DGX Cloud resources remain scalable and maintainable.
  • Drive the evolution of the production model to automate and streamline large-scale cluster operations.

What we're looking for

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience.
  • 18+ years of experience building and operating large-scale distributed systems, infrastructure platforms, or production environments.
  • Confirmed company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms.
  • Experience establishing operating models, architectural direction, and engineering standards across various technical domains and organizations.
  • Proven track record of leading large, cross-team technical efforts from concept through production to deliver measurable outcomes.
  • Expertise in Kubernetes service management and managing on-premise, bare-metal, and cloud infrastructure operations.
  • Experience developing automation and engineering interfaces that connect platform teams, infrastructure teams, and service owners.
  • Ability to define long-range technical strategy and architectural vision for cluster lifecycle and runtime delivery.

More like this

Similar roles

Controls Engineer, Data Center Engineering

Apple Inc

Sparks, NV 45 days ago
Ignition MQTT PLC DDC SCADA HMI Python BACnet Modbus OPC UA IEC 61850 IEC 61131-3 Structured Text Ladder Diagram Function Block Diagram Schneider Electric Modicon M580 Siemens Desigo Simatic S7 Johnson Controls Metasys Tridium Niagara SEL RTAC Woodward easYgen DEIF AGC GitHub GitLab Fault Detection and Diagnostics (FDD)
5+ yrs exp