Senior Datacenter Technical Program Manager, At-Scale AI Clusters

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$168,000–$258,750 / yr
Posted
6 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $211k
This role $213k
$157k most similar roles pay here $270k

This role pays more than 51% of similar roles. Most pay $180,412–$241,562 — the shaded band above. At the midpoint, this role pays about $213k versus about $211k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Datacenter Technical Program Manager, At-Scale AI Clusters

Senior Datacenter Technical Program Manager, At-Scale AI Clusters will join the Applied Systems Engineering Team to drive datacenter integration for next-generation AI supercomputing systems. This role involves managing the full lifecycle of AI systems at scale, including design, requirements definition, and system integration into datacenter environments. The manager will collaborate with engineers and architects to build and deploy large-scale GPU computing systems based on reference architectures while coordinating the design and fit-out of new builds involving power, cooling, and instrumentation. Key responsibilities include producing detailed documentation for end-to-end processes and communicating with leadership to resolve critical issues. Required expertise includes high-performance computing systems and GPU clusters in on-premises datacenters. Preferred technical skills include monitoring and instrumentation tools such as Prometheus, Grafana, Splunk, Modbus, and BACNet to solve complex infrastructure challenges.

What you'll do

  • Drive the integration of next-generation AI supercomputing systems into datacenter environments.
  • Manage the lifecycle of AI clusters from design and requirements definition to production support.
  • Lead the integration of new AI clusters with specific power, cooling, and instrumentation requirements.
  • Coordinate the design and fit-out of new datacenter builds with internal teams and external contractors.
  • Produce detailed documentation for end-to-end datacenter fit-out and integration processes.
  • Communicate with engineering leadership to prioritize and resolve critical issues for major customers.
  • Develop reference architectures to advise and support customers and partners.

What we're looking for

  • BS in Applied Science or Engineering (or equivalent experience).
  • 8+ years of overall experience.
  • Experience with high-performance computing systems and GPU clusters deployed in on-premises datacenters.
  • Strong teamwork and interpersonal skills to facilitate collaboration between multiple teams.
  • Understanding of datacenter design, including power and cooling technologies.
  • Expertise in system monitoring and instrumentation using tools like Prometheus, Grafana, Splunk, Modbus, and BACNet.
  • Experience supporting high-performance computing or deep learning for the engineering or academic research community.

More like this

Similar roles

Senior Technical Program Manager, AI Acceleration

Nvidia

Santa Clara, CA 72 days ago $168,000$258,750
AI LLM ASIC RTL Design Simulation Formal Verification System Architecture Scalable Infrastructure Data Pipelines Software Engineering Program Management Pre-silicon Workflows
10+ yrs exp

Senior Technical Program Manager, Data Center Engineering Operations

Nvidia

Santa Clara, CA +1 26 days ago $168,000$258,750
Program Management Engineering Operations Infrastructure Deployment Service Management Product Roadmap KPIs System Architecture Parallel Computing QA Best Practices Process Automation Software Engineering Principles
10+ yrs exp

Datacenter Software Program Manager

Qualcomm

San Diego, CA 18 days ago $154,400$231,600
x86 ARM Linux AI Gantt Charts Dashboards Program Management Semiconductor Data Center Software Stack Risk Management
4+ yrs exp

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote