Senior Manager, Validation and HPC

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
TXNCWACASC
Salary
$216,000–$345,000 / yr
Posted
14 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $200k
This role $280k
$122k most similar roles pay here $369k

This role pays more than 94% of similar roles. Most pay $172,900–$226,350 — the shaded band above. At the midpoint, this role pays about $280k versus about $200k for comparable roles.

Based on 239 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Manager, Validation and HPC

Senior Manager, Validation and HPC - NVIS joins the NVIDIA Infrastructure Specialists division to lead service engineering functions for Customer AI High-Performance Computing systems. This role involves supervising a multi-layered team of professionals to design, develop, install, and validate hardware and software while managing project planning and implementation. The manager will oversee system validation procedures, ensure operational reliability through custom scripts, and drive the deployment of network and server platforms. Key technical requirements include expertise in InfiniBand and Ethernet technologies, HPC storage, and cluster management tools like Slurm, Salt, and xCAT. Candidates must possess skills in OpenMP, MPI, NCCL, HPL, and GPU accelerators, alongside proficiency in Bash, Perl, or Python scripting. The role focuses on solving complex infrastructure challenges within the domain of AI supercomputing and large-scale data center projects.

What you'll do

  • Supervise the design, installation, and validation of hardware and software for Customer AI HPC systems.
  • Manage the planning, implementation, and performance tracking of large-scale HPC projects.
  • Configure and maintain high-performance AI network and server platforms using InfiniBand and Ethernet technologies.
  • Develop and deploy procedures for system validation, custom scripts, and testing to ensure operational reliability.
  • Lead the professional development and career growth goals of a multi-layered team of HPC service professionals.
  • Act as the primary domain authority for customers during planning calls and implementation phases.
  • Collaborate with internal teams to develop strategies for service quality and continuous improvement.

What we're looking for

  • Bachelor's degree in computer science, information systems, or a related field or equivalent experience.
  • Experience in IT, high-performance computing, or other related fields.
  • 3+ years of experience in a management or leadership role.
  • Expertise in HPC systems design configuration, planning, and storage.
  • Proficiency with low latency/high-bandwidth interconnect infrastructure including InfiniBand and Ethernet.
  • Expertise with cluster management tools (Slurm, Salt, xCAT) and programming fundamentals with scripting skills (Bash, Perl, Python).
  • Proficiency with shared and distributed memory parallelism (OpenMP, MPI, NCCL, HPL) and GPU accelerators.
  • Expertise in administering secure Linux/Unix operating systems and managing multi-vendor hardware/software.

More like this

Similar roles

Senior AI Compute Engineer

Nvidia

Remote 66 days ago $148,000$235,750
Linux Python Bash Ansible Kubernetes InfiniBand MPI SLURM LSF UGE HPC NCCL MLPerf HPL Lustre GPFS Networking GPU System Administration Automation
8+ yrs exp Remote

Senior AI Compute Engineer

Nvidia

Remote 151 days ago $148,000$235,750
Linux Python Bash Ansible Kubernetes SLURM LSF UGE InfiniBand MPI HPL NCCL MLPerf Lustre GPFS GPU Networking HPC System Administration Automation
8+ yrs exp Remote

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote

Senior Solutions Architect, AI Compute

Nvidia

Remote 21 days ago $184,000$287,500
Linux Python Bash Ansible Kubernetes SLURM LSF UGE InfiniBand MPI HPL NCCL MLPerf Lustre GPFS GPU HPC Networking System Administration Automation
8+ yrs exp Remote