Senior HPC AI Cluster Engineer

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$176,000–$276,000 / yr
Posted
22 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $210k
This role $226k
$161k most similar roles pay here $288k

This role pays more than 57% of similar roles. Most pay $172,981–$246,150 — the shaded band above. At the midpoint, this role pays about $226k versus about $210k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior HPC AI Cluster Engineer

As a Senior HPC AI Cluster Engineer on the Networking Clusters Solutions Infrastructure team, you will design, implement, and maintain large-scale HPC and AI clusters while managing Linux job scheduling and orchestration tools. You will develop continuous integration pipelines, create automation tools for infrastructure management, and provide troubleshooting from bare metal through the application level. The role involves collaborating with researchers and specialists to architect performance platforms and support R&D activities. Key technical requirements include experience with Slurm, K8s, Python, Bash, and configuration management tools like Jenkins, Ansible, or Puppet. You will work with Linux and Windows systems, InfiniBand and Ethernet networking protocols, and storage solutions such as Lustre, GPFS, and Weka.io. The role focuses on solving complex problems in accelerated computing, GPU compute, and large-scale system tuning for advanced AI workloads.

What you'll do

  • Design, implement, and maintain large-scale HPC and AI clusters with monitoring, logging, and alerting systems.
  • Manage Linux job scheduling and orchestration tools to handle complex workloads.
  • Develop and maintain continuous integration and delivery pipelines for infrastructure.
  • Create automation tools for deployment, management, and self-service resource consumption of large-scale environments.
  • Deploy and manage monitoring solutions across servers, network components, and storage systems.
  • Perform full-stack troubleshooting from bare metal and operating systems up to the application level.
  • Develop and document standard methodologies and technical resources for internal teams.
  • Support R&D activities and participate in proofs of concept for future infrastructure improvements.

What we're looking for

  • A degree in Computer Science, Engineering, or a related field (or equivalent experience).
  • 8+ years of experience in the field.
  • Knowledge of HPC and AI solution technologies including CPUs, GPUs, and high-speed interconnects.
  • Experience with job scheduling workloads and orchestration tools such as Slurm and K8s.
  • Excellent knowledge of Windows and Linux networking, internals, and OS level security.
  • Experience with multiple storage solutions like Lustre, GPFS, or Weka.io.
  • Proficiency in Python programming and bash scripting.
  • Familiarity with automation/configuration tools (Jenkins, Ansible, Puppet/Chef) and cloud platforms (AWS, Azure, Google Cloud).
  • Knowledge of CPU/GPU architecture, Kubernetes, container microservices, GPU-focused hardware/software (DGX, CUDA), and RDMA fabrics (preferred).

More like this

Similar roles

Senior Solutions Architect, AI Compute

Nvidia

Remote 21 days ago $184,000$287,500
Linux Python Bash Ansible Kubernetes SLURM LSF UGE InfiniBand MPI HPL NCCL MLPerf Lustre GPFS GPU HPC Networking System Administration Automation
8+ yrs exp Remote

Senior AI Compute Engineer

Nvidia

Remote 66 days ago $148,000$235,750
Linux Python Bash Ansible Kubernetes InfiniBand MPI SLURM LSF UGE HPC NCCL MLPerf HPL Lustre GPFS Networking GPU System Administration Automation
8+ yrs exp Remote

Senior AI Compute Engineer

Nvidia

Remote 151 days ago $148,000$235,750
Linux Python Bash Ansible Kubernetes SLURM LSF UGE InfiniBand MPI HPL NCCL MLPerf Lustre GPFS GPU Networking HPC System Administration Automation
8+ yrs exp Remote

Senior Manager, Validation and HPC

Nvidia

Remote (TX) +4 14 days ago $216,000$345,000
HPC InfiniBand Ethernet GPU MPI NCCL OpenMP Slurm Salt xCAT Python Bash Perl Linux Unix Ansible Puppet Lustre GPFS
10+ yrs exp Remote