Senior GPU Supercomputer Scheduler Engineer

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CARedmond, WA
Salary
$152,000–$241,500 / yr
Posted
29 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $206k
This role $197k
$140k most similar roles pay here $264k

This role pays less than 56% of similar roles. Most pay $177,250–$235,750 — the shaded band above. At the midpoint, this role pays about $197k versus about $206k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior GPU Supercomputer Scheduler Engineer

As a Senior GPU Supercomputer Scheduler Engineer on the Managed AI Research Superclusters team, you will design and develop new scheduling features and add-on services to improve GPU compute clusters across dimensions like resource usage fairness, occupancy, waste, resilience, performance, and power usage. You will build batch workload management and orchestration services while providing support to resolve scheduler issues for demanding deep learning and high-performance computing workloads. The role involves performing performance analysis of deep learning workflows, developing large-scale automation solutions, and conducting root cause analysis. Key technical requirements include proficiency in C/C++, Go, Python, and bash, along with experience in Linux environments, container technologies like Docker, Singularity, or Podman. You will also utilize knowledge of SLURM or K8s batch schedulers to manage complex multi-node GPU workloads and system co-design challenges.

What you'll do

  • Design and develop new scheduling features to improve GPU occupancy, resource fairness, and power usage.
  • Develop batch workload management and orchestration services for large multi-node GPU workloads.
  • Provide technical support to staff and end users to resolve complex batch scheduler issues.
  • Build and improve the ecosystem surrounding GPU-accelerated computing infrastructure.
  • Perform performance analysis and optimizations for deep learning workflows.
  • Develop large-scale automation solutions to streamline system operations.
  • Conduct root cause analysis and implement corrective actions for hardware and software problems.
  • Identify architectural directions and new approaches for AI workload scheduling.

What we're looking for

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field (or equivalent experience).
  • 5+ years of work experience.
  • Strong understanding of batch scheduling with experience in schedulers like SLURM or K8s batch schedulers.
  • Significant experience in systems programming languages including C/C++, Go, Python, and bash.
  • Established experience with Linux operating systems, environments, and tools.
  • Experience analyzing and tuning performance for a variety of AI workloads.
  • In-depth understanding of container technologies such as Docker, Singularity, and Podman.
  • Excellent communication, interpersonal, and customer collaboration skills.

More like this

Similar roles

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote

Senior System Software Engineer, GPU Performance

Nvidia

Remote 6 days ago $152,000$241,500
HPC C++ Python CUDA NCCL UCX NVSHMEM MPI Infiniband Ethernet RDMA PyTorch TensorFlow Kubernetes Docker SLURM Ansible Performance Engineering Parallel Programming
Remote