Site Reliability Engineer, HPC & Automation
SpaceX
Quick summary
Market check
How this pay compares to similar roles
This role pays more than 61% of similar roles. Most pay $151,612–$215,000 — the shaded band above. At the midpoint, this role pays about $197k versus about $183k for comparable roles.
Based on 240 similar postings.
Employer
Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing
Nvidia currently has 896 open roles on FindRole.
Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.
Most-posted roles
At a glance
The Senior Site Reliability Engineer - HPC joins the Compute Farm team to build and maintain a global services platform for high-performance computing. This role involves owning SRE solutions from design through implementation, ensuring seamless integration with HPC schedulers, storage, and network fabrics. The engineer will automate provisioning using Infrastructure as Code and configuration management across multi-cloud hybrid environments including AWS, GCP, and OCI. Key responsibilities include capacity planning, incident review, root cause analysis, and designing for failure with redundancy and strict change control. Required skills include proficiency in Python, Go, Perl, or Ruby, along with experience in Slurm, LSF, or Kubernetes. The candidate must possess expertise in CI/CD techniques, container management, log collection, and AIOps to manage large-scale infrastructure platforms while ensuring high quality of service for internal customers within a complex HPC environment.
What does a Site Reliability Engineer earn in California?
Median $214000 from 54 postings across 15 companies.
What you'll do
What we're looking for
More like this
SpaceX
Salesforce
Nvidia
Nvidia
The Federal Reserve
Autodesk