Senior Site Reliability Engineer, HPC

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CAAustin, TXDurham, NC
Salary
$152,000–$241,500 / yr
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $183k
This role $197k
$134k most similar roles pay here $253k

This role pays more than 61% of similar roles. Most pay $151,612–$215,000 — the shaded band above. At the midpoint, this role pays about $197k versus about $183k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Site Reliability Engineer, HPC

The Senior Site Reliability Engineer - HPC joins the Compute Farm team to build and maintain a global services platform for high-performance computing. This role involves owning SRE solutions from design through implementation, ensuring seamless integration with HPC schedulers, storage, and network fabrics. The engineer will automate provisioning using Infrastructure as Code and configuration management across multi-cloud hybrid environments including AWS, GCP, and OCI. Key responsibilities include capacity planning, incident review, root cause analysis, and designing for failure with redundancy and strict change control. Required skills include proficiency in Python, Go, Perl, or Ruby, along with experience in Slurm, LSF, or Kubernetes. The candidate must possess expertise in CI/CD techniques, container management, log collection, and AIOps to manage large-scale infrastructure platforms while ensuring high quality of service for internal customers within a complex HPC environment.

What does a Site Reliability Engineer earn in California?

Median $214000 from 54 postings across 15 companies.

See salary data

What you'll do

  • Own SRE solutions end-to-end from design and implementation to operation and continuous improvement.
  • Automate provisioning across a multi-cloud hybrid environment using Infrastructure-as-Code and configuration management.
  • Design for high availability by implementing redundancy, failure domains, and strict change control.
  • Manage capacity planning and performance monitoring to ensure high Quality of Service for internal customers.
  • Participate in on-call rotations, incident reviews, and the production of high-quality root cause analysis reports.
  • Develop automated host lifecycle management and self-healing systems to reduce manual intervention.
  • Mentor other engineers and influence technical direction through design reviews and architecture documentation.

What we're looking for

  • Bachelor's degree in Computer Science or a related technical field (or equivalent experience).
  • 5+ years of professional experience building and supporting critical services.
  • Experience supporting large-scale HPC clusters using Slurm, LSF, or Kubernetes for setup, tuning, and troubleshooting.
  • Proficiency in modern CI/CD techniques and Infrastructure as Code (IaC) for managing services.
  • Experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability, and E2E observability.
  • 5+ years of coding and scripting experience in at least two high-level languages such as Python, Go, Perl, or Ruby.
  • Experience mentoring other engineers and influencing technical direction through design reviews and architecture documents.
  • Strong communication, documentation, and debugging skills to solve complex problems.

More like this

Similar roles

Site Reliability Engineer, HPC & Automation

SpaceX

Redmond, WA 72 days ago $125,000$150,000
HPC Python Bash Linux Docker Kubernetes Terraform Ansible Puppet Prometheus Grafana CI/CD MySQL PostgreSQL SQLite TCP/IP Slurm LSF REST API NFS Cadence Synopsys
2+ yrs exp

Senior Site Reliability Engineer

Salesforce

Remote (San Francisco, CA) 58 days ago $148,500$223,900
SRE Python Go Docker Kubernetes CI/CD Prometheus Grafana ELK Splunk Datadog Temporal Airflow Argo Workflows AWS GCP Linux Unix LLM Prompt Engineering
5+ yrs exp Remote

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 10 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 32 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp

Senior Site Reliability Engineer

The Federal Reserve

Boston, MA 15 days ago $140,000$210,900
AWS EKS Terraform Python Java Go Docker Ansible CI/CD IaC Linux Shell Scripting Prometheus Grafana CloudWatch OpenSearch Dynatrace Consul Vault S3 RDS Aurora Route 53 ELB ECR

Senior Site Reliability Engineer

Autodesk

Remote (ID) +1 28 days ago $117,000$209,330
Site Reliability Engineering AWS Kubernetes Python Go Java Infrastructure as Code CI/CD CloudWatch Splunk Datadog Dynatrace Bash PowerShell FedRAMP Distributed Systems Load Balancing DNS
7+ yrs exp Remote