Site Reliability Engineer (High Performance Computing)

SpaceX

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Hawthorne, CA
Salary
$125,000–$160,000 / yr
Posted
10 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $183k
This role $142k
$112k most similar roles pay here $243k

This role pays less than 79% of similar roles. Most pay $149,300–$216,675 — the shaded band above. At the midpoint, this role pays about $142k versus about $183k for comparable roles.

Based on 240 similar postings.

Employer

About SpaceX

SpaceX designs, manufactures, and launches advanced rockets and spacecraft with the mission of enabling humans to become a multi-planetary species. It operates the Falcon 9, Falcon Heavy, and Starship launch vehicles, as well as the Starlink satellite internet constellation.

SpaceX currently has 654 open roles on FindRole.

Listed pay typically runs $130,000–$165,000 across 468 roles with salary data.

Most-posted roles

View all roles at SpaceX

At a glance

TL;DR · Site Reliability Engineer (High Performance Computing)

Site Reliability Engineer (High Performance Computing) joins the HPC team to manage a shared compute platform used for vehicle and structures simulation, machine learning, AI inference, and other critical engineering tasks. This role focuses on implementing an SRE model by reducing toil through automation, improving observability, and establishing sustainable incident processes. The engineer will manage node lifecycles using infrastructure as code, oversee storage and compute resources, perform capacity planning, and build monitoring for both administrators and end users. Key technical requirements include experience with Linux operating systems, production infrastructure, and configuration management tools like Ansible, Puppet, or Terraform. Candidates should be proficient in Python or similar languages for automation, and have familiarity with containers, Kubernetes, Prometheus, Grafana, and high-performance storage. The role addresses the challenge of providing reliable, scalable compute services for complex engineering workflows.

What does a Site Reliability Engineer earn in California?

Median $214000 from 59 postings across 16 companies.

See salary data

What you'll do

  • Manage node lifecycles using infrastructure as code for OS images, firmware, and configuration management.
  • Build observability systems to monitor cluster health for operators and job performance for end users.
  • Develop automation scripts and software to reduce manual toil in production operations.
  • Perform sustainable incident response and conduct blameless postmortems during on-call rotations.
  • Manage compute and storage resources to ensure high availability for the HPC platform.
  • Lead capacity planning by translating user requirements into concrete plans for future infrastructure needs.
  • Maintain and optimize the entire ecosystem of Linux machines, storage, and user-facing applications.

What we're looking for

  • Bachelor's degree in computer science, engineering, math, or a scientific discipline, or 2+ years of professional experience operating production infrastructure.
  • 2+ years of experience with Linux operating systems in production.
  • 2+ years of experience operating production infrastructure including monitoring, debugging, and repairing.
  • Eligibility for access to classified material up to TS/SCI with polygraph.
  • Must be a U.S. citizen, national, lawful permanent resident, or eligible for required export control authorizations.
  • 2+ years of professional experience in SRE, DevOps, or production infrastructure engineering (preferred).
  • Experience with monitoring and alerting tools like Prometheus, Grafana, or Nagios (preferred).
  • Experience with configuration management or infrastructure as code such as Ansible, Puppet, or Terraform (preferred).

More like this

Similar roles

Site Reliability Engineer, HPC & Automation

SpaceX

Redmond, WA 94 days ago $125,000–$150,000
HPC Python Bash Linux Docker Kubernetes Terraform Ansible Puppet Prometheus Grafana CI/CD MySQL PostgreSQL SQLite TCP/IP Slurm LSF REST API NFS Cadence Synopsys
2+ yrs exp