Site Reliability Engineer, Data Center Infrastructure

SpaceX

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Bastrop, TXHawthorne, CARedmond, WACape Canaveral, FLStarbase, TX
Posted
2 days ago
Freshness
Confirmed live today

Market check

Salary context

How this pay compares to similar roles

Similar $182k
$135k $237k
below market most similar roles pay here above market

This listing doesn't post a salary. Most similar roles pay $145,000–$219,237.

Based on 240 similar postings.

Employer

About SpaceX

SpaceX designs, manufactures, and launches advanced rockets and spacecraft with the mission of enabling humans to become a multi-planetary species. It operates the Falcon 9, Falcon Heavy, and Starship launch vehicles, as well as the Starlink satellite internet constellation.

SpaceX currently has 796 open roles on FindRole.

Listed pay typically runs $130,000–$165,000 across 555 roles with salary data.

Most-posted roles

View all roles at SpaceX

At a glance

TL;DR · Site Reliability Engineer, Data Center Infrastructure

SR. SITE RELIABILITY ENGINEER, DATA CENTER INFRASTRUCTURE joins the application software team to manage the compute, storage, and networking infrastructure supporting Starship, Starlink, Starshield, and Terafab. This role focuses on ensuring factory uptime and production scale by deploying, upgrading, and maintaining manufacturing systems. Key responsibilities include managing infrastructure as code, utilizing observability for platform health, performing capacity planning, and conducting sustainable incident response. The candidate will use Python, Linux, and infrastructure as code tools like Terraform, Ansible, or Puppet. They may also work with containers and virtualization such as Docker, Kubernetes, vSphere, QEMU, and KVM, alongside databases like Postgres and Clickhouse. The position solves critical reliability and scalability problems for mission-critical manufacturing systems across multiple production programs.

What does a Site Reliability Engineer earn in California?

Median $214000 from 72 postings across 20 companies.

See salary data

What you'll do

  • Deploy, upgrade, operate, and scale compute, storage, and networking for manufacturing systems.
  • Manage infrastructure as code and use observability tools to monitor platform health.
  • Design systems for reliability and scale while identifying and removing performance bottlenecks.
  • Perform proactive maintenance including capacity planning and lifecycle management to reduce toil.
  • Improve the full infrastructure lifecycle from initial design through deployment and continuous refinement.
  • Execute sustainable incident response and conduct blameless postmortems.
  • Provide high-quality technical support to manufacturing and engineering users.
  • Participate in on-call rotations and travel to sites for deployments and incident response.

What we're looking for

  • Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 7+ years of professional experience in SRE or DevOps in lieu of a degree.
  • 3+ years of experience with Python and Python-based development frameworks.
  • Experience with Linux operating systems.
  • Experience with compute, storage, and/or networking infrastructure in production (preferred).
  • Experience with Infrastructure as Code (Terraform, Ansible, Puppet, or similar) (preferred).
  • Experience with containers and virtualization (Docker, Kubernetes, vSphere, QEMU, KVM, etc.) (preferred).
  • Experience with databases and data modeling (Postgres, Clickhouse, etc.) (preferred).
  • Must be a U.S. citizen, national, lawful permanent resident, refugee, asylee, or eligible for Department of State authorization.

More like this

Similar roles

Site Reliability Engineer, HPC & Automation

SpaceX

Redmond, WA 102 days ago $125,000–$150,000
High Performance Computing Python Bash Linux Docker Kubernetes Terraform Ansible Puppet CI/CD Jenkins Prometheus Grafana MySQL PostgreSQL SQLite TCP/IP REST API NFS Cadence Synopsys Ansys Keysight Siemens
2+ yrs exp

Site Reliability Engineer, High Performance Computing

SpaceX

Hawthorne, CA 18 days ago $125,000–$160,000
Linux Infrastructure as Code Python Kubernetes Docker Terraform Ansible Prometheus Grafana Nagios Slurm PBS LSF PyTorch TensorFlow CUDA HIGH PERFORMANCE COMPUTING
2+ yrs exp

Site Reliability Engineer, Application Software

SpaceX

Hawthorne, CA 80 days ago $125,000–$160,000
Python Linux Docker Kubernetes Terraform Ansible MySQL ClickHouse JavaScript C# C++ Bazel Buck Make Infrastructure as Code Observability vSphere QEMU KVM