Senior Site Reliability Engineer, AI Infrastructure

SpaceX

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Hawthorne, CA
Salary
$165,000–$265,000 / yr
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $182k
This role $215k
$133k most similar roles pay here $279k

This role pays more than 73% of similar roles. Most pay $147,200–$216,250 — the shaded band above. At the midpoint, this role pays about $215k versus about $182k for comparable roles.

Based on 240 similar postings.

Employer

About SpaceX

SpaceX designs, manufactures, and launches advanced rockets and spacecraft with the mission of enabling humans to become a multi-planetary species. It operates the Falcon 9, Falcon Heavy, and Starship launch vehicles, as well as the Starlink satellite internet constellation.

SpaceX currently has 704 open roles on FindRole.

Listed pay typically runs $130,000–$165,000 across 465 roles with salary data.

Most-posted roles

View all roles at SpaceX

At a glance

TL;DR · Senior Site Reliability Engineer, AI Infrastructure

Sr. Site Reliability Engineer, AI Infrastructure (Starshield) joins the team focused on Starshield's software and GPU infrastructure to support critical national security missions. This role involves designing, operating, and scaling infrastructure for large-scale AI clusters, including managing GPU/CPU deployments in Top Secret data centers and providing GPU as a service on bare metal and virtualized platforms. The engineer will develop automation for Kubernetes and AI clusters, manage distributed storage, databases, and monitoring systems while collaborating with AI engineers to build maintainable products. Required technical skills include Linux operating systems, Kubernetes, Terraform, Ansible, and containerization technologies. Candidates should be proficient in scripting with Bash or Python and have development experience in Python, C++, or Go. The role addresses the challenge of providing immediate access to intelligence data through a massive satellite constellation infrastructure.

What does a Site Reliability Engineer earn in California?

Median $214000 from 61 postings across 17 companies.

See salary data

What you'll do

  • Manage GPU and CPU infrastructure deployments within Top Secret data centers.
  • Provide "GPU as a service" for external customers on bare metal and virtualized platforms.
  • Design, validate, and productize solutions for large-scale AI clusters exceeding 100k GPUs.
  • Develop automation to deploy and manage on-premise Kubernetes, AI clusters, and operating systems.
  • Deploy and manage core infrastructure including databases, monitoring systems, and distributed storage.
  • Identify performance bottlenecks and create innovative solutions to ensure high system availability.
  • Lead the team toward technical excellence by making critical architectural and engineering decisions.

What we're looking for

  • Bachelor's degree in computer science, information systems, engineering, or 7+ years of experience in software, DevOps, or SRE.
  • 5+ years of professional experience with Linux operating systems.
  • 5+ years of experience with Kubernetes.
  • Experience with infrastructure tools like Terraform and Ansible.
  • Experience with containerization technologies such as OCI containers.
  • Development experience in Python, C++, or Go.
  • Ability to obtain and maintain a Top Secret Security Clearance.
  • Must be a U.S. citizen, permanent resident, or eligible for export control authorizations.

More like this

Similar roles

Site Reliability Engineer, AI Infrastructure

SpaceX

Washington, DC 3 days ago $125,000$160,000
Kubernetes Python Terraform Ansible Linux C++ Go Bash NVIDIA GPU TCP/IP Bare Metal Bazel Makefiles Monitoring Virtualization Data Modeling
1+ yrs exp

Senior Site Reliability Engineer

SpaceX

Redmond, WA 92 days ago $165,000$230,000
Kubernetes Linux Python Bash Terraform Ansible TCP/IP Bazel Makefiles Distributed Databases Monitoring Site Reliability Engineering DevOps
5+ yrs exp