Site Reliability Engineer, AI Infrastructure

SpaceX

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Palo Alto, CAHawthorne, CARedmond, WAWashington, DC
Salary
$125,000–$160,000 / yr
Posted
3 days ago
Freshness
Confirmed live today

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $181k
This role $142k
$112k most similar roles pay here $243k

This role pays less than 80% of similar roles. Most pay $145,000–$216,675 — the shaded band above. At the midpoint, this role pays about $142k versus about $181k for comparable roles.

Based on 240 similar postings.

Employer

About SpaceX

SpaceX designs, manufactures, and launches advanced rockets and spacecraft with the mission of enabling humans to become a multi-planetary species. It operates the Falcon 9, Falcon Heavy, and Starship launch vehicles, as well as the Starlink satellite internet constellation.

SpaceX currently has 709 open roles on FindRole.

Listed pay typically runs $130,000–$165,000 across 461 roles with salary data.

Most-posted roles

View all roles at SpaceX

At a glance

TL;DR · Site Reliability Engineer, AI Infrastructure

Site Reliability Engineer, AI Infrastructure (Starshield) joins the team focused on software and GPU infrastructure to support critical national security missions. This role involves designing, operating, and scaling infrastructure for large-scale AI clusters, specifically managing GPU/CPU deployments in Top Secret data centers. You will develop automation for on-premise Kubernetes and AI clusters, manage distributed storage and monitoring systems, and collaborate with engineers to create maintainable software products. The position requires proficiency in Linux operating systems, containerization technologies like OCI containers, and infrastructure tools such as Terraform or Ansible. Candidates should possess scripting skills in Bash or Python, along with development experience in Python, C++, or Go. Technical requirements include knowledge of TCP/IP networking, distributed databases, and NVIDIA GPU deployment stacks to solve complex problems regarding high availability for government satellite data systems.

What does a Site Reliability Engineer earn in California?

Median $214000 from 60 postings across 16 companies.

See salary data

What you'll do

  • Manage GPU and CPU infrastructure deployments within Top Secret data centers.
  • Provide "GPU as a service" for external customers on bare metal and virtualized platforms.
  • Design, validate, and productize solutions for large-scale AI clusters exceeding 100k GPUs.
  • Develop automation to deploy and manage on-premise Kubernetes, AI clusters, and operating systems.
  • Deploy and manage core infrastructure including databases, monitoring systems, and distributed storage.
  • Implement monitoring and alerting systems to ensure high availability of critical services.
  • Identify system bottlenecks and create innovative solutions to improve overall reliability.

What we're looking for

  • Bachelor's degree in computer science, information systems, engineering, or 3+ years of professional experience in site reliability engineering or DevOps.
  • At least 1 year of professional experience with Linux operating systems.
  • Experience with infrastructure tools such as Terraform and Ansible.
  • Experience with containerization technologies like OCI containers and Kubernetes.
  • Experience scripting in Bash, Python, or other similar languages.
  • Development experience in Python, C++, or Go.
  • Must be a U.S. citizen, national, lawful permanent resident, or eligible for export control authorizations.
  • Security Clearance as a condition of employment.

More like this

Similar roles

Site Reliability Engineer, AI Infrastructure

SpaceX

Washington, DC 10 days ago $125,000–$160,000
Kubernetes Python Terraform Ansible Linux C++ Go Bash NVIDIA GPU TCP/IP Bare Metal Bazel Makefiles Monitoring Virtualization Data Modeling
1+ yrs exp