Site Reliability Engineer, HPC & Automation

SpaceX

Confirmed live yesterday Trusted

Quick summary

Work type
On-site
Location
Redmond, WA
Salary
$125,000–$150,000 / yr
Posted
72 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $184k
This role $138k
$113k most similar roles pay here $241k

This role pays less than 85% of similar roles. Most pay $152,496–$216,250 — the shaded band above. At the midpoint, this role pays about $138k versus about $184k for comparable roles.

Based on 238 similar postings.

Employer

About SpaceX

SpaceX designs, manufactures, and launches advanced rockets and spacecraft with the mission of enabling humans to become a multi-planetary species. It operates the Falcon 9, Falcon Heavy, and Starship launch vehicles, as well as the Starlink satellite internet constellation.

SpaceX currently has 680 open roles on FindRole.

Listed pay typically runs $130,000–$165,000 across 442 roles with salary data.

Most-posted roles

View all roles at SpaceX

At a glance

TL;DR · Site Reliability Engineer, HPC & Automation

Site Reliability Engineer — HPC & Automation (Silicon Engineering) joins the Silicon Engineering team to design, operate, scale, and automate high performance computing infrastructure used to develop chips for a large satellite constellation. The role involves deploying and maintaining clusters, managing infrastructure as code, and operating continuous integration pipelines and build systems. You will collaborate with cross-disciplinary teams to create automated turnkey solutions for silicon simulation workflows while identifying and eliminating performance bottlenecks through measurement. Required skills include proficiency in Linux, Bash, and Python. Preferred qualifications include experience with Docker, Kubernetes, Slurm, Terraform, Ansible, and Prometheus. The role also values knowledge of TCP/IP networking, REST APIs, and enterprise storage automation. This position specifically addresses the technical challenge of accelerating design iterations, simulations, and regression turnaround times for critical silicon hardware components.

What does a Site Reliability Engineer earn?

Median $186200 from 131 postings across 36 companies.

See salary data

What you'll do

  • Deploy, upgrade, operate, maintain, and scale a suite of clusters and services.
  • Develop automated, full turnkey solutions for silicon simulation workflows to accelerate project timelines.
  • Manage infrastructure as code using modern tools to monitor cluster and system health.
  • Operate continuous integration pipelines, build systems, and manage version control across the environment.
  • Identify and eliminate performance bottlenecks through measurement and creative engineering.
  • Automate enterprise and networked storage systems for high-performance computing environments.

What we're looking for

  • Bachelor's degree in computer science, information systems, or an engineering discipline.
  • 2+ years of professional experience in system administration, high performance computing, or site reliability engineering.
  • 1+ years of development experience with Bash, Python, and/or other programming languages.
  • 1+ years of experience with Linux operating systems.
  • Experience with infrastructure as code (IaC) tools like Terraform, Ansible, or Puppet.
  • Familiarity with containerization technologies such as Docker and Kubernetes.
  • Knowledge of high performance computing, workload managers, and networking protocols like TCP/IP.
  • Must be a U.S. citizen, permanent resident, or eligible for required export control authorizations.

More like this

Similar roles

Site Reliability Engineer

Morgan Stanley

Alpharetta, GA 22 days ago
Python Shell Scripting Perl Ruby Java C# AWS Azure Jenkins Splunk DB2 Oracle Sybase Autosys Linux Unix Windows Agile Scrum Web Services MQ
5+ yrs exp

Site Reliability Engineer

Berkeley Research Group

Remote 77 days ago $130,000$160,000
Azure Kubernetes CI/CD GitHub Actions GitLab CI Golang Ruby Python AWS GCP Infrastructure as Code Datadog OpsGenie PagerDuty SRE Incident Management
5+ yrs exp Remote

Site Reliability Engineer

Balyasny Asset Management

Warsaw, Poland 78 days ago
Prometheus Grafana Loki Tempo OTEL Kubernetes Docker AWS Python Bash Go CI/CD DevOps SRE Agile
5+ yrs exp

Site Reliability Engineer

Booz Allen Hamilton

McLean, VA 8 days ago $86,800$198,000
AWS Kubernetes Terraform Ansible CI/CD Python Bash PowerShell Docker Prometheus Grafana Loki Elasticsearch Kibana GitLab GitHub CloudFormation OpenShift Jenkins REST JSON YAML XML Agile
6+ yrs exp