Senior High Performance Computing (HPC) Systems Engineer

SpaceX

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Hawthorne, CA
Salary
$165,000–$230,000 / yr
Posted
64 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $190k
This role $198k
$141k most similar roles pay here $240k

This role pays more than 61% of similar roles. Most pay $153,100–$227,875 — the shaded band above. At the midpoint, this role pays about $198k versus about $190k for comparable roles.

Based on 240 similar postings.

Employer

About SpaceX

SpaceX designs, manufactures, and launches advanced rockets and spacecraft with the mission of enabling humans to become a multi-planetary species. It operates the Falcon 9, Falcon Heavy, and Starship launch vehicles, as well as the Starlink satellite internet constellation.

SpaceX currently has 680 open roles on FindRole.

Listed pay typically runs $130,000–$165,000 across 442 roles with salary data.

Most-posted roles

View all roles at SpaceX

At a glance

TL;DR · Senior High Performance Computing (HPC) Systems Engineer

Sr. High Performance Computing (HPC) Systems Engineer will join the HPC team to support personnel and proprietary systems within a high-paced engineering environment. The role involves administering and managing HPC clusters, storage systems, and high-speed networks while providing application support across various engineering disciplines. Key responsibilities include installing and integrating Linux-based compute clusters and creating technical documentation for non-technical audiences. The candidate will utilize technologies including Kubernetes, Docker, Podman, Singularity, and configuration management tools like Puppet and Ansible. Technical expertise is required in Bash or Python scripting, cluster resource managers such as Slurm, PBS, or LSF, and monitoring tools like Prometheus, Grafana, and Nagios. The work involves supporting scientific and engineering computing tasks including CFD, FEA, large-scale AI training, and GPU usage with Cuda within a mission-critical infrastructure context.

What you'll do

  • Administer and manage HPC clusters, storage systems, and high-speed networks.
  • Provide technical application support to SpaceX employees across various engineering disciplines.
  • Install and integrate Linux-based compute clusters into the infrastructure.
  • Write instructional documentation to communicate complex technical ideas to non-technical audiences.
  • Automate recurring tasks and solve problems using scripting languages like Bash or Python.
  • Manage cluster resource managers such as Slurm, PBS, or LSF.
  • Deploy and maintain automated configuration management software like Puppet or Ansible.
  • Monitor and alert on system performance using tools like Prometheus, Grafana, or Nagios.

What we're looking for

  • Bachelor's degree in computer science, engineering, math, or a scientific discipline and 5+ years of systems engineering experience.
  • 7+ years of professional experience building software in lieu of a degree.
  • 5+ years of hands-on experience with client/server hardware, management tools, enterprise networking, virtualization, and security technologies.
  • Experience with Kubernetes.
  • 5+ years of professional experience building, deploying, and troubleshooting Linux systems.
  • Experience with scripting languages like Bash or Python to automate tasks.
  • Eligibility for access to classified material up to TS/SCI with Polygraph.
  • Must be a U.S. citizen, national, lawful permanent resident, refugee, or asylee per ITAR requirements.

More like this

Similar roles

Site Reliability Engineer

SpaceX

Hawthorne, CA 59 days ago $125,000$150,000
High Performance Computing Linux Windows Server Infiniband Python Bash Puppet Ansible Kubernetes Docker Networking Virtualization CFD FEA ANSYS StarCCM+
1+ yrs exp

Site Reliability Engineer, HPC & Automation

SpaceX

Redmond, WA 72 days ago $125,000$150,000
HPC Python Bash Linux Docker Kubernetes Terraform Ansible Puppet Prometheus Grafana CI/CD MySQL PostgreSQL SQLite TCP/IP Slurm LSF REST API NFS Cadence Synopsys
2+ yrs exp

System Software Engineer, HPC Performance

Nvidia

Champaign, IL +2 17 days ago $152,000$241,500
C C++ Python HPC Machine Learning Deep Learning Artificial Intelligence x86 ARM Linux Windows macOS Profiling Tools Cloud Computing
5+ yrs exp

HPC Systems Engineer, Modeling & Simulation

Anduril Industries

Costa Mesa, CA 92 days ago $132,000$198,000
HPC Linux Unix Python Bash Slurm MPI OpenMP CMake NFS NAS TCP/IP ParaView VisIt Cubit CTH ALE3D Sierra Data Workflows
5+ yrs exp

AI Systems Engineer, HPC

Amd

San Jose, CA 24 days ago $173,600$260,400
GPU HPC Kubernetes Python SLURM RoCEv2 KVM Ubuntu Shell Ansible Saltstack Terraform Prometheus Grafana Distributed ML LLMs 400G Networking