Software Engineer, Inference (AI Data Engineering)

SpaceX

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Palo Alto, CA
Salary
$135,000–$175,000 / yr
Posted
23 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $182k
This role $155k
$124k most similar roles pay here $236k

This role pays less than 63% of similar roles. Most pay $142,487–$221,000 — the shaded band above. At the midpoint, this role pays about $155k versus about $182k for comparable roles.

Based on 240 similar postings.

Employer

About SpaceX

SpaceX designs, manufactures, and launches advanced rockets and spacecraft with the mission of enabling humans to become a multi-planetary species. It operates the Falcon 9, Falcon Heavy, and Starship launch vehicles, as well as the Starlink satellite internet constellation.

SpaceX currently has 680 open roles on FindRole.

Listed pay typically runs $130,000–$165,000 across 442 roles with salary data.

Most-posted roles

View all roles at SpaceX

At a glance

TL;DR · Software Engineer, Inference (AI Data Engineering)

Software Engineer, Inference (AI Data Engineering) joins the application software team to maintain a high-performance AI inference platform serving internal models. This role involves designing and optimizing large-scale model serving systems end-to-end, including distributed infrastructure, load balancing, auto-scaling, and low-level GPU kernel work. The engineer will build reliable, high-throughput systems with features like quantization, speculative decoding, and custom tools for tracing and replaying issues across the full stack. Key responsibilities include benchmarking inference engines such as SGLang, vLLM, and TensorRT-LLM while developing CI/CD infrastructure for seamless deployment. Required technical skills include Rust or C++, Python, Go, gRPC, Docker, and Kubernetes. The role addresses the critical challenge of providing reliable, high-concurrency inference to power mission-critical applications, ensuring high availability and low tail latency for internal engineering goals.

What does a Software Engineer earn in California?

Median $214000 from 775 postings across 63 companies.

See salary data

What you'll do

  • Develop high-throughput inference systems to serve AI models across internal SpaceX platforms.
  • Architect and implement scalable distributed infrastructure including load balancing, auto-scaling, and batch scheduling.
  • Optimize model inference performance using GPU kernels, quantization, and speculative decoding techniques.
  • Build high-concurrency serving systems with 100% uptime and low tail latency.
  • Own end-to-end components such as request routing, SDK development, and rate limiting.
  • Benchmark and accelerate inference engines like SGLang, vLLM, and TensorRT-LLM.
  • Create custom tools for tracing, replaying, and resolving issues across the full stack.
  • Build robust CI/CD infrastructure for seamless deployment of endpoints and inference engine updates.

What we're looking for

  • Bachelor's degree in computer science, engineering, math, or a scientific discipline, or 2+ years of professional software experience.
  • Experience designing, implementing, and maintaining reliable and horizontally scalable distributed systems.
  • At least 1 year of experience in full stack or backend development with production systems.
  • At least 1 year of experience with Rust or C++.
  • Experience with LLM inference engines and serving frameworks like SGLang, vLLM, Triton, or TensorRT-LLM (preferred).
  • Deep low-level systems programming and optimizations including GPU kernels, quantization, and speculative decoding (preferred).
  • Proficiency in Python, Go, or similar languages; expert knowledge of gRPC; and experience with Docker/Kubernetes (preferred).
  • Must be a U.S. citizen, national, lawful permanent resident, or eligible for required export control authorizations.

More like this

Similar roles

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 136 days ago $184,000$287,500
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid

AI Inference Engineer

F5 Inc

San Jose, CA +1 15 days ago $176,600$265,000
vLLM TGI NVIDIA Triton Python C++ Rust TensorRT Llama.cpp Ollama Kubernetes Docker AWS GCP Azure CUDA TPUs MLOps
Hybrid

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp

Senior AI Inference Platform Engineer

Apple Inc

Seattle, WA 31 days ago $175,000$308,500
Python Go C++ Triton TensorRT-LLM vLLM Kubernetes Prometheus Grafana Splunk CI/CD Data Pipelines Performance Benchmarking Distributed Systems Capacity Planning
7+ yrs exp

Senior AI Inference Platform Engineer

Apple Inc

Seattle, WA 30 days ago $175,000$308,500
Python Go C++ Triton TensorRT-LLM vLLM Kubernetes Prometheus Grafana Splunk CI/CD Data Pipelines Performance Benchmarking Distributed Systems Capacity Planning
7+ yrs exp

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (CA) +4 15 days ago $152,000$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization Profiling Agile
5+ yrs exp Remote