Staff Site Reliability Engineer

Anduril Industries

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Costa Mesa, CA
Salary
$191,000–$253,000 / yr
Posted
31 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $188k
This role $222k
$137k most similar roles pay here $265k

This role pays more than 82% of similar roles. Most pay $161,278–$215,000 — the shaded band above. At the midpoint, this role pays about $222k versus about $188k for comparable roles.

Based on 238 similar postings.

Employer

About Anduril Industries

Anduril Industries is a defense technology company that builds advanced hardware and software systems for national security, including autonomous drones, surveillance systems, and the Lattice AI command platform.

Anduril Industries currently has 1697 open roles on FindRole.

Listed pay typically runs $146,000–$194,000 across 1504 roles with salary data.

Most-posted roles

View all roles at Anduril Industries

At a glance

TL;DR · Staff Site Reliability Engineer

Staff Site Reliability Engineer The Staff Site Reliability Engineer joins the CorpTech Platform team to establish the reliability architecture for internal business and manufacturing systems. This role focuses on building infrastructure that makes reliability a core platform property rather than an individual effort. You will design and operate observability platforms, manage deployment safety through canary analysis and automated rollbacks, define SLO frameworks, and lead incident response for complex multi-system failures. The position involves developing reliability patterns specifically for AI-enabled systems to monitor model behavior drift and ensure graceful fallbacks. Required skills include deep expertise in distributed systems, Kubernetes, cloud platforms like AWS, GCP, or Azure, and proficiency in Go, Python, or Rust. You will solve critical problems regarding production readiness, capacity planning, and cost optimization while ensuring the stability of internal enterprise software and manufacturing operations.

What does a Site Reliability Engineer earn in California?

Median $214000 from 54 postings across 15 companies.

See salary data

What you'll do

  • Design and manage the reliability architecture for production environments including observability, deployment systems, and incident management.
  • Build and operate an enterprise-scale observability platform featuring metrics, distributed tracing, and structured logging.
  • Develop release-safety mechanisms such as canary analysis, automated rollbacks, and progressive rollout systems.
  • Define and govern SLO frameworks to make reliability measurable for engineering teams and leadership.
  • Identify systemic risks and implement automation to eliminate entire classes of failures rather than patching symptoms.
  • Establish production-readiness standards and review processes to embed reliability into the software development lifecycle.
  • Lead incident response for complex, multi-system failures and drive durable improvements through post-incident analysis.
  • Develop reliability patterns specifically for AI-enabled systems, including monitoring for model behavior drift and degradation.

What we're looking for

  • 10+ years of experience in site reliability engineering, production engineering, infrastructure engineering, or a related discipline at an architecture or platform-wide scope.
  • Demonstrated experience designing and owning reliability infrastructure such as observability platforms, deployment systems, or incident management tooling at scale.
  • Deep technical fluency in distributed systems, Kubernetes, cloud platforms (AWS, GCP, or Azure), networking, and storage.
  • Proficiency in systems programming languages including Go, Python, Rust, or equivalent for building production infrastructure and automation.
  • Demonstrated experience defining SRE standards and influencing adoption across engineering teams without formal authority.
  • Track record of leading incident response for complex failures and converting findings into systemic infrastructure improvements.
  • Degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent practical experience.
  • U.S. Person status is required to access export controlled data.

More like this

Similar roles

Staff Site Reliability Engineer

CME Group

Chicago, IL 16 days ago $132,100$220,100
SRE Google Cloud GKE Terraform CloudFormation Chef Python Java Go Bash TypeScript Rust CI/CD Infrastructure as Code Prometheus OpenTelemetry Performance Testing Configuration Management
Hybrid

Staff DevOps Engineer

Anduril Industries

Costa Mesa, CA 29 days ago $191,000$253,000
SRE Kubernetes AWS GCP Azure Go Python Rust Prometheus Grafana OpenTelemetry Datadog Distributed Systems Canary Analysis Feature Flagging Capacity Planning Cost Optimization
10+ yrs exp

Staff Site Reliability Engineer

TransUnion

Chicago, IL +4 140 days ago $112,500$187,500
GCP Kubernetes CI/CD Datadog Prometheus Grafana PagerDuty Linux PostgreSQL MySQL Cloud SQL Bigtable Firestore Redis Terraform Pulumi Python Bash Go Infrastructure-as-Code
5+ yrs exp Hybrid

Staff Site Reliability Engineer

CME Group

Chicago, IL 42 days ago $132,100$220,100
Python Go Kubernetes GCP GKE Kafka Terraform ArgoCD Node.js Gemini Distributed Systems GitOps SRE
10+ yrs exp Hybrid

Staff Site Reliability Engineer

MongoDB

Bengaluru, India 52 days ago
Kubernetes Python Go AWS GCP Azure Multi-cloud Distributed Systems Virtual Machines Capacity Planning Incident Response SLO Observability Alerting Automation Networking Infrastructure Architecture
10+ yrs exp Hybrid

Staff Site Reliability Engineer

Circle

Remote (San Francisco, CA) 37 days ago $195,000$257,500
Kubernetes Terraform Pulumi Go Python CI/CD Infrastructure as Code Blockchain Distributed Systems SQL Helm SRE Chaos Engineering Cloud Networking DNS
6+ yrs exp Remote