Staff Site Reliability Engineer

Anduril Industries

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Costa Mesa, CA
Salary
$191,000–$253,000 / yr
Posted
8 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $190k
This role $222k
$136k most similar roles pay here $266k

This role pays more than 80% of similar roles. Most pay $160,962–$219,012 — the shaded band above. At the midpoint, this role pays about $222k versus about $190k for comparable roles.

Based on 240 similar postings.

Employer

About Anduril Industries

Anduril Industries is a defense technology company that builds advanced hardware and software systems for national security, including autonomous drones, surveillance systems, and the Lattice AI command platform.

Anduril Industries currently has 850 open roles on FindRole.

Listed pay typically runs $166,000–$220,000 across 821 roles with salary data.

Most-posted roles

View all roles at Anduril Industries

At a glance

TL;DR · Staff Site Reliability Engineer

Staff Site Reliability Engineer As part of the CorpTech Platform team, you will establish the reliability architecture for systems powering business and manufacturing operations. You will not simply respond to tickets; instead, you will design and operate the observability platform, manage deployment infrastructure including canary analysis and automated rollbacks, and define SLO frameworks to make reliability measurable. Your work involves identifying systemic risks, establishing production-readiness standards, and developing reliability patterns for AI-enabled systems to monitor model behavior drift. You will utilize Go, Python, or Rust to build production tools while managing Kubernetes, cloud platforms like AWS, GCP, or Azure, and complex networking and storage layers. This role solves the challenge of scaling infrastructure for internal corporate systems, ensuring that high-growth business operations remain stable through automated mechanisms rather than manual intervention.

What does a Site Reliability Engineer earn in California?

Median $214000 from 56 postings across 13 companies.

See salary data

What you'll do

  • Design the reliability architecture for production environments including observability, deployment systems, and incident management.
  • Build and operate an enterprise-scale observability platform featuring metrics, distributed tracing, and structured logging.
  • Develop release-safety mechanisms such as canary analysis, automated rollbacks, and progressive rollout systems.
  • Define and govern SLO frameworks to make reliability measurable for engineering teams and leadership.
  • Identify systemic risks and drive infrastructure investments to eliminate entire classes of failures.
  • Establish production-readiness standards and review processes for all software development lifecycles.
  • Lead incident response for complex multi-system failures and implement durable, systemic improvements.
  • Develop reliability patterns specifically for AI-enabled systems including monitoring for model behavior drift.

What we're looking for

  • 10+ years of experience in site reliability, production engineering, infrastructure engineering, or a related discipline at an architecture or platform-wide scope.
  • Demonstrated experience designing and owning reliability infrastructure such as observability platforms, deployment systems, or incident management tooling at scale.
  • Deep technical fluency in distributed systems, Kubernetes, cloud platforms (AWS, GCP, or Azure), networking, and storage.
  • Proficiency in systems programming languages including Go, Python, Rust, or equivalent for building production infrastructure and automation.
  • Demonstrated experience defining SRE standards and influencing adoption across engineering teams without formal authority.
  • Track record of leading incident response for complex failures and converting findings into systemic infrastructure improvements.
  • Degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent practical experience.
  • U.S. Person status is required to access export controlled data; eligibility to obtain a U.S. Secret security clearance (preferred).

More like this

Similar roles

Staff Site Reliability Engineer

CME Group

Chicago, IL 63 days ago $132,100–$220,100
Python Go Kubernetes GCP GKE Kafka Terraform ArgoCD Node.js Gemini Distributed Systems GitOps SRE
10+ yrs exp Hybrid

Staff Site Reliability Engineer, Ads

Reddit

San Francisco, CA 147 days ago $217,000–$303,900
Site Reliability Engineering Distributed Systems Go Python Kubernetes Prometheus Grafana OpenTelemetry Envoy Kafka ClickHouse Cassandra Redis Linux SLOs CDN Cloud Infrastructure Traffic Engineering
8+ yrs exp

Staff Site Reliability Engineer

Okta Inc

Washington, DC 63 days ago $174,000–$238,000
AWS GCP Kubernetes Terraform Helm Go Python PostgreSQL Redis OpenSearch MySQL Cassandra CI/CD GitOps ArgoCD DNS TLS FedRAMP

Principal Site Reliability Engineer

Nvidia

Santa Clara, CA 24 days ago $248,000–$396,750
Kubernetes Distributed Systems Python Go Terraform AWS Azure GCP OpenTelemetry infrastructure-as-code Linux TypeScript JavaScript Java Crossplane AWS CDK CloudFormation AI/ML Platforms High-Performance Computing
10+ yrs exp Hybrid

Staff Site Reliability Engineer

Circle

Remote (San Francisco, CA) 58 days ago $195,000–$257,500
Kubernetes Terraform Pulumi Go Python CI/CD Infrastructure as Code Blockchain Distributed Systems SQL Helm SRE Chaos Engineering Cloud Networking DNS
6+ yrs exp Remote

Senior Staff Site Reliability Engineer

Early Warning Services

Scottsdale, AZ +3 2 days ago $150,000–$200,000
Site Reliability Engineering AWS CI/CD Infrastructure as Code Linux Unix Distributed Systems Monitoring Logging Tracing Azure GCP Oracle Cloud Infrastructure
10+ yrs exp Hybrid