Director, Site Reliability Engineering

Anduril Industries

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Costa Mesa, CA
Salary
$253,000–$336,000 / yr
Posted
25 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $191k
This role $294k
$128k most similar roles pay here $358k

This role pays more than 97% of similar roles. Most pay $160,259–$221,600 — the shaded band above. At the midpoint, this role pays about $294k versus about $191k for comparable roles.

Based on 238 similar postings.

Employer

About Anduril Industries

Anduril Industries is a defense technology company that builds advanced hardware and software systems for national security, including autonomous drones, surveillance systems, and the Lattice AI command platform.

Anduril Industries currently has 1697 open roles on FindRole.

Listed pay typically runs $146,000–$194,000 across 1504 roles with salary data.

Most-posted roles

View all roles at Anduril Industries

At a glance

TL;DR · Director, Site Reliability Engineering

Director, Site Reliability Engineering The Director, Site Reliability Engineering leads the SRE organization within the CorpTech Platform team, overseeing the reliability systems for software powering business and manufacturing operations. This leader is responsible for building the SRE team, establishing a portfolio-wide reliability strategy, and fostering shared accountability between SRE and software engineering teams. Key responsibilities include defining observability, deployment safety, incident management, capacity planning, and disaster recovery protocols while reducing operational toil through automation and architectural improvements. The role specifically addresses the challenges of maintaining high availability for business-critical systems, including AI-enabled production environments with non-deterministic behaviors. Required expertise includes distributed systems, cloud infrastructure, container orchestration, networking, storage, and deployment systems. The candidate will manage complex dependencies across enterprise domains such as ERP, MES, WMS, and CRM to ensure robust performance and production readiness for the company's internal infrastructure.

What you'll do

  • Build and lead the Site Reliability Engineering organization through hiring, coaching, and performance management.
  • Develop and execute a portfolio-wide reliability strategy based on operational risk and business criticality.
  • Establish shared accountability between SRE and software engineering via SLOs and production-readiness standards.
  • Define technical standards for observability, deployment safety, incident management, and disaster recovery.
  • Lead the organizational response to critical incidents and ensure durable post-incident improvements.
  • Reduce operational toil by converting recurring failure patterns into automated platform capabilities and architectural changes.
  • Create reliability metrics and operating reviews to guide investment decisions and identify systemic risks.
  • Establish reliable on-call systems that provide adequate staffing, tooling, and training to prevent burnout.

What we're looking for

  • 12+ years of experience in site reliability, production engineering, infrastructure engineering, software engineering, or a related discipline.
  • Experience leading engineering teams and managers including hiring, coaching, and performance management.
  • Proven success leading an SRE, infrastructure, or production engineering organization for business-critical systems at scale.
  • Deep technical fluency in distributed systems, cloud infrastructure, container orchestration, networking, storage, observability, and deployment systems.
  • Experience establishing SLOs, production-readiness standards, incident management, on-call practices, and operational maturity mechanisms.
  • Degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent practical experience.
  • U.S. Person status is required to access export controlled data.
  • Eligible to obtain and maintain a U.S. Secret security clearance (preferred).

More like this

Similar roles

Site Reliability Engineer

Anduril Industries

Waltham, MA 3 days ago $112,000$149,000
SRE DevOps Linux Python Bash Nix NixOS systemd Networking PagerDuty Observability Firmware
3+ yrs exp

Senior Engineering Manager, Site Reliability

Upstart

Remote (Canada) 55 days ago $195,300$270,400
Site Reliability Engineering Distributed Systems Cloud Infrastructure Datadog Grafana Prometheus OpenTelemetry Kubernetes AWS Incident Management Observability Service Level Objectives Postmortem Cloud Native Architecture
7+ yrs exp Remote

Principal Site Reliability Engineer

Oracle

Nashville, TN 57 days ago $84,900$209,500
Site Reliability Engineering Oracle Cloud Infrastructure Linux Windows Server Python PowerShell Bash Ansible Chef infrastructure-as-code Networking DNS Firewalls Load Balancing Certificates incident-management Observability Capacity Planning
3+ yrs exp

Principal Site Reliability Engineering Manager

Microsoft

15 days ago $142,800$274,800
Site Reliability Engineering Distributed Systems Incident Management Disaster Recovery Automation Telemetry Security Compliance Infrastructure Network Engineering Systems Administration
5+ yrs exp

Manager, Site Reliability Engineering

Okta Inc

San Francisco, CA 43 days ago $204,000$306,000
AWS Kubernetes Terraform CI/CD Grafana Splunk APM DevOps SaaS Cloud-native Architecture Containerization Edge Networking SDLC
3+ yrs exp Hybrid

Manager, Site Reliability Engineering

Oracle

Reston, VA +1 50 days ago $102,000$234,600
Site Reliability Engineering Incident Response Automation Monitoring Provisioning Decommissioning Scalability Data Collection
6+ yrs exp