Senior Production Engineer

Anduril Industries

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Costa Mesa, CAWashington, DC
Salary
$166,000–$220,000 / yr
Posted
105 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $158k
This role $193k
$115k most similar roles pay here $231k

This role pays more than 82% of similar roles. Most pay $138,804–$177,430 — the shaded band above. At the midpoint, this role pays about $193k versus about $158k for comparable roles.

Based on 240 similar postings.

Employer

About Anduril Industries

Anduril Industries is a defense technology company that builds advanced hardware and software systems for national security, including autonomous drones, surveillance systems, and the Lattice AI command platform.

Anduril Industries currently has 1697 open roles on FindRole.

Listed pay typically runs $146,000–$194,000 across 1504 roles with salary data.

Most-posted roles

View all roles at Anduril Industries

At a glance

TL;DR · Senior Production Engineer

The Senior Production Engineer joins the SRE team to manage reliability and infrastructure for cloud deployments across multiple production environments. This role focuses on diagnosing and fixing stability vulnerabilities in core platform services that cause cascading failures within multi-tenant systems. The engineer will write production Go code to implement critical resilience patterns, such as leader election, circuit breakers, and failure domain isolation, while designing multi-replica support for single-instance services. Key responsibilities include tracing failures across service boundaries, collaborating on contract testing, and performing infrastructure tasks using Terraform and Kubernetes. The role requires expertise in distributed systems, consensus, replication, and gRPC architectures. Candidates should be proficient in Go and potentially Rust, with experience in observability platforms and solving complex reliability problems within live production environments rather than building greenfield projects.

What you'll do

  • Diagnose and fix stability vulnerabilities in core platform services causing cascading failures.
  • Implement resilience patterns like leader election and circuit breakers directly in Go service code.
  • Design multi-replica support for services currently operating as single instances.
  • Trace cascading failures across service boundaries to identify and implement root-cause fixes.
  • Improve the observability platform to enhance overall service stability.
  • Perform infrastructure updates using Terraform and Kubernetes to support service fixes.
  • Collaborate with service owners on contract testing and upgrade validation.

What we're looking for

  • Must possess production-quality Go programming skills for modifying core platform services.
  • Must have practical experience with distributed systems including leader election, consensus, replication, and failure modes.
  • Must have knowledge of Kubernetes to understand how services run.
  • Must be able to debug complex systems by tracing cascading failures across service boundaries.
  • Must have 4+ years of experience in SRE, platform engineering, or backend development roles.
  • Must be a U.S. Person due to access requirements for export-controlled information or facilities.
  • Must be eligible to obtain and maintain an active U.S. Secret security clearance.
  • Experience with Rust, gRPC, HashiCorp Consul, FedRAMP/IL5, or ArgoCD is preferred.

More like this

Similar roles

Senior Site Reliability Engineer

Anduril Industries

Costa Mesa, CA +1 81 days ago $166,000$220,000
Kubernetes AWS Azure Terraform Python Go Rust C++ Docker Helm ArgoCD CI/CD Linux KubeVirt qemu Infrastructure as Code SRE Networking Cybersecurity
6+ yrs exp

Senior Site Reliability Engineer, Infra Ops

Circle

Remote (San Francisco, CA) 28 days ago $152,500$205,000
Kubernetes Terraform Go Python JavaScript TypeScript CI/CD Infrastructure as Code Distributed Systems Cloud Infrastructure Observability GitOps SRE DevOps Networking Security incident-management Capacity Planning
5+ yrs exp Remote

Senior Production Support Engineer

Marqeta

Remote (Oakland, CA) 25 days ago $86,500$108,100
SQL Datadog Jira Salesforce cURL HTTP API Incident Management Escalation Management Root Cause Analysis Automation AI Tools Tokenization 3DS
8+ yrs exp Remote

Senior Production Software Engineer

Anduril Industries

Lexington, MA 3 days ago $191,000$253,000
Python MATLAB C/C++ JavaScript Linux Git SQL Networking Computer Vision System Architecture Automated Test Systems Hardware-Software Integration Data Analysis Nix
8+ yrs exp

Senior Engineer

GEICO

Bethesda, MD +2 57 days ago $100,000$215,000
Azure DevOps Python C# Java Kubernetes Docker Terraform CI/CD SQL NoSQL Ansible Chef Helm GitHub PowerShell Splunk App Insights Infrastructure as Code
4+ yrs exp