Senior AI Infrastructure Engineer

Anduril Industries

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Costa Mesa, CASeattle, WAWashington, DCBoston, MA
Salary
$191,000–$253,000 / yr
Posted
8 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $187k
This role $222k
$128k most similar roles pay here $266k

This role pays more than 71% of similar roles. Most pay $146,125–$227,750 — the shaded band above. At the midpoint, this role pays about $222k versus about $187k for comparable roles.

Based on 240 similar postings.

Employer

About Anduril Industries

Anduril Industries is a defense technology company that builds advanced hardware and software systems for national security, including autonomous drones, surveillance systems, and the Lattice AI command platform.

Anduril Industries currently has 850 open roles on FindRole.

Listed pay typically runs $166,000–$220,000 across 821 roles with salary data.

Most-posted roles

View all roles at Anduril Industries

At a glance

TL;DR · Senior AI Infrastructure Engineer

As a Senior AI Infrastructure Engineer on the Air Dominance & Strike team, you will build, scale, and optimize the end-to-end machine learning platform powering autonomous systems like unmanned fighter jets and cruise missiles. You will own critical components of the MLOps tooling, developing infrastructure to train, evaluate, host, and serve complex models including LLMs, computer vision, and RL agents across cloud and air-gapped edge networks. Your daily work involves building scalable training pipelines, managing multi-modal data ETL processes for video and telemetry, and implementing CI/CD workflows for safety-critical deployments. You will utilize Python, Go, or C++ alongside tools like Docker, Kubernetes, PyTorch Distributed, Ray, and Slurm. The role focuses on solving the technical challenge of deploying high-throughput, low-latency model serving in resource-constrained environments to ensure reliable performance in mission-critical defense operations.

What you'll do

  • Build and maintain scalable training, orchestration, and experimentation infrastructure for machine learning models.
  • Develop tools for experiment tracking, automated profiling, and hyperparameter tuning to resolve lifecycle bottlenecks.
  • Implement robust ETL pipelines to process multi-modal data from physical assets and test sites.
  • Deploy high-throughput, low-latency model serving frameworks for cloud and air-gapped tactical edge hardware.
  • Create CI/CD pipelines for ML models including automated regression testing and safe rollout strategies.
  • Develop evaluation and reinforcement learning alignment loops to ensure safety in mission-critical deployments.
  • Translate research requirements into scalable infrastructure while mentoring peers and conducting code reviews.

What we're looking for

  • 5+ years of software engineering experience in building and operating production-scale machine learning infrastructure or distributed systems.
  • Proficiency in Python, Go, or C++ with strong knowledge of systems design and concurrent programming.
  • Hands-on experience with container orchestration (Docker, Kubernetes) and distributed training frameworks like PyTorch Distributed, Ray, Slurm, or Megatron-LM.
  • Experience building and maintaining distributed data pipelines for large-scale unstructured or multi-modal datasets.
  • Proven track record of owning projects from technical design through production deployment and operational monitoring.
  • Eligible to obtain and maintain an active U.S. Top Secret security clearance.
  • Experience deploying ML infrastructure in secure, air-gapped, or regulated environments (preferred).
  • Experience with GPU profiling, multi-tenant cluster management, or supporting LLM and Reinforcement Learning pipelines (preferred).

More like this

Similar roles

Staff AI Infrastructure Engineer

Anduril Industries

Costa Mesa, CA +2 44 days ago $220,000–$292,000
MLOps Python Go C++ Docker Kubernetes PyTorch Distributed Ray Slurm Megatron-LM Computer Vision RLHF DPO ETL CI/CD AWS Trainium Google TPU
7+ yrs exp

Senior AI Infrastructure Engineer, EDA Infrastructure

Nvidia

Remote (Westford, MA) +2 4 days ago $184,000–$287,500
Python Go TypeScript Java Telemetry Pipelines Metrics Logs Traces Observability SaaS ML Models AI Agent Frameworks CMDB incident-management Configuration Management
8+ yrs exp Remote

AI Infrastructure Engineer

Fortinet

New York, NY 35 days ago $215,000–$350,000
Linux GPU Docker Kubernetes Python Bash CI/CD KVM FortiGate FortiManager FortiAnalyzer Monitoring Logging Alerting Networking Virtualization Performance Testing Benchmarking Automation

AI Infrastructure Engineer

Amd

San Jose, CA 27 days ago $204,000–$306,000
Kubernetes Terraform GitOps ArgoCD Flux Helm Prometheus Grafana Loki CSI CNI Slurm PyTorch vLLM SGLang Infrastructure as Code
Hybrid