Senior AI Infrastructure Engineer, Physical Infrastructure

Anduril Industries

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Costa Mesa, CA
Salary
$166,000–$220,000 / yr
Posted
8 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $190k
This role $193k
$113k most similar roles pay here $249k

This role pays more than 58% of similar roles. Most pay $144,350–$235,750 — the shaded band above. At the midpoint, this role pays about $193k versus about $190k for comparable roles.

Based on 240 similar postings.

Employer

About Anduril Industries

Anduril Industries is a defense technology company that builds advanced hardware and software systems for national security, including autonomous drones, surveillance systems, and the Lattice AI command platform.

Anduril Industries currently has 1697 open roles on FindRole.

Listed pay typically runs $146,000–$194,000 across 1504 roles with salary data.

Most-posted roles

View all roles at Anduril Industries

At a glance

TL;DR · Senior AI Infrastructure Engineer, Physical Infrastructure

Senior AI Infrastructure Engineer, Physical Infrastructure joins the CorpTech Infrastructure Engineering team to lead the vision and execution for large-scale GPU training infrastructure. This hands-on role involves taking ownership of cluster robustness by building self-healing mechanisms, automating deployment pipelines, and ensuring high availability for ML platform teams. The engineer will rack, stack, and cable GPU compute systems including H200, B200, and B300 models while tuning interconnect fabrics like NVLink, InfiniBand, RoCE, and Spectrum-X. Responsibilities include integrating parallel storage solutions such as VAST, DDN, or Weka, and managing Kubernetes and Run:AI environments for workload isolation. The role requires expertise in firmware management, network fabric tuning, and hardware fault triage to support distributed training and inference. This position solves the technical challenge of providing scalable, automated infrastructure to support advanced AI and autonomy development.

What you'll do

  • Rack, stack, cable, and bring up high-performance GPU compute systems including firmware and BIOS configuration.
  • Build and tune interconnect fabrics like NVLink, InfiniBand, and RoCE for large-scale training clusters.
  • Integrate high-performance parallel storage to support distributed training and multi-modal datasets.
  • Automate end-to-end cluster deployment using infrastructure as code for hardware and network configurations.
  • Manage Kubernetes and Run:AI environments for GPU scheduling, quota management, and workload isolation.
  • Monitor fleet health by proactively detecting and triaging hardware faults and network congestion issues.
  • Onboard internal researchers and engineers while serving as the primary escalation point for infrastructure bottlenecks.
  • Translate emerging compute needs from product teams into scalable platform capabilities.

What we're looking for

  • Must have 10+ years of experience in infrastructure, HPC, or datacenter engineering supporting GPU compute at scale.
  • Must have hands-on experience with H200/B200/B300 (or comparable) GPU systems including bring up and firmware management.
  • Must have experience with high-performance interconnects like NVLink, InfiniBand, RoCE, or Spectrum-X in large clusters.
  • Must have experience with high-performance parallel storage systems such as VAST, DDN, Weka, or Lustre.
  • Must have experience with Kubernetes and a strong background in building automated deployment pipelines.
  • Must be able to lift 50+ lbs and perform physical datacenter tasks like racking and cabling.
  • Must be eligible to obtain and maintain an active U.S. Top Secret clearance.

More like this

Similar roles

AI Infrastructure Operations Engineer

Accenture

Remote (Boston, MA) +4 10 days ago $94,400$266,300
GPU Kubernetes Slurm Run:ai Terraform Ansible Python Bash CUDA-X NCCL REST API JSON YAML NVIDIA HPC Bare-metal NVMe-oF Parallel File Systems
5+ yrs exp Remote

AI Infrastructure Engineer

Fortinet

New York, NY 14 days ago $215,000$350,000
Linux GPU Docker Kubernetes Python Bash CI/CD KVM FortiGate FortiManager FortiAnalyzer Monitoring Logging Alerting Networking Virtualization Performance Testing Benchmarking Automation

Senior AI Infrastructure Engineer, EDA Infrastructure

Nvidia

Remote (Westford, MA) +2 8 days ago $184,000$287,500
Python Go TypeScript Java Telemetry Pipelines Metrics Logs Traces Observability SaaS ML Models AI Agent Frameworks CMDB incident-management Configuration Management
8+ yrs exp Remote

Staff AI Infrastructure Engineer

Anduril Industries

Costa Mesa, CA +2 23 days ago $220,000$292,000
MLOps Python Go C++ Docker Kubernetes PyTorch Distributed Ray Slurm Megatron-LM Computer Vision RLHF DPO ETL CI/CD AWS Trainium Google TPU
7+ yrs exp

Senior Solutions Architect, Generative AI

Nvidia

Remote (Santa Clara, CA) +1 43 days ago $184,000$287,500
GPU InfiniBand RoCE RDMA NCCL NVLink NVSwitch Kubernetes Slurm Python Linux High-Performance Computing Distributed Systems Shell Scripting DCGM Nsight Systems
6+ yrs exp Remote