Principal Software Engineer, AI Compute Infrastructure

Arm Holdings

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Seattle, WA
Salary
$262,700–$355,400 / yr
Posted
2 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $208k
This role $309k
$117k most similar roles pay here $381k

This role pays more than 92% of similar roles. Most pay $174,600–$241,600 — the shaded band above. At the midpoint, this role pays about $309k versus about $208k for comparable roles.

Based on 240 similar postings.

Employer

About Arm Holdings

Arm Holdings plc is a leading British semiconductor and software design firm, established in 1990 and recognized for developing energy-efficient processor architectures that power nearly all smartphones and a vast range of IoT and computing devices.

Arm Holdings currently has 48 open roles on FindRole.

Listed pay typically runs $209,100–$282,900 across 47 roles with salary data.

Most-posted roles

View all roles at Arm Holdings

At a glance

TL;DR · Principal Software Engineer, AI Compute Infrastructure

As a Principal Software Engineer, AI Compute Infrastructure, you will join the AI Compute Infra team to design, build, and operate large-scale infrastructure for AI training, fine-tuning, evaluation, and inference. You will manage Kubernetes clusters, accelerator enablement, workload scheduling, high-performance networking, storage, and capacity management while partnering with researchers to improve reliability and developer productivity. Your daily work involves integrating and validating CPU and GPU systems, investigating performance issues across cloud infrastructure and hardware, and automating cluster provisioning and maintenance. The role requires proficiency in Go, Python, Linux, and Kubernetes, along with experience in distributed machine-learning workloads. Preferred technical skills include NVIDIA technologies like CUDA and NVLink, as well as tools such as Terraform, Argo CD, Prometheus, and Grafana. You will solve complex problems related to scaling infrastructure for large-scale AI models.

What does a Software Engineer earn in Washington?

Median $202800 from 318 postings across 31 companies.

See salary data

What you'll do

  • Build and operate large-scale Kubernetes clusters for AI training, fine-tuning, and inference.
  • Optimize workload scheduling, topology-aware placement, and capacity management across compute clusters.
  • Integrate and validate drivers, networking, storage, and health checks for new CPU and GPU systems.
  • Investigate and resolve performance and reliability issues across hardware, cloud infrastructure, and applications.
  • Automate cluster provisioning, upgrades, monitoring, and maintenance tasks to support AI research teams.
  • Develop reliable infrastructure software using languages like Go or Python.
  • Tune distributed machine learning workloads and qualify specialized accelerators.

What we're looking for

  • 8+ years of experience building or operating cloud, compute, HPC, or distributed infrastructure in a production environment.
  • Programming experience in Go, Python, or another systems language for developing reliable infrastructure software.
  • Practical knowledge of Kubernetes, containers, Linux, networking, and storage.
  • Experience supporting GPU, accelerator, or distributed machine-learning workloads.
  • Ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds.
  • Familiarity with Kubernetes scheduling, operators, quotas, or resource management (preferred).
  • Experience with NVIDIA technologies such as CUDA, NVLink, NVSwitch, NCCL, EFA, or DCGM (preferred).
  • Knowledge of AWS EKS, Terraform, Argo CD, Helm, Prometheus, Grafana, or specific AI frameworks like PyTorch and Ray (preferred).

More like this

Similar roles

Principal Software Engineer, AI Compute Platform

Arm Holdings

Seattle, WA 2 days ago $262,700$355,400
Kubernetes Go Python Distributed Systems APIs Containers PostgreSQL Kafka Prometheus Grafana OpenTelemetry PyTorch Ray vLLM GitOps RPC Asynchronous Processing
8+ yrs exp

Staff Software Engineer, AI Compute Platform

Arm Holdings

Seattle, WA 2 days ago $209,100$282,900
Kubernetes Go Python Distributed Systems APIs Containers PostgreSQL Kafka Prometheus Grafana OpenTelemetry PyTorch Ray vLLM GitOps RPC Asynchronous Processing
5+ yrs exp