Staff Software Engineer, AI Compute Infrastructure

Arm Holdings

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Seattle, WA
Salary
$209,100–$282,900 / yr
Posted
2 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $221k
This role $246k
$167k most similar roles pay here $295k

This role pays more than 72% of similar roles. Most pay $191,468–$251,000 — the shaded band above. At the midpoint, this role pays about $246k versus about $221k for comparable roles.

Based on 240 similar postings.

Employer

About Arm Holdings

Arm Holdings plc is a leading British semiconductor and software design firm, established in 1990 and recognized for developing energy-efficient processor architectures that power nearly all smartphones and a vast range of IoT and computing devices.

Arm Holdings currently has 48 open roles on FindRole.

Listed pay typically runs $209,100–$282,900 across 47 roles with salary data.

Most-posted roles

View all roles at Arm Holdings

At a glance

TL;DR · Staff Software Engineer, AI Compute Infrastructure

Staff Software Engineer, AI Compute Infrastructure will join the AI Compute Infra team to design, build, and operate large-scale infrastructure for AI training, fine-tuning, evaluation, and inference. The role involves building and operating Kubernetes clusters, improving workload scheduling, and enabling new CPU and GPU systems by integrating drivers, networking, storage, and monitoring. You will investigate performance issues across cloud infrastructure and hardware while partnering with AI teams to automate cluster provisioning and maintenance. Required skills include experience in Go or Python, along with practical knowledge of Kubernetes, containers, Linux, and distributed machine-learning workloads. Preferred expertise includes NVIDIA technologies like CUDA and NVLink, as well as tools such as Terraform, Argo CD, Prometheus, and Grafana. The role focuses on solving complex infrastructure challenges to improve reliability, performance, and scalability for large-scale AI compute systems.

What does a Software Engineer earn in Washington?

Median $202800 from 318 postings across 31 companies.

See salary data

What you'll do

  • Build and operate large-scale Kubernetes clusters for AI training, fine-tuning, and inference.
  • Optimize workload scheduling, topology-aware placement, and capacity management across clusters.
  • Integrate and validate drivers, networking, storage, and health checks for new CPU and GPU systems.
  • Investigate and resolve performance and reliability issues across hardware, cloud infrastructure, and applications.
  • Automate cluster provisioning, upgrades, monitoring, and maintenance based on AI team requirements.
  • Develop reliable infrastructure software using languages like Go or Python.
  • Tune distributed workloads and qualify accelerators for machine learning tasks.

What we're looking for

  • 5+ years of experience building or operating cloud, compute, HPC, or distributed infrastructure in a production environment.
  • Programming experience in Go, Python, or another systems language.
  • Practical knowledge of Kubernetes, containers, Linux, networking, and storage.
  • Experience supporting GPU, accelerator, or distributed machine-learning workloads.
  • Ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds.
  • Familiarity with Kubernetes scheduling, operators, quotas, or resource management (preferred).
  • Experience with NVIDIA technologies such as CUDA, NVLink, NVSwitch, NCCL, EFA, or DCGM (preferred).
  • Knowledge of AWS EKS, Terraform, Argo CD, Helm, Prometheus, or Grafana; familiarity with frameworks like PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM (preferred).

More like this

Similar roles

Staff Software Engineer, AI Compute Platform

Arm Holdings

Seattle, WA 2 days ago $209,100$282,900
Kubernetes Go Python Distributed Systems APIs Containers PostgreSQL Kafka Prometheus Grafana OpenTelemetry PyTorch Ray vLLM GitOps RPC Asynchronous Processing
5+ yrs exp

Staff Software Engineer, AI Inference Cloud

Arm Holdings

Seattle, WA 2 days ago $209,100$282,900
Kubernetes Go C++ Rust Python AI Inference Distributed Systems PyTorch Ray vLLM SGLang TensorRT-LLM Observability Networking Capacity Planning
5+ yrs exp

Principal Software Engineer, AI Compute Platform

Arm Holdings

Seattle, WA 2 days ago $262,700$355,400
Kubernetes Go Python Distributed Systems APIs Containers PostgreSQL Kafka Prometheus Grafana OpenTelemetry PyTorch Ray vLLM GitOps RPC Asynchronous Processing
8+ yrs exp

Staff Software Engineer, Core AI Infrastructure

Coinbase

Remote 52 days ago $218,025$256,500
AWS Kubernetes Terraform Go Python Docker CI/CD Ansible Chef Puppet Salt Git Bash Ruby Distributed Systems Data Pipelines infrastructure-as-code
Remote

Staff Software Engineer, AI Platform

JLL (Jones Lang LaSalle)

Remote (Chicago, IL) 4 days ago $160,000$260,000
LLM AI Engineering Distributed Systems Kubernetes Infrastructure as Code Observability Retrieval Guardrails
8+ yrs exp Remote