Staff Software Engineer, AI Inference Runtime

Arm Holdings

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Seattle, WA
Salary
$209,100–$282,900 / yr
Posted
22 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $223k
This role $246k
$164k most similar roles pay here $296k

This role pays more than 71% of similar roles. Most pay $192,070–$254,750 — the shaded band above. At the midpoint, this role pays about $246k versus about $223k for comparable roles.

Based on 240 similar postings.

Employer

About Arm Holdings

Arm Holdings plc is a leading British semiconductor and software design firm, established in 1990 and recognized for developing energy-efficient processor architectures that power nearly all smartphones and a vast range of IoT and computing devices.

Arm Holdings currently has 109 open roles on FindRole.

Listed pay typically runs $209,100–$282,900 across 108 roles with salary data.

Most-posted roles

View all roles at Arm Holdings

At a glance

TL;DR · Staff Software Engineer, AI Inference Runtime

Staff Software Engineer, AI Inference Runtime will join the AI Inference Runtime team to set technical direction for critical components of distributed inference runtimes for state-of-the-art models. The role involves leading hands-on work in scheduling, batching, KV-cache management, memory allocation, and kernel development while profiling system bottlenecks to optimize data-movement paths across compute, memory, and networking. You will define architectures for inference capabilities, enable new model architectures through operator support, and build benchmarking and validation systems to improve latency and throughput. Required expertise includes experience in ML systems, high-performance systems, compilers, or kernel development. Candidates should possess strong programming skills in C++, Rust, or Python, with deep knowledge of concurrency and parallel programming. The work focuses on the technical challenge of optimizing how efficiently new models utilize available compute resources within an AI platform.

What does a Software Engineer earn in Washington?

Median $202800 from 341 postings across 36 companies.

See salary data

What you'll do

  • Define architecture, interfaces, and roadmaps for AI inference runtime capabilities and abstractions.
  • Enable new model architectures through operator support and optimization of scheduling and batching.
  • Manage memory allocation, KV-cache efficiency, and distributed workload execution.
  • Profile system bottlenecks to develop optimized kernels and data-movement paths across hardware and networks.
  • Build benchmarking, regression, and validation systems to improve latency, throughput, and reliability.
  • Evaluate new inference techniques to enhance performance and resource efficiency on the Arm platform.
  • Lead technical reviews and establish meticulous performance-engineering practices for the team.

What we're looking for

  • 5+ years of experience in ML systems, high-performance systems, compilers, kernel development, or production AI inference.
  • Deep understanding of modern AI inference including model execution, Attention, MoE, batching, and KV-cache behavior.
  • Strong programming skills in C++, Rust, Python, or comparable languages with knowledge of concurrency and memory management.
  • Proven ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, operating systems, and hardware.
  • Experience developing inference schedulers, cache managers, batching systems, or distributed execution paths (preferred).
  • Experience optimizing kernels using accelerator programming tools, assembly, intrinsics, and low-precision execution (preferred).
  • Familiarity with model parallelism, collective communication, high-performance networking, compilers, or graph optimization (preferred).
  • Contributions to open-source ML runtimes, frameworks, compilers, or kernel libraries (preferred).

More like this

Similar roles

Principal Software Engineer, AI Inference Runtime

Arm Holdings

Seattle, WA 22 days ago $262,700–$355,400
C++ Rust Python Kernel Development ML Systems Compiler Assembly performance-engineering Parallel Programming Memory Management Collective Communication Graph Optimization Benchmarking
8+ yrs exp

Staff Software Engineer, AI Inference Cloud

Arm Holdings

Seattle, WA 22 days ago $209,100–$282,900
Kubernetes Go C++ Rust Python AI Inference Distributed Systems PyTorch Ray vLLM SGLang TensorRT-LLM Observability Networking Capacity Planning
5+ yrs exp

Staff Software Engineer, Physical AI

Arm Holdings

Seattle, WA 22 days ago $209,100–$282,900
C++ Rust Go Python Linux Edge Computing Distributed Systems ARM IPC Robotics Performance Optimization
6+ yrs exp Hybrid

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 158 days ago
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid