Principal Software Engineer, AI Inference Runtime

Arm Holdings

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Seattle, WA
Salary
$262,700–$355,400 / yr
Posted
22 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $209k
This role $309k
$117k most similar roles pay here $381k

This role pays more than 96% of similar roles. Most pay $174,600–$242,500 — the shaded band above. At the midpoint, this role pays about $309k versus about $209k for comparable roles.

Based on 240 similar postings.

Employer

About Arm Holdings

Arm Holdings plc is a leading British semiconductor and software design firm, established in 1990 and recognized for developing energy-efficient processor architectures that power nearly all smartphones and a vast range of IoT and computing devices.

Arm Holdings currently has 109 open roles on FindRole.

Listed pay typically runs $209,100–$282,900 across 108 roles with salary data.

Most-posted roles

View all roles at Arm Holdings

At a glance

TL;DR · Principal Software Engineer, AI Inference Runtime

As a Principal Software Engineer on the AI Inference Runtime team, you will set the technical direction for critical components of distributed inference runtimes used to execute state-of-the-art AI models. You will lead hands-on development involving scheduling, batching, KV-cache management, memory allocation, and kernel optimization while defining architectures and roadmaps for inference capabilities. Your daily work involves profiling system bottlenecks, developing data-movement paths, and building benchmarking and validation systems to improve latency and throughput. The role requires expertise in C++, Rust, or Python, alongside deep knowledge of concurrency, parallel programming, and high-performance systems. You will solve complex problems regarding how efficiently new models utilize available compute by optimizing execution across kernels, runtimes, frameworks, and hardware, while collaborating with infrastructure and compiler teams to enhance the performance of the underlying AI platform.

What does a Software Engineer earn in Washington?

Median $202800 from 341 postings across 36 companies.

See salary data

What you'll do

  • Define architecture, interfaces, and roadmaps for AI inference runtime capabilities and abstractions.
  • Enable new model architectures through operator support and optimization of scheduling, batching, and memory management.
  • Profile system bottlenecks to develop optimized kernels and data-movement paths across compute and networking.
  • Build benchmarking, regression, and validation systems to improve latency, throughput, and resource efficiency.
  • Manage KV-cache efficiency and distributed workload execution for large-scale AI models.
  • Lead technical reviews and establish performance-engineering practices to optimize the Arm AI platform.

What we're looking for

  • 8+ years of experience in ML systems, high-performance systems, compilers, kernel development, or production AI inference.
  • Deep understanding of modern AI inference including model execution, Attention, MoE, batching, and KV-cache behavior.
  • Strong programming skills in C++, Rust, Python, or comparable languages with knowledge of concurrency and memory management.
  • Proven ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, operating systems, and hardware.
  • Experience developing inference schedulers, cache managers, batching systems, or distributed execution paths (preferred).
  • Experience optimizing kernels using accelerator programming tools, assembly, intrinsics, and low-precision execution (preferred).
  • Familiarity with model parallelism, collective communication, high-performance networking, compilers, or graph optimization (preferred).
  • Contributions to open-source ML runtimes, frameworks, compilers, or kernel libraries (preferred).

More like this

Similar roles

Staff Software Engineer, AI Inference Runtime

Arm Holdings

Seattle, WA 22 days ago $209,100–$282,900
C++ Rust Python Kernel Development ML Systems performance-engineering Assembly Parallel Programming High-Performance Networking Matrix Multiplication Graph Optimization Benchmarking
5+ yrs exp Hybrid

Principal Software Engineer, AI Compute Platform

Arm Holdings

Seattle, WA 22 days ago $262,700–$355,400
Kubernetes Go Python Distributed Systems APIs Containers PostgreSQL Kafka Prometheus Grafana OpenTelemetry PyTorch Ray vLLM GitOps RPC Asynchronous Processing
8+ yrs exp

Staff Software Engineer, AI Inference Cloud

Arm Holdings

Seattle, WA 22 days ago $209,100–$282,900
Kubernetes Go C++ Rust Python AI Inference Distributed Systems PyTorch Ray vLLM SGLang TensorRT-LLM Observability Networking Capacity Planning
5+ yrs exp