Principal Software Engineer, AI Inference Cloud

Arm Holdings

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Seattle, WA
Salary
$262,700–$355,400 / yr
Posted
2 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $208k
This role $309k
$118k most similar roles pay here $381k

This role pays more than 93% of similar roles. Most pay $174,600–$241,937 — the shaded band above. At the midpoint, this role pays about $309k versus about $208k for comparable roles.

Based on 240 similar postings.

Employer

About Arm Holdings

Arm Holdings plc is a leading British semiconductor and software design firm, established in 1990 and recognized for developing energy-efficient processor architectures that power nearly all smartphones and a vast range of IoT and computing devices.

Arm Holdings currently has 48 open roles on FindRole.

Listed pay typically runs $209,100–$282,900 across 47 roles with salary data.

Most-posted roles

View all roles at Arm Holdings

At a glance

TL;DR · Principal Software Engineer, AI Inference Cloud

As a Principal Software Engineer on the AI Inference Cloud team, you will shape technical direction and develop highly available, scalable services for running AI inference workloads. You will define architecture and build Kubernetes controllers and platform capabilities to support workload deployment, scheduling, recovery, scaling, and lifecycle management. Your daily work involves establishing production practices for health validation, progressive rollouts, and observability while resolving complex issues across networking and compute infrastructure. The role requires extensive experience in distributed systems and deep expertise with Kubernetes operators and resource management. You will utilize programming languages such as Go, C++, Rust, or Python to build reliable services. Preferred technical skills include familiarity with PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM. This role focuses on solving the technical challenges of model serving and accelerator-backed workloads within an AI inference platform.

What does a Software Engineer earn in Washington?

Median $202800 from 318 postings across 31 companies.

See salary data

What you'll do

  • Define and build the architecture for cloud-based AI inference services.
  • Develop Kubernetes controllers to manage workload deployment, scheduling, scaling, and lifecycle management.
  • Establish production practices for health validation, progressive rollouts, and service observability.
  • Improve platform reliability, scalability, performance, and resource efficiency.
  • Resolve complex issues across services, networking, and compute infrastructure to create durable improvements.
  • Lead design and build reviews while mentoring engineers and driving technical alignment.
  • Partner with AI compute and inference teams to enhance the performance of Arm's AI platform.

What we're looking for

  • 8+ years of experience building distributed systems, cloud platforms, or production infrastructure.
  • Deep production experience with Kubernetes including controllers, operators, scheduling, and workload lifecycle management.
  • Strong software engineering skills in Go, C++, Rust, Python, or similar languages.
  • Experience with reliable services, APIs, concurrency, observability, deployment safety, capacity planning, and incident response.
  • A track record of leading complex technical initiatives while remaining hands-on in architecture, implementation, and mentoring.
  • Ability to troubleshoot complex systems and communicate clearly across different technical backgrounds.
  • Experience with AI infrastructure, model serving, or accelerator-backed workloads (preferred).
  • Familiarity with frameworks like PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM (preferred).

More like this

Similar roles

Staff Software Engineer, AI Inference Cloud

Arm Holdings

Seattle, WA 2 days ago $209,100$282,900
Kubernetes Go C++ Rust Python AI Inference Distributed Systems PyTorch Ray vLLM SGLang TensorRT-LLM Observability Networking Capacity Planning
5+ yrs exp

Senior AI Inference Platform Engineer

Apple Inc

Seattle, WA 34 days ago $175,000$308,500
Python Go C++ Triton TensorRT-LLM vLLM Kubernetes Prometheus Grafana Splunk CI/CD Data Pipelines Performance Benchmarking Distributed Systems Capacity Planning
7+ yrs exp

Senior AI Inference Platform Engineer

Apple Inc

Seattle, WA 33 days ago $175,000$308,500
Python Go C++ Triton TensorRT-LLM vLLM Kubernetes Prometheus Grafana Splunk CI/CD Data Pipelines Performance Benchmarking Distributed Systems Capacity Planning
7+ yrs exp

Principal Software Engineer, AI Compute Platform

Arm Holdings

Seattle, WA 2 days ago $262,700$355,400
Kubernetes Go Python Distributed Systems APIs Containers PostgreSQL Kafka Prometheus Grafana OpenTelemetry PyTorch Ray vLLM GitOps RPC Asynchronous Processing
8+ yrs exp

Principal Platform Software Engineer, AI

Oracle

Nashville, TN 30 days ago $104,500$234,600
Python Go Java C++ Kubernetes OCI LLM Machine Learning PyTorch TensorFlow Hugging Face CI/CD Microservices Distributed Systems Infrastructure Automation
8+ yrs exp