Senior AI Inference Platform Engineer

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Seattle, WA
Salary
$175,000–$308,500 / yr
Posted
31 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $204k
This role $242k
$144k most similar roles pay here $326k

This role pays more than 64% of similar roles. Most pay $162,000–$246,271 — the shaded band above. At the midpoint, this role pays about $242k versus about $204k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Senior AI Inference Platform Engineer

As a Sr. AI Inference Platform Engineer within the next-generation datacenter engineering team, you will build tooling, automation, and analysis capabilities to strengthen the AI inference platform. You will develop sophisticated performance benchmarking systems, capacity projection models, and data analysis pipelines to inform infrastructure teams and capacity planners regarding scaling and optimization. Your daily work involves designing automations to evaluate performance across hardware generations, identifying bottlenecks in utilization data, and creating workflows to improve data reliability. The role requires expertise in Python, Go, or C++, along with a deep understanding of AI/ML inference architecture, GPU accelerator architecture, and distributed systems. You will utilize tools such as Triton, TensorRT-LLM, vLLM, Nsight, Prometheus, and Grafana to solve complex problems involving performance trends, regression detection, and long-term capacity forecasting for large-scale infrastructure.

What you'll do

  • Build automations to evaluate AI inference performance across various hardware generations and configurations.
  • Develop tooling to surface performance trends, regressions, and insights for infrastructure teams.
  • Create projection and forecasting models to support long-term capacity planning decisions.
  • Analyze performance and utilization data to identify system bottlenecks and optimization opportunities.
  • Design and enhance performance analysis workflows to improve team velocity and data reliability.
  • Improve the accuracy and coverage of performance measurement and analysis systems.

What we're looking for

  • BS or MS in Computer Science or a related technical field.
  • 7 or more years of experience with performance and infrastructure engineering in distributed systems.
  • 7 years of experience coding in Python, Go, C++, or other programming languages.
  • Solid understanding of AI/ML inference architecture and the performance characteristics of serving systems.
  • Strong knowledge of GPU/accelerator architecture as it relates to AI workloads.
  • Experience with automation engineering, tooling, and data pipelines to support engineering workflows.
  • Practical statistical knowledge applicable to performance analysis and forecasting.
  • Experience with benchmarking, capacity planning, GPU profiling, ML serving frameworks, or cluster orchestration (preferred).

More like this

Similar roles

Senior AI Inference Platform Engineer

Apple Inc

Seattle, WA 30 days ago $175,000$308,500
Python Go C++ Triton TensorRT-LLM vLLM Kubernetes Prometheus Grafana Splunk CI/CD Data Pipelines Performance Benchmarking Distributed Systems Capacity Planning
7+ yrs exp

AI Inference Engineer

F5 Inc

San Jose, CA +1 15 days ago $176,600$265,000
vLLM TGI NVIDIA Triton Python C++ Rust TensorRT Llama.cpp Ollama Kubernetes Docker AWS GCP Azure CUDA TPUs MLOps
Hybrid

Senior AI/ML Platform Engineer

Amd

Santa Clara, CA 51 days ago $204,000$306,000
Python C++ Go Rust Kubernetes Ray Slurm vLLM SGLang Triton PyTorch JAX ROCm HIP CUDA MLflow Weights & Biases CI/CD Distributed Systems GPU Infrastructure
Hybrid

Software Engineer, Inference (AI Data Engineering)

SpaceX

Palo Alto, CA 23 days ago $135,000$175,000
SGLang vLLM TensorRT-LLM Triton Rust C++ Python Go gRPC REST Docker Kubernetes PostgreSQL ClickHouse MongoDB CI/CD GPU Kernels Quantization Speculative Decoding
2+ yrs exp