Senior Machine Learning Engineer, Foundation Models Inference

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$184,700–$324,800 / yr
Posted
28 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $230k
This role $255k
$166k most similar roles pay here $342k

This role pays more than 81% of similar roles. Most pay $204,450–$254,750 — the shaded band above. At the midpoint, this role pays about $255k versus about $230k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Senior Machine Learning Engineer, Foundation Models Inference

Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference joins the Foundation Model Inference team within the Cloud OS and AI Inference organization. This role sits at the intersection of research and production, where you will partner with research teams to move cutting-edge model architectures from prototype to large-scale deployment. You will be responsible for optimizing inference for language, vision, and speech models while building profiling tools, simulators, and high-throughput serving systems. The work involves solving complex problems in inference efficiency, hardware/software codesign, and systems architecture on privacy-preserving cloud infrastructure. Required skills include experience with LLM inference stacks, GPU or TPU programming, PyTorch, JAX, or TensorFlow. Preferred qualifications include proficiency in Go or Python, knowledge of Transformer architectures, and experience with frameworks like TensorRT-LLM, vLLM, SGLang, TGI, or Triton for custom CUDA kernels.

What does a Machine Learning Engineer earn in California?

Median $246394 from 172 postings across 28 companies.

See salary data

What you'll do

  • Optimize inference for the latest language, vision, and speech model architectures.
  • Design and ship production-grade inference systems serving millions of customers in real time.
  • Build profiling tools and simulators to identify and resolve performance bottlenecks across hardware configurations.
  • Drive technical decisions regarding high-throughput, low-latency serving at supercomputing scale.
  • Develop and optimize inference systems on privacy-preserving cloud infrastructure.
  • Lead the transition of cutting-edge model architectures from prototype to planetary-scale deployment.
  • Mentor and grow engineers across the organization.

What we're looking for

  • 5+ years of experience leading complex, ambiguous technical projects from end to end.
  • Hands-on experience with LLM inference stacks.
  • Working knowledge of GPU or TPU programming concepts.
  • Proficiency with PyTorch, JAX, or TensorFlow.
  • Experience building and operating high-throughput services at large distributed scale.
  • Proficiency deploying applications on cloud platforms using Kubernetes and Docker.
  • BS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field.
  • Experience with Go/Python, deep learning architectures, inference optimization frameworks, or custom CUDA kernels (preferred); MS degree (preferred).

More like this

Similar roles

Machine Learning Engineer, Foundation Model Services

Apple Inc

Santa Clara, CA 113 days ago $184,700$324,800
LLMs Machine Learning NLP Information Retrieval Python Golang Kubernetes Docker AWS Azure PyTorch TensorFlow Transformers Nvidia TensorRT-LLM DeepSpeed Nvidia Triton Server
5+ yrs exp

Machine Learning Engineer, Foundation Model Services

Apple Inc

Seattle, WA 113 days ago $175,000$308,500
LLMs Machine Learning NLP Python Golang Kubernetes Docker AWS Azure PyTorch TensorFlow Transformers Nvidia TensorRT-LLM DeepSpeed Nvidia Triton Server Information Retrieval Statistics
5+ yrs exp

Senior AI Inference Platform Engineer

Apple Inc

Seattle, WA 31 days ago $175,000$308,500
Python Go C++ Triton TensorRT-LLM vLLM Kubernetes Prometheus Grafana Splunk CI/CD Data Pipelines Performance Benchmarking Distributed Systems Capacity Planning
7+ yrs exp

Senior AI Inference Platform Engineer

Apple Inc

Seattle, WA 30 days ago $175,000$308,500
Python Go C++ Triton TensorRT-LLM vLLM Kubernetes Prometheus Grafana Splunk CI/CD Data Pipelines Performance Benchmarking Distributed Systems Capacity Planning
7+ yrs exp