ML Infrastructure Engineer, ML Compute Capacity

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$184,700–$324,800 / yr
Posted
13 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $220k
This role $255k
$156k most similar roles pay here $343k

This role pays more than 86% of similar roles. Most pay $186,147–$254,750 — the shaded band above. At the midpoint, this role pays about $255k versus about $220k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 2202 open roles on FindRole.

Listed pay typically runs $175,000–$278,800 across 1802 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · ML Infrastructure Engineer, ML Compute Capacity

ML Infrastructure Engineer - ML Compute Capacity joins the Machine Learning Platform Technologies organization to manage large-scale training and inference workloads across a massive accelerator fleet. You will design, build, and operate production systems for demand and capacity planning, including data pipelines and telemetry systems that ingest and serve utilization and cost data. The role involves developing observability infrastructure, optimization algorithms, and self-service platforms with defined schema contracts to balance usability and costs. Key responsibilities include creating end-to-end tooling from data models to interactive dashboards. Required skills include proficiency in Python or Go, experience with Trino, PostgreSQL, Elasticsearch, Prometheus, and Grafana. Preferred qualifications include Kubernetes operations at scale, React for web frameworks, and expertise in accelerator utilization patterns for ML training and inference across multi-tenant and heterogeneous fleets.

What you'll do

  • Build and operate demand and capacity planning systems for large-scale accelerator fleets.
  • Develop data pipelines and telemetry systems to ingest and serve fleet-wide utilization and cost data.
  • Create observability infrastructure including monitoring, alerting, and dashboards for real-time fleet health.
  • Drive innovation in forecasting, optimization, and supply chain management tooling at scale.
  • Build end-to-end tools from data models and APIs to interactive dashboards for leadership insights.
  • Develop self-service platforms with defined schema contracts to balance usability, utilization, and costs.
  • Manage the distribution of compute resources across multi-tenant and heterogeneous hardware environments.

What we're looking for

  • 7+ years of experience in relevant areas.
  • Experience with machine learning infrastructure on GPUs or TPUs.
  • Proficiency in Python and/or Go for production backend and data engineering work.
  • Experience building data pipelines and crafting robust queries over large-scale, multi-source data.
  • Experience with observability tools like Prometheus, Grafana, or equivalent monitoring systems.
  • Strong CS fundamentals and excellent problem-framing and problem-solving skills.
  • Bachelor's degree or higher in Engineering, Mathematics, Economics, or a related quantitative field.
  • Experience with Kubernetes at production scale, modern web frameworks like React, or familiarity with accelerator utilization patterns (preferred).

More like this

Similar roles

ML Infrastructure Engineer, ML Compute Capacity

Apple Inc

Santa Clara, CA 21 days ago $184,700–$324,800
Python Go Kubernetes Prometheus Grafana Trino PostgreSQL Elasticsearch React Data Pipelines Distributed Systems Machine Learning Infrastructure FinOps Monitoring Alerting Capacity Planning
7+ yrs exp

Senior Machine Learning Infrastructure Engineer

Apple Inc

Santa Clara, CA 8 days ago $150,400–$277,600
Python Go Kubernetes Ray Beam Flink JAX TensorFlow PyTorch TensorRT vLLM Distributed Systems Containerization Cloud Computing GPU TPU AWS Trainium ML Training fine tuning
4+ yrs exp