SRE Manager, ML Operations

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
New York, NY
Salary
$237,600–$356,400 / yr
Posted
77 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $219k
This role $297k
$151k most similar roles pay here $378k

This role pays more than 92% of similar roles. Most pay $183,875–$254,750 — the shaded band above. At the midpoint, this role pays about $297k versus about $219k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · SRE Manager, ML Operations

SRE Manager, ML Operations will lead and grow a Site Reliability Engineering team focused on the reliability, performance, and scalability of ML Platforms and Services. This leader will manage engineers through mentorship and goal-setting while championing SRE best practices such as SLOs, SLAs, error budgets, observability, incident management, and fault analysis. The role involves collaborating with Product, ML Platform, Ads Serving, and Data Science teams to execute high-impact initiatives and shape long-term platform strategy. Candidates should possess expertise in large-scale distributed systems, operating system principles, networking fundamentals, and systems management. Preferred qualifications include experience managing GPU-based clusters, building large-scale ML infrastructure, and managing AWS cloud infrastructure. The role specifically addresses the technical demands of the digital advertising ecosystem to ensure the reliability of the Ad Serving infrastructure.

What you'll do

  • Manage and scale SRE teams responsible for the reliability, performance, and availability of ML Platforms and Services.
  • Mentor engineers through clear goal-setting and career development programs.
  • Define a team vision and drive execution toward high-quality, measurable outcomes.
  • Champion SRE best practices including SLOs/SLAs, error budgets, observability, and incident management.
  • Partner with technical leadership on architecture decisions, platform strategy, and long-term roadmaps.
  • Collaborate cross-functionally with Product, ML Platform, and Data Science teams to deliver complex initiatives.
  • Oversee the reliability of high-scale infrastructure for global ad serving systems.

What we're looking for

  • 10+ years of experience with large-scale distributed systems.
  • 5+ years of experience in an engineering leadership role, ideally managing SRE or Production Engineering teams.
  • Proven track record of building and leading high-performing engineering teams.
  • Strong grasp of operating system principles, networking fundamentals, and systems management.
  • Deep understanding of SRE principles including monitoring, alerting, error budgets, fault analysis, capacity planning, and incident response.
  • Experience managing and optimizing GPU-based clusters in production environments.
  • Experience building and operating large-scale ML systems or infrastructure at scale.
  • Bachelor's or Master's degree in Computer Science or a related field.

More like this

Similar roles

Senior Staff Machine Learning Engineer, ML Platform

Apple Inc

New York, NY 107 days ago $216,200$394,000
Machine Learning Deep Learning Transformers LLMs TensorFlow PyTorch Distributed Training Model Pruning Quantization Distillation Agentic AI Federated Learning Differential Privacy SFT Agile
10+ yrs exp

Site Reliability Engineering Manager

Apple Inc

Cupertino, CA 42 days ago $267,800$401,700
SRE Distributed Systems Linux Kubernetes AWS GCP AI/ML LLM AIOps Capacity Planning Performance Engineering Infrastructure Architecture Networking Automation
10+ yrs exp