Senior Machine Learning Engineering Manager, Evaluation

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Cupertino, CA
Salary
$237,600–$401,700 / yr
Posted
14 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $229k
This role $320k
$159k most similar roles pay here $428k

This role pays more than 97% of similar roles. Most pay $202,412–$254,750 — the shaded band above. At the midpoint, this role pays about $320k versus about $229k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Senior Machine Learning Engineering Manager, Evaluation

AIML - Sr Machine Learning Engineering Manager, Evaluation joins the AIML Evaluation team to lead a small team focused on model evaluation, agent optimization, and data generation. This hands-on leadership role involves building scalable evaluation systems for foundation models and agents, including LLM-based evaluators, simulation environments, and trajectory analysis. The manager will develop methods to convert evaluation findings into actionable signals like reward models, synthetic trajectories, and optimized prompts or contexts. Key responsibilities include designing automated optimization pipelines and scaling synthetic data generation for post-training while ensuring quality and privacy. The role requires expertise in Python, large language models, agentic systems, and multi-turn behavior. By bridging research and production, the manager creates a feedback loop to improve intelligence experiences by addressing failure modes through systematic improvements to tools, rubrics, and model training data.

What you'll do

  • Architect and build scalable evaluation systems including benchmarks, LLM-based evaluators, and simulation environments.
  • Establish an end-to-end evaluation flywheel that connects observed failures to targeted model and agent refinements.
  • Lead and mentor a team of machine learning engineers while remaining involved in technical design and code reviews.
  • Define the technical strategy for automatic optimization of prompts, context, tools, and rubrics.
  • Convert evaluation findings into actionable signals such as reward models, preference signals, and synthetic trajectories.
  • Design and scale synthetic data generation pipelines for both model evaluation and post-training.
  • Translate recent research in LLM evaluation and agentic systems into production-quality workflows.

What we're looking for

  • Master's or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field.
  • 8+ years of professional experience in machine learning, applied research, or software engineering.
  • 3+ years of technical leadership experience including direct people management and mentoring of engineers.
  • Strong hands-on programming skills in Python and building reliable ML pipelines with modern frameworks.
  • Deep experience with large language models or agentic systems, including multi-turn behavior and tool use.
  • Experience building automated evaluation methods such as LLM-based judges, reward models, or simulation-based evaluation.
  • Experience with model refinement areas like automatic prompt optimization, post-training, or reinforcement learning.
  • Proven ability to communicate complex technical trade-offs and align research, engineering, and product teams.

More like this

Similar roles

Senior Machine Learning Engineer, Evaluation

Apple Inc

Cupertino, CA 91 days ago $216,200$394,000
Machine Learning LLM Python PyTorch Reward Modeling RLHF SFT DPO Prompt Optimization Distributed Systems Agentic Systems Simulation Environments ML Infrastructure Data Generation
8+ yrs exp

Senior Manager, Evaluation - Data Science & Insights

Apple Inc

Seattle, WA 22 days ago $225,600$381,600
Machine Learning LLM Data Science Human Evaluation Autograders Rubrics Agentic Systems Logging Infrastructure Instrumentation Evaluation Frameworks Statistics Computer Science
6+ yrs exp

Senior Data Scientist, AIML Evaluation

Apple Inc

Cupertino, CA 76 days ago $150,400$277,600
Machine Learning LLMs Python SQL Spark R Scala Prompt Engineering Statistical Data Analysis Data Processing Experimentation

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 65 days ago $184,700$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Software Engineer - AI, Evaluation

Apple Inc

Cupertino, CA 2 days ago $150,400$277,600
Python LLM MLOps CI/CD System Design API Design Machine Learning AI Testing Monitoring Debugging