Machine Learning Engineer, Agentic AI Evaluation Frameworks

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Cupertino, CA
Salary
$184,700–$277,600 / yr
Posted
2 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $204k
This role $231k
$138k most similar roles pay here $293k

This role pays more than 70% of similar roles. Most pay $162,000–$246,150 — the shaded band above. At the midpoint, this role pays about $231k versus about $204k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Machine Learning Engineer, Agentic AI Evaluation Frameworks

Machine Learning Engineer - Agentic AI Evaluation Frameworks joins the Channel Sales AI Product Engineering team to build and scale evaluation capabilities for next-generation AI-powered experiences. The role involves designing automated evaluation frameworks, pipelines, and high-quality datasets including golden sets and regression suites to improve Generative AI and LLM-powered products. You will develop Auto Eval capabilities, implement model-based approaches like LLM-as-a-Judge, and establish Human-in-the-Loop workflows for complex quality dimensions. Key responsibilities include defining metrics for accuracy, groundedness, and consistency while performing detailed error analysis to improve retrieval systems and agentic task execution. The position requires proficiency in Python and experience with RAG, prompt engineering, and multi-turn conversations. This role solves the challenge of ensuring AI experiences across the Commerce domain are reliable, accurate, and useful by integrating evaluation into CI/CD workflows and production signals.

What does a Machine Learning Engineer earn in California?

Median $246394 from 172 postings across 28 companies.

See salary data

What you'll do

  • Design and develop automated evaluation frameworks and pipelines for LLM, Generative AI, and Agentic AI products.
  • Define quality metrics such as accuracy, groundedness, consistency, and instruction following across various AI dimensions.
  • Build and maintain high-quality evaluation datasets including golden sets, benchmark suites, and adversarial scenarios.
  • Implement model-based evaluation techniques like LLM-as-a-Judge with appropriate calibration and validation methodologies.
  • Develop Human-in-the-Loop (HITL) evaluation processes for complex or subjective quality dimensions.
  • Perform detailed error analysis and failure-mode investigations to identify improvements for models, prompts, and retrieval systems.
  • Build scalable evaluation infrastructure, APIs, and dashboards to support multiple product teams.
  • Integrate automated evaluation gates into CI/CD workflows to ensure release readiness and detect regressions.

What we're looking for

  • Minimum of 7 years of experience in Machine Learning Engineering, ML Evaluation, Software Engineering, Data Science, Quality Engineering, or a related technical field.
  • Bachelor's degree in Computer Science, Machine Learning, AI, Data Science, Statistics, Electrical Engineering, or a related technical field (or equivalent experience).
  • Strong programming skills in Python and experience developing production-quality software, ML systems, data pipelines, or evaluation infrastructure.
  • Experience developing or evaluating LLMs, Generative AI, Conversational AI, NLP, recommendation systems, or other machine-learning-driven products.
  • Experience designing automated ML evaluation frameworks, metrics, benchmarks, datasets, or experimentation methodologies.
  • Understanding of modern LLM application architectures including prompting, embeddings, RAG, tool use, and agentic workflows.
  • Experience with model-based evaluation techniques like LLM-as-a-Judge and Human-in-the-Loop evaluation/annotation workflows.
  • Strong understanding of statistical analysis, experimentation, sampling, measurement methodologies, and performing model error or root-cause analysis.

More like this

Similar roles

Machine Learning Engineer, AI & ML Evaluation Frameworks

Apple Inc

Cupertino, CA 94 days ago $150,400$277,600
Python LLMs Diffusion Models Machine Learning Deep Learning CI/CD Git Spark Kubernetes Airflow RAG Prompt Engineering Synthetic Data Generation Federated Learning Data Pipelines Model Interpretability AI Safety
3+ yrs exp

Machine Learning Engineer, Agentic AI

Apple Inc

Sunnyvale, CA 124 days ago $150,400$277,600
LLMs Agentic Systems Machine Learning Computer Vision Data Pipelines System Design Monitoring Observability

Machine Learning Engineer, AI Evaluation & LLM Systems

Apple Inc

Cupertino, CA 43 days ago $150,400$225,300
Python C++ PyTorch TensorFlow JAX LLM Multimodal AI Generative AI Git CI/CD Distributed Computing Cloud Platforms Data Processing Statistical Analysis Machine Learning Software Engineering
1+ yrs exp

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 65 days ago $184,700$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Machine Learning Evaluation Engineer

Apple Inc

Sunnyvale, CA 13 days ago $150,400$277,600
Machine Learning Computer Vision Python Data Analysis Statistical Modeling Deep Learning Data Pipelines Visualization Synthetic Data Data Augmentation Model Monitoring Generative AI Root-Cause Analysis Evaluation Frameworks
3+ yrs exp

Principal Machine Learning Engineer, Agentic AI

Zillow

Remote 79 days ago $204,400$326,600
Agentic AI LLM LangChain LangGraph AgentSDK Reinforcement Learning Natural Language Processing GenAI A/B Testing Multi-agent Systems Machine Learning Scalable Infrastructure OpenAI Voice API
7+ yrs exp Remote