ML Engineer, Automated Evaluation and Adversarial Design

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Cupertino, CA
Salary
$150,400–$277,600 / yr
Posted
165 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $221k
This role $214k
$135k most similar roles pay here $293k

This role pays more than 59% of similar roles. Most pay $187,850–$254,750 — the shaded band above. At the midpoint, this role pays about $214k versus about $221k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 3587 open roles on FindRole.

Listed pay typically runs $165,800–$277,600 across 2768 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · ML Engineer, Automated Evaluation and Adversarial Design

ML Engineer - Automated Evaluation and Adversarial Design joins the Productivity and Machine Learning Evaluation team to ensure quality across various AI-powered features. This role focuses on building and scaling automated evaluation systems, designing adversarial test suites, and developing stress-testing methodologies for multi-turn conversation flows and agentic experience chains. The candidate will translate qualitative quality notions into measurable assessments, identify failure modes like context degradation or goal drift, and create robust evaluation rubrics and report findings to cross-functional partners. Required skills include Python and ML frameworks such as PyTorch or TensorFlow. Preferred expertise includes familiarity with agent orchestration frameworks like LangChain, LangGraph, CrewAI, or AutoGen, along with observability tools like LangSmith, Braintrust, or Arize. The role addresses the technical challenge of ensuring model reliability across complex, multi-step task sequences and tool-use interactions.

What you'll do

  • Define and own automated evaluation approaches that translate qualitative quality into measurable metrics for single-turn and multi-turn AI experiences.
  • Build adversarial test suites targeting model failure modes such as context loss, instruction forgetting, and cascading errors.
  • Develop and execute stress test protocols to validate performance under atypical conditions like extended conversation lengths and complex tool-use sequences.
  • Ensure alignment between automated and human evaluation methods by identifying and resolving systematic disagreements.
  • Scale the generation of adversarial test cases and stress tests using programmatic automation for multi-turn scenarios.
  • Communicate evaluation findings and model readiness assessments to cross-functional partners to influence product decisions.

What we're looking for

  • Bachelor's degree in Computer Science, Machine Learning, Statistics, or a related field.
  • 4+ years of experience building or extending ML evaluation systems and designing quality assessment frameworks.
  • Experience independently defining evaluation architecture for AI systems where the unit of analysis is a conversation or session.
  • Experience designing adversarial or red-teaming test methodologies for ML models, specifically targeting multi-turn interactions.
  • Experience with Python and ML frameworks like PyTorch or TensorFlow in production or near-production settings.
  • Track record of owning technical direction for evaluation efforts across multiple features or product areas.
  • Experience evaluating user-facing AI features in consumer applications (preferred).
  • Familiarity with productivity software, creative tools, or agent orchestration frameworks such as LangChain and LangGraph (preferred).

More like this

Similar roles

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 88 days ago $184,700–$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc

San Diego, CA 88 days ago $175,000–$308,500
Machine Learning Generative AI LLM-as-a-judge Python MLflow Data Pipelines OCR Confidence Calibration Uncertainty Quantification Prompt Regression Testing offline metrics AUC Robustness Testing
7+ yrs exp

Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc

New York, NY 88 days ago $184,700–$324,800
Machine Learning Generative AI LLM-as-a-judge Python MLflow Data Pipelines OCR Confidence Calibration Uncertainty Quantification offline metrics Model Evaluation
7+ yrs exp