ML Engineer, Automated Evaluation and Adversarial Design

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Culver City, CA
Salary
$142,300–$263,300 / yr
Posted
165 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $221k
This role $203k
$127k most similar roles pay here $286k

This role pays less than 62% of similar roles. Most pay $187,850–$254,750 — the shaded band above. At the midpoint, this role pays about $203k versus about $221k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 3587 open roles on FindRole.

Listed pay typically runs $165,800–$277,600 across 2768 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · ML Engineer, Automated Evaluation and Adversarial Design

The ML Engineer - Automated Evaluation and Adversarial Design joins the Productivity and Machine Learning Evaluation team to ensure quality across AI-powered features in productivity and creative applications. This role focuses on building and scaling automated evaluation systems, designing adversarial test suites, and developing stress-testing methodologies for multi-turn conversation flows and agentic experiences. The candidate will translate qualitative quality into measurable assessments, create evaluation rubrics, and identify failure modes like context degradation or goal drift. Key responsibilities include managing automated and human evaluation alignment and providing readiness assessments to cross-functional partners. Required skills include Python and ML frameworks such as PyTorch or TensorFlow. Preferred experience includes familiarity with agent orchestration frameworks like LangChain, LangGraph, CrewAI, or AutoGen, along with observability tools like LangSmith, Braintrust, or Arize to evaluate multi-step agent runs and tool-use reliability.

What you'll do

  • Define and own automated evaluation approaches that translate qualitative quality into measurable metrics for multi-turn agentic experiences.
  • Build adversarial test suites targeting model failure modes like context loss, instruction forgetting, and cascading errors.
  • Develop and execute stress test protocols to validate performance under atypical conditions and complex tool-use sequences.
  • Ensure alignment between automated and human evaluation methods by identifying and resolving systematic discrepancies.
  • Scale the generation of adversarial test cases and stress tests using programmatic automation for multi-turn scenarios.
  • Communicate evaluation findings and model readiness assessments to cross-functional partners to influence product decisions.

What we're looking for

  • Bachelor's degree in Computer Science, Machine Learning, Statistics, or a related field.
  • 4+ years of experience building or extending ML evaluation systems and designing quality assessment frameworks.
  • Experience independently defining evaluation architecture for AI systems where the unit of analysis is a conversation or session.
  • Experience designing adversarial or red-teaming test methodologies for ML models, specifically targeting multi-turn interaction failures.
  • Experience with Python and ML frameworks like PyTorch or TensorFlow in production or near-production settings.
  • Track record of owning technical direction for evaluation efforts across multiple features or product areas.
  • Graduate degree in a relevant field (preferred).
  • Experience evaluating user-facing AI features, productivity tools, or agentic workflows (preferred).

More like this

Similar roles

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 88 days ago $184,700–$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Machine Learning Evaluation Engineer

Apple Inc

Sunnyvale, CA 36 days ago $150,400–$277,600
Machine Learning Computer Vision Python Data Analysis Statistical Modeling Deep Learning Data Pipelines Visualization Synthetic Data Data Augmentation Model Monitoring Generative AI Root-Cause Analysis Evaluation Frameworks
3+ yrs exp

Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc

San Diego, CA 88 days ago $175,000–$308,500
Machine Learning Generative AI LLM-as-a-judge Python MLflow Data Pipelines OCR Confidence Calibration Uncertainty Quantification Prompt Regression Testing offline metrics AUC Robustness Testing
7+ yrs exp