ML Engineer, Automated Evaluation and Adversarial Design

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Seattle, WA
Salary
$142,300–$263,300 / yr
Posted
165 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $221k
This role $203k
$127k most similar roles pay here $286k

This role pays less than 62% of similar roles. Most pay $187,850–$254,750 — the shaded band above. At the midpoint, this role pays about $203k versus about $221k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 3587 open roles on FindRole.

Listed pay typically runs $165,800–$277,600 across 2768 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · ML Engineer, Automated Evaluation and Adversarial Design

ML Engineer - Automated Evaluation and Adversarial Design joins the Productivity and Machine Learning Evaluation team to ensure quality across AI-powered features. This role focuses on building and scaling automated evaluation systems, designing adversarial test suites, and developing stress-testing methodologies for multi-turn conversation flows and agentic experiences. The engineer will translate qualitative quality into measurable assessments, identify failure modes like context degradation or goal drift, and create robust evaluation frameworks and rubrics. Key responsibilities include managing automated and human evaluation alignment, scaling test case generation, and providing readiness assessments to cross-functional partners. Required skills include Python, PyTorch, or TensorFlow, along with experience in ML evaluation systems, red-teaming methodologies, and agentic workflows. The role addresses the technical challenge of evaluating complex, multi-step interaction chains rather than just single outputs within productivity and creative applications.

What you'll do

  • Define and own automated evaluation approaches that translate qualitative quality into measurable metrics for single-turn and multi-turn AI experiences.
  • Build adversarial test suites to identify model failure modes such as context loss, instruction forgetting, and cascading errors.
  • Develop and execute stress test protocols to validate performance under atypical conditions like extended conversation lengths and complex tool-use sequences.
  • Ensure alignment between automated and human evaluation methods by identifying and resolving systematic disagreements.
  • Scale the generation of adversarial test cases and stress tests using programmatic automation for multi-turn scenarios.
  • Communicate evaluation findings and model readiness assessments to cross-functional partners to influence product decisions.

What we're looking for

  • Bachelor's degree in Computer Science, Machine Learning, Statistics, or a related field.
  • 4+ years of experience building or extending ML evaluation systems and designing quality assessment frameworks.
  • Experience independently defining evaluation architecture for AI systems where the unit of analysis is a conversation or session.
  • Experience designing adversarial or red-teaming test methodologies for ML models, specifically targeting multi-turn interactions.
  • Experience with Python and ML frameworks like PyTorch or TensorFlow in production or near-production settings.
  • Track record of owning technical direction for evaluation efforts across multiple features or product areas.
  • Graduate degree in a relevant field (preferred).
  • Experience evaluating user-facing AI features, productivity tools, or agentic workflows (preferred).

More like this

Similar roles

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 88 days ago $184,700–$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Machine Learning Evaluation Engineer

Apple Inc

Sunnyvale, CA 36 days ago $150,400–$277,600
Machine Learning Computer Vision Python Data Analysis Statistical Modeling Deep Learning Data Pipelines Visualization Synthetic Data Data Augmentation Model Monitoring Generative AI Root-Cause Analysis Evaluation Frameworks
3+ yrs exp

Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc

San Diego, CA 88 days ago $175,000–$308,500
Machine Learning Generative AI LLM-as-a-judge Python MLflow Data Pipelines OCR Confidence Calibration Uncertainty Quantification Prompt Regression Testing offline metrics AUC Robustness Testing
7+ yrs exp