Evaluation Science Lead

Apple Inc

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Austin, TX
Posted
2 days ago
Freshness
Confirmed live today

Market check

Salary context

How this pay compares to similar roles

Similar $203k
$136k $258k
below market most similar roles pay here above market

This listing doesn't post a salary. Most similar roles pay $169,955–$236,062.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 3595 open roles on FindRole.

Listed pay typically runs $166,600–$277,600 across 2772 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Evaluation Science Lead

JOB TITLE: Evaluation Science Lead The Evaluation Science Lead joins the Globalization Quality and Operations team to define how multilingual Globalization AI solutions are measured and validated across various services. As a strategic individual contributor, you will own the scientific rigor behind measuring AI performance across over 50 languages and dozens of markets. You will design statistically grounded evaluation frameworks, architecting scalable workflows that blend human annotation with automated systems like autograders and LLM-as-judge. Key responsibilities include designing long-term strategies, performing power analysis, and building drift-detection systems. You must possess proficiency in R, Python, and SQL to translate complex data into actionable recommendations. This role solves the technical challenge of capturing cultural and linguistic nuances through rigorous, culturally-informed evaluation to improve global AI language quality.

What you'll do

  • Define the long-term evaluation science strategy and roadmap for Globalization AI solutions across 50+ languages.
  • Design robust statistical methodologies including power analysis, sampling, and significance thresholds to evaluate models.
  • Architect scalable evaluation workflows that blend human annotation protocols with automated systems like LLM-as-judge.
  • Expand evaluation criteria to incorporate behavioral and user engagement signals beyond core linguistic quality.
  • Translate complex evaluation data into actionable recommendations and go/no-go evidence for Engineering and leadership.
  • Build continuous performance monitoring and drift-detection systems to identify model regressions.
  • Establish feedback loops with Quality Operations to refine evaluation rubrics as models evolve.

What we're looking for

  • 5+ years of experience in evaluation science, data science, or ML systems development owning evaluation systems at scale end-to-end.
  • Experience applying statistical methodology including sampling, significance testing, and confidence intervals.
  • Practical understanding of measurement validity principles.
  • Hands-on experience measuring annotator agreement, diagnosing divergence, and improving annotation protocols.
  • Proficiency with statistical tools and languages including R, Python, and SQL.
  • Experience designing evaluation frameworks adopted and scaled by operational teams.
  • Strong written and verbal communication skills for technical and non-technical audiences.
  • Master’s, PhD, or comparable experience in Statistics, Computational Linguistics, Computer Science, Psychometrics, Data Science, or related quantitative field (preferred).

More like this

Similar roles

Evaluation Science Lead

Apple Inc

Cupertino, CA 2 days ago $175,500–$311,700
Python R SQL Generative AI LLM-as-judge A/B Testing Causal Inference Data Science ML Systems Computational Linguistics drift-detection Synthetic Data Evaluation Sampling Confidence Intervals Significance Testing
5+ yrs exp

Director, Evaluation, Data Science and Insights

Apple Inc

Cupertino, CA 142 days ago $311,100–$496,900
Data Science Machine Learning LLM A/B Testing Human Evaluation Logging Synthetic Data Agentic Systems Statistics Computer Science Instrumentation Machine Translation
10+ yrs exp

Director, Evaluation, Data Science and Insights

Apple Inc

Cupertino, CA 127 days ago $311,100–$496,900
Data Science Machine Learning LLM A/B Testing Human Evaluation Logging Synthetic Data Agentic Systems Statistics Computer Science Instrumentation Machine Translation
10+ yrs exp

Director, Evaluation, Data Science and Insights

Apple Inc

Cupertino, CA 29 days ago $311,100–$496,900
Data Science Machine Learning LLM A/B Testing Human Evaluation Logging Synthetic Data Agentic Systems Statistics Computer Science Instrumentation Machine Translation
10+ yrs exp

AIML Data Scientist, Evaluation

Apple Inc

Cupertino, CA 16 days ago $150,400–$277,600
Python SQL LLM NLP Statistics Data Manipulation Natural Language Processing Metrics Development Exploratory Data Analysis Machine Learning Data Science
3+ yrs exp

AIML, Senior Lead Engineer, Evaluation

Apple Inc

Seattle, WA 108 days ago $205,400–$374,300
A/B Testing Service-Oriented Architecture Machine Learning AI Operating System Layer Statistical Foundations
10+ yrs exp