AI/ML Evaluation Specialist, Human Data

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Cupertino, CA
Salary
$144,600–$263,800 / yr
Posted
58 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $220k
This role $204k
$129k most similar roles pay here $294k

This role pays less than 61% of similar roles. Most pay $185,643–$254,750 — the shaded band above. At the midpoint, this role pays about $204k versus about $220k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · AI/ML Evaluation Specialist, Human Data

The AI/ML Evaluation Specialist, Human Data joins the Human-centered AI team within the Data Quality and Operations division. This role focuses on spearheading complex, multi-stakeholder operations involving data collection, curation, annotation, and human evaluation for services like Apple Music, App Store, TV+, Podcasts, and Books. You will own the operational strategy for large-scale, multilingual human data programs, designing onboarding scaffolds, analyzing behavior patterns to identify automation opportunities, and enforcing quality frameworks. The role involves building data pipelines, ETL services, and measurement frameworks to ensure generative AI features are safe and reliable. Key responsibilities include managing project timelines, costs, and compliance with privacy standards while acting as the connective tissue between engineering, data science, and legal teams. Required skills include proficiency in Python, R, and SQL, along with expertise in NLP/NLU environments and human-in-the-loop evaluation strategies.

What you'll do

  • Execute end-to-end human data collection programs for multilingual, multimodal, and multi-turn AI features.
  • Design measurement frameworks and reporting systems to track spend, speed, inter-rater reliability, and volume.
  • Build and implement quality frameworks using statistical process controls to detect and remediate data degradation.
  • Analyze workflows to identify inefficiencies and implement automated pipelines or agent-in-the-loop solutions.
  • Develop onboarding, calibration, and performance programs for internal and external workforces.
  • Apply human-centered AI principles to reduce cognitive burden on annotators during task design.
  • Ensure all data collection programs comply with legal, privacy, and security standards.
  • Serve as the primary liaison between engineering, product, legal, and procurement teams to align on shared standards.

What we're looking for

  • Bachelor's degree or higher in Cognitive Science, Linguistics, or a related field with an experimental or empirical component.
  • 4+ years of experience leading cross-team human data programs for AI/ML in NLP/NLU or generative AI environments.
  • Proficiency in programming and data languages including Python, R, and SQL to analyze large datasets and automate tasks.
  • Hands-on experience managing 0→1 human-in-the-loop data collection, annotation, and evaluation initiatives using agentic workflows.
  • Experience working with diverse data types such as speech, text, and multimodal content across multiple languages.
  • Expertise in end-to-end data annotation quality management, including statistical process controls and quality metrics.
  • Familiarity with privacy-preserving data handling practices and compliance frameworks.
  • Proven ability to work cross-functionally with engineering, data science, legal, privacy, and third-party suppliers.

More like this

Similar roles

AI/ML Evaluation Specialist, Human Data

Apple Inc

Cupertino, CA 71 days ago $144,600$263,800
Python R SQL Generative AI NLP NLU Data Pipelines ETL Human-in-the-loop Machine Learning Data Annotation Statistical Process Control AI Safety Responsible AI Data Quality Frameworks
4+ yrs exp

Machine Learning Engineer, Agentic AI Evaluation Frameworks

Apple Inc

Cupertino, CA 2 days ago $184,700$277,600
Python LLM Generative AI RAG NLP Agentic AI CI/CD Machine Learning Data Pipelines Statistical Analysis Human-in-the-Loop Evaluation Frameworks Multi-turn Conversation Recommendation Systems
7+ yrs exp

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 65 days ago $184,700$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Senior Data Scientist, AIML Evaluation

Apple Inc

Cupertino, CA 76 days ago $150,400$277,600
Machine Learning LLMs Python SQL Spark R Scala Prompt Engineering Statistical Data Analysis Data Processing Experimentation

Machine Learning Engineer, AI & ML Evaluation Frameworks

Apple Inc

Cupertino, CA 94 days ago $150,400$277,600
Python LLMs Diffusion Models Machine Learning Deep Learning CI/CD Git Spark Kubernetes Airflow RAG Prompt Engineering Synthetic Data Generation Federated Learning Data Pipelines Model Interpretability AI Safety
3+ yrs exp