AI/ML Evaluation Specialist, Human Data

Apple Inc

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Cupertino, CA
Salary
$144,600–$263,800 / yr
Posted
71 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $220k
This role $204k
$129k most similar roles pay here $294k

This role pays less than 61% of similar roles. Most pay $185,643–$254,750 — the shaded band above. At the midpoint, this role pays about $204k versus about $220k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · AI/ML Evaluation Specialist, Human Data

AI/ML Evaluation Specialist, Human Data joins the Human-centered AI team within the Data Quality and Operations division. This role focuses on spearheading complex, multi-stakeholder operations involving data collection, curation, annotation, and human evaluation for services like Apple Music, App Store, TV+, Podcasts, and Books. The specialist will own the operational strategy for large-scale, multilingual human data programs, designing onboarding scaffolds, analyzing behavior patterns to identify automation opportunities, and enforcing quality frameworks. Key responsibilities include managing project timelines, building measurement frameworks, and developing statistical process controls to ensure high-quality deliverables. The role requires proficiency in Python, R, and SQL to analyze large datasets and build ETL services. Candidates must possess expertise in human-in-the-loop data collection for multimodal AI features while ensuring compliance with privacy and safety standards across engineering, legal, and product teams.

What you'll do

  • Execute end-to-end human data collection programs for multilingual, multimodal, and multi-turn AI features.
  • Design measurement frameworks and reporting systems to track spend, speed, inter-rater reliability, and volume.
  • Build and implement quality frameworks using statistical process controls to detect and remediate data degradation.
  • Analyze workflows to identify inefficiencies and implement automated pipelines or agentic workflows to improve scalability.
  • Develop onboarding, calibration, and performance programs for internal and external workforces.
  • Apply human-centered AI principles to reduce cognitive burden on annotators during task design.
  • Partner with legal and security teams to ensure all data collection complies with privacy and regulatory standards.
  • Serve as the primary liaison between engineering, product, legal, and procurement teams to align on shared standards.

What we're looking for

  • Bachelor's degree or higher in Cognitive Science, Linguistics, or a related field with an experimental or empirical component.
  • 4+ years of experience leading cross-team human data programs for AI/ML in NLP/NLU or generative AI environments.
  • Proficiency in programming and data languages including Python, R, and SQL to analyze large datasets and automate tasks.
  • Hands-on experience designing and managing 0→1 human-in-the-loop data collection, annotation, and evaluation initiatives.
  • Experience working with diverse data types such as speech, text, and multimodal content across multiple languages.
  • Expertise in end-to-end data annotation quality management, including statistical process controls and quality metrics.
  • Familiarity with privacy-preserving data handling practices and compliance frameworks.
  • Experience working cross-functionally with engineering, data science, legal, privacy, and third-party suppliers.

More like this

Similar roles

AI/ML Evaluation Specialist, Human Data

Apple Inc

Cupertino, CA 58 days ago $144,600$263,800
Python R SQL Generative AI NLP NLU Data Pipelines ETL Human-in-the-loop Machine Learning Data Annotation Statistical Process Control AI Safety Responsible AI Data Quality Frameworks
4+ yrs exp

Machine Learning Engineer, Agentic AI Evaluation Frameworks

Apple Inc

Cupertino, CA 2 days ago $184,700$277,600
Python LLM Generative AI RAG NLP Agentic AI CI/CD Machine Learning Data Pipelines Statistical Analysis Human-in-the-Loop Evaluation Frameworks Multi-turn Conversation Recommendation Systems
7+ yrs exp

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 65 days ago $184,700$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Senior Data Scientist, AIML Evaluation

Apple Inc

Cupertino, CA 76 days ago $150,400$277,600
Machine Learning LLMs Python SQL Spark R Scala Prompt Engineering Statistical Data Analysis Data Processing Experimentation

Machine Learning Engineer, AI & ML Evaluation Frameworks

Apple Inc

Cupertino, CA 94 days ago $150,400$277,600
Python LLMs Diffusion Models Machine Learning Deep Learning CI/CD Git Spark Kubernetes Airflow RAG Prompt Engineering Synthetic Data Generation Federated Learning Data Pipelines Model Interpretability AI Safety
3+ yrs exp