Machine Learning Evaluation Engineer

Apple Inc

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Sunnyvale, CA
Salary
$150,400–$277,600 / yr
Posted
1 day ago
Freshness
Confirmed live today

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $226k
This role $214k
$135k most similar roles pay here $293k

This role pays more than 52% of similar roles. Most pay $198,000–$254,750 — the shaded band above. At the midpoint, this role pays about $214k versus about $226k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 3518 open roles on FindRole.

Listed pay typically runs $166,600–$277,600 across 2727 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Machine Learning Evaluation Engineer

The Machine Learning Evaluation Engineer joins the Data, Analytics and Quality team to evaluate and elevate advanced sensing technologies. This role involves leading the benchmarking of state-of-the-art multi-modal models by designing comprehensive evaluation systems, including metric design, large-scale data preparation, and automated failure analysis. You will build cloud-based, LLM-powered workflows and dashboards to streamline reporting and isolate problems within complex algorithm stacks. Key responsibilities include curating high-quality datasets, performing deep failure analysis to identify root causes in models or prompts, and providing data-driven insights to guide model training. You will utilize Python, computer vision, and vision-language models to assess model capabilities. The work focuses on solving technical challenges in multi-modal foundation models and task-specific CV/ML algorithms to ensure high-quality performance across various product features.

What you'll do

  • Design and build scalable evaluation pipelines for multi-modal models across text, image, and video tasks.
  • Develop and calibrate subjective evaluation methods using LLM-as-Judge and human grading techniques.
  • Perform deep failure analysis to identify root causes in models, prompts, and upstream software components.
  • Create failure taxonomies and automated dashboards to streamline reporting and iterative experimentation.
  • Curate high-quality evaluation datasets using auto-labeling, manual annotation, and synthetic data generation.
  • Analyze dataset quality, diversity, and coverage gaps to ensure evaluation sets reflect real-world usage.
  • Provide data-driven insights and metrics to model training teams to guide fine-tuning and iterations.

What we're looking for

  • Bachelor's degree and a minimum of 3 years of relevant industry experience.
  • 3+ years of applied experience in Machine Learning, Computer Vision, or AI System Evaluation.
  • Deep understanding of core Machine Learning principles including probability, statistics, data distributions, and model bias/variance.
  • Experience analyzing data quality, diversity, representativeness, and coverage gaps in large datasets to inform curation strategies.
  • Applied experience evaluating generative models, including prompt tuning, subjective evaluation of open-ended outputs, and human-in-the-loop methodologies.
  • Proven track record of defining robust metrics and designing evaluation frameworks for foundation models using LLM/VLM-as-a-Judge methodologies.
  • Strong proficiency in Python and building automated pipelines for model endpoint integration, multi-step evaluation orchestration, and logging.
  • Theoretical and practical understanding of Computer Vision and Vision-Language Models (preferred).

More like this

Similar roles

Machine Learning Evaluation Engineer

Apple Inc

Sunnyvale, CA 40 days ago $150,400–$277,600
Machine Learning Computer Vision Python Data Analysis Statistical Modeling Deep Learning Data Pipelines Visualization Synthetic Data Data Augmentation Model Monitoring Generative AI Root-Cause Analysis Evaluation Frameworks
3+ yrs exp

Machine Learning Engineer, AI & ML Evaluation Frameworks

Apple Inc

Cupertino, CA 121 days ago $150,400–$277,600
Python LLMs Diffusion Models Machine Learning Deep Learning CI/CD Git Spark Kubernetes Airflow RAG Prompt Engineering Synthetic Data Generation Federated Learning Data Pipelines Model Interpretability AI Safety
3+ yrs exp

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 92 days ago $184,700–$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Machine Learning Engineer, AI Evaluation & LLM Systems

Apple Inc

Cupertino, CA 70 days ago $150,400–$225,300
Python C++ PyTorch TensorFlow JAX LLM Multimodal AI Generative AI Git CI/CD Distributed Computing Cloud Platforms Data Processing Statistical Analysis Machine Learning Software Engineering
1+ yrs exp