Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
San Diego, CA
Salary
$175,000–$308,500 / yr
Posted
88 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $217k
This role $242k
$154k most similar roles pay here $325k

This role pays more than 74% of similar roles. Most pay $187,850–$246,150 — the shaded band above. At the midpoint, this role pays about $242k versus about $217k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 3587 open roles on FindRole.

Listed pay typically runs $165,800–$277,600 across 2768 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Machine Learning Engineer, ML/GenAI Evaluation

Machine Learning Engineer, ML/GenAI Evaluation joins the Wallet team to establish evaluation criteria, metrics frameworks, and quality standards for models powering financial features. The role involves designing test frameworks, adversarial corpora, and benchmarks to identify failure modes before launch. Key responsibilities include developing methodologies for robustness testing, such as distribution shift and temporal drift, while owning end-to-end fairness evaluations across protected attributes. The engineer will evaluate generative and agentic model outputs using LLM-as-a-judge frameworks, human evaluation protocols, and prompt regression testing to measure hallucination rates and groundedness. Required skills include Python programming, experience with data pipelines, and experiment tracking tools like MLflow or W&B. The role addresses the critical challenge of ensuring accuracy, reliability, and fairness in high-stakes payment and commerce products by translating complex metrics into actionable product quality narratives for cross-functional stakeholders.

What does a Machine Learning Engineer earn in California?

Median $231150 from 200 postings across 30 companies.

See salary data

What you'll do

  • Define evaluation criteria and quality metrics for machine learning models powering Wallet features.
  • Design and maintain structured test sets covering diverse real-world scenarios, including edge cases and adversarial inputs.
  • Develop methodologies for robustness testing, including distribution shift, out-of-distribution generalization, and temporal drift.
  • Execute end-to-end fairness evaluations by building bias test suites across protected attributes and user populations.
  • Evaluate generative model outputs using LLM-as-a-judge frameworks to measure hallucination rates and groundedness.
  • Own the final quality sign-off process to determine if models meet readiness standards before shipping.
  • Translate evaluation results into actionable insights to guide model development priorities and product decisions.

What we're looking for

  • M.S. in Machine Learning, Computer Science, Statistics, Applied Mathematics, or a related technical field (preferred).
  • PhD in Computer Science, Data Science, Statistics, AI/ML, or a related field (preferred).
  • Bachelor's degree with 7+ years of experience in ML evaluation, model quality, or applied research.
  • 5+ years of hands-on ML experience with expertise in model evaluation, offline metrics design, and behavioral testing.
  • Strong programming skills in Python and fluency with evaluation tooling, data pipelines, and experiment tracking.
  • Proven ability to construct adversarial test suites and evaluate models for distribution shift, out-of-distribution generalization, and temporal drift.
  • Experience owning model quality sign-off in a cross-functional launch process.
  • Experience with structured/semi-structured document understanding, OCR, financial data extraction, or causal fairness approaches (preferred).

More like this

Similar roles

Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc

New York, NY 88 days ago $184,700–$324,800
Machine Learning Generative AI LLM-as-a-judge Python MLflow Data Pipelines OCR Confidence Calibration Uncertainty Quantification offline metrics Model Evaluation
7+ yrs exp

Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc

Austin, TX 88 days ago
Machine Learning Generative AI LLM-as-a-judge Python MLflow Data Pipelines OCR Model Evaluation Robustness Testing Fairness Evaluation Confidence Calibration Uncertainty Quantification
7+ yrs exp

Machine Learning Engineer, AI & ML Evaluation Frameworks

Apple Inc

Cupertino, CA 117 days ago $150,400–$277,600
Python LLMs Diffusion Models Machine Learning Deep Learning CI/CD Git Spark Kubernetes Airflow RAG Prompt Engineering Synthetic Data Generation Federated Learning Data Pipelines Model Interpretability AI Safety
3+ yrs exp

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 88 days ago $184,700–$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Machine Learning Evaluation Engineer

Apple Inc

Sunnyvale, CA 36 days ago $150,400–$277,600
Machine Learning Computer Vision Python Data Analysis Statistical Modeling Deep Learning Data Pipelines Visualization Synthetic Data Data Augmentation Model Monitoring Generative AI Root-Cause Analysis Evaluation Frameworks
3+ yrs exp

Machine Learning Engineer, AI Evaluation & LLM Systems

Apple Inc

Cupertino, CA 66 days ago $150,400–$225,300
Python C++ PyTorch TensorFlow JAX LLM Multimodal AI Generative AI Git CI/CD Distributed Computing Cloud Platforms Data Processing Statistical Analysis Machine Learning Software Engineering
1+ yrs exp