Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Austin, TX
Posted
88 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $217k
$161k most similar roles pay here $275k

This listing doesn't post a salary. Most similar roles pay $187,850–$246,150.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 3587 open roles on FindRole.

Listed pay typically runs $165,800–$277,600 across 2768 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Machine Learning Engineer, ML/GenAI Evaluation

As a Machine Learning Engineer, ML/GenAI Evaluation, you will join the team responsible for Wallet, Payments, and Commerce features to establish evaluation criteria, metrics frameworks, and quality standards for machine learning models. You will manage the full evaluation lifecycle by designing test frameworks, adversarial corpora, and benchmarks that account for diverse document formats, languages, and edge cases. Your daily work involves developing methodologies for robustness testing, including distribution shift and temporal drift, while owning end-to-end fairness evaluations to gate model launches. You will evaluate generative and agentic outputs using LLM-as-a-judge frameworks and prompt regression testing. The role requires proficiency in Python and experience with tools like MLflow or W&B. You will solve critical problems regarding model reliability, hallucination rates, and groundedness within the specific domain of financial features and document understanding.

What does a Machine Learning Engineer earn?

Median $230800 from 335 postings across 52 companies.

See salary data

What you'll do

  • Define evaluation criteria and quality metrics for machine learning models powering Wallet features.
  • Design and maintain structured test sets covering diverse real-world scenarios, including edge cases and adversarial inputs.
  • Develop methodologies for robustness testing, including distribution shift, out-of-distribution generalization, and temporal drift.
  • Own the end-to-end fairness evaluation process by building bias test suites across protected attributes and user populations.
  • Evaluate generative model outputs using LLM-as-a-judge frameworks to measure hallucination rates and groundedness.
  • Build persona-stratified benchmarks that reflect a global user base across various locales and spending patterns.
  • Provide final quality sign-off on model readiness before features are launched to the public.
  • Translate evaluation results into actionable insights to guide model development priorities and product decisions.

What we're looking for

  • M.S. in Machine Learning, Computer Science, Statistics, Applied Mathematics, or a related technical field (preferred).
  • PhD in Computer Science, Data Science, Statistics, AI/ML, or a related field (preferred).
  • Bachelor's degree with 7+ years of experience in ML evaluation, model quality, or applied research.
  • 5+ years of hands-on ML experience with expertise in model evaluation, offline metrics design, and behavioral testing.
  • Proven ability to design evaluation frameworks for production systems including precision-recall tradeoffs, calibration, fairness, and task-specific dimensions.
  • Experience testing for distribution shift, out-of-distribution generalization, and temporal drift in deployed models.
  • Ability to construct adversarial test suites, aggressor scenarios, and edge-case corpora to identify model failure modes.
  • Proficiency in Python and experience with evaluation tooling, data pipelines, and experiment tracking.
  • Experience with structured/semi-structured document understanding, OCR, or financial data extraction (preferred).
  • Experience with Bayesian/causal approaches to data generation or fairness evaluation (preferred).
  • Experience evaluating models under privacy constraints or on-device inference settings (preferred).
  • Familiarity with confidence calibration techniques and uncertainty quantification (preferred).
  • Background in financial services, fintech, or consumer payment products (preferred).

More like this

Similar roles

Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc

San Diego, CA 88 days ago $175,000–$308,500
Machine Learning Generative AI LLM-as-a-judge Python MLflow Data Pipelines OCR Confidence Calibration Uncertainty Quantification Prompt Regression Testing offline metrics AUC Robustness Testing
7+ yrs exp

Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc

New York, NY 88 days ago $184,700–$324,800
Machine Learning Generative AI LLM-as-a-judge Python MLflow Data Pipelines OCR Confidence Calibration Uncertainty Quantification offline metrics Model Evaluation
7+ yrs exp

Machine Learning Engineer, AI & ML Evaluation Frameworks

Apple Inc

Cupertino, CA 118 days ago $150,400–$277,600
Python LLMs Diffusion Models Machine Learning Deep Learning CI/CD Git Spark Kubernetes Airflow RAG Prompt Engineering Synthetic Data Generation Federated Learning Data Pipelines Model Interpretability AI Safety
3+ yrs exp

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 88 days ago $184,700–$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Machine Learning Engineer, AI Evaluation & LLM Systems

Apple Inc

Cupertino, CA 66 days ago $150,400–$225,300
Python C++ PyTorch TensorFlow JAX LLM Multimodal AI Generative AI Git CI/CD Distributed Computing Cloud Platforms Data Processing Statistical Analysis Machine Learning Software Engineering
1+ yrs exp