Evaluation & Insights Machine Learning Engineer

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Cupertino, CA
Salary
$184,700–$324,800 / yr
Posted
65 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $225k
This role $255k
$162k most similar roles pay here $342k

This role pays more than 78% of similar roles. Most pay $194,283–$254,750 — the shaded band above. At the midpoint, this role pays about $255k versus about $225k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Evaluation & Insights Machine Learning Engineer

The Evaluation & Insights Machine Learning Engineer joins the Human-Centered AI team to evaluate and improve AI systems by combining data science, model behavior analysis, and qualitative insights. This role involves architecting evaluation suites for LLMs and multimodal models, developing advanced scoring frameworks like LLM-as-a-judge, and identifying edge cases in reasoning, factuality, and safety. The engineer will translate qualitative failure modes into quantifiable loss patterns to inform prompt engineering, RAG strategies, and model fine-tuning. Key responsibilities include building MLOps workflows, designing distributed inference pipelines using Ray or vLLM, and mapping error taxonomies through embedding-based clustering. Required skills include Python, PyTorch, JAX, Hugging Face, and experience with RLHF/DPO. The role addresses the challenge of ensuring AI experiences are reliable, safe, and aligned with human expectations across various product features.

What you'll do

  • Architect and execute comprehensive evaluation suites for LLMs and multimodal models to identify edge cases in reasoning, factuality, and safety.
  • Develop deterministic, heuristic, and LLM-assisted scoring frameworks to quantify human-perceived quality metrics like helpfulness and hallucination rates.
  • Translate qualitative failure modes into quantifiable loss patterns and actionable data-mixture adjustments for model training and inference.
  • Use advanced ML techniques like embedding-based clustering and perturbation analysis to map error taxonomies in model outputs.
  • Build automated evaluation pipelines using LLMs to assess outputs at scale with high correlation to human baselines.
  • Develop MLOps workflows to codify metrics, automate regression testing, and integrate assessments into CI/CD pipelines.
  • Architect scalable, distributed inference and processing pipelines for high-throughput model evaluation and automated annotation.
  • Partner with engineering teams to refine model behavior through prompt engineering, RAG strategies, and fine-tuning based on evaluation telemetry.

What we're looking for

  • Bachelor's or Master's degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or a related technical field.
  • 8+ years of relevant industry experience in ML Engineering or Applied Research.
  • Advanced proficiency in Python and modern deep learning ecosystems including PyTorch, JAX, and Hugging Face.
  • Proven experience building scalable ML inference pipelines, model evaluation workflows, and structured rating frameworks for large-scale AI systems.
  • Hands-on experience developing, fine-tuning, or evaluating LLMs, multimodal models, and NLP systems.
  • Deep familiarity with AI quality metrics, hallucination detection, model alignment (RLHF/DPO), and LLM-as-a-judge frameworks.
  • Strong familiarity with advanced prompt engineering, RAG architectures, vector databases, and fine-tuning techniques.
  • Experience building internal tools or automated pipelines for ML workflows using platforms like MLflow or Weights & Biases.

More like this

Similar roles

Machine Learning Evaluation Engineer

Apple Inc

Sunnyvale, CA 13 days ago $150,400$277,600
Machine Learning Computer Vision Python Data Analysis Statistical Modeling Deep Learning Data Pipelines Visualization Synthetic Data Data Augmentation Model Monitoring Generative AI Root-Cause Analysis Evaluation Frameworks
3+ yrs exp

Machine Learning Engineer, AI & ML Evaluation Frameworks

Apple Inc

Cupertino, CA 94 days ago $150,400$277,600
Python LLMs Diffusion Models Machine Learning Deep Learning CI/CD Git Spark Kubernetes Airflow RAG Prompt Engineering Synthetic Data Generation Federated Learning Data Pipelines Model Interpretability AI Safety
3+ yrs exp

Senior Machine Learning Engineer, Evaluation

Apple Inc

Cupertino, CA 91 days ago $216,200$394,000
Machine Learning LLM Python PyTorch Reward Modeling RLHF SFT DPO Prompt Optimization Distributed Systems Agentic Systems Simulation Environments ML Infrastructure Data Generation
8+ yrs exp

Machine Learning Engineer

Booz Allen Hamilton

Arlington, VA +2 81 days ago $77,600$176,000
Python TensorFlow PyTorch scikit-learn C++ Rust Java LLMs LangChain LangGraph Computer Vision Generative AI Deep Learning Cursor Supervised Learning Unsupervised Learning
2+ yrs exp

Machine Learning Engineer

Booz Allen Hamilton

Arlington, VA +2 81 days ago $77,600$176,000
Python TensorFlow PyTorch scikit-learn C++ Rust Java Deep Learning Computer Vision Generative AI LLMs LangChain LangGraph MCP Cursor Supervised Learning Unsupervised Learning
2+ yrs exp