Research Scientist / Engineer, Foundation Model Evaluation

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Cupertino, CA
Salary
$184,700–$324,800 / yr
Posted
148 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $221k
This role $255k
$157k most similar roles pay here $343k

This role pays more than 81% of similar roles. Most pay $186,787–$254,750 — the shaded band above. At the midpoint, this role pays about $255k versus about $221k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Research Scientist / Engineer, Foundation Model Evaluation

Research Scientist / Engineer, Foundation Model Evaluation joins the team developing frontier foundation models optimized for Apple silicon and integrated OS experiences. This role focuses on building evaluation systems that provide actionable signals to drive model improvement across reasoning, code, knowledge, and agentic workflows. You will design and implement benchmarks, develop product-aligned evaluation methods, and research state-of-the-art techniques like model-based judging and contamination-resistant designs. Day-to-day tasks include building reusable scorer libraries, performing experimental analysis to identify development gaps, and collaborating with training and product teams to translate insights into data decisions. The role requires proficiency in Python and machine learning frameworks such as PyTorch or JAX. You will solve complex problems in reward modeling, handling sparse rewards, and aligning models for both creative tasks and precise action-taking workflows within the foundation model lifecycle.

What you'll do

  • Design and implement evaluation benchmarks, metrics, and test suites for reasoning, knowledge, code, and agentic workflows.
  • Develop evaluation methods that capture model behavior in real product settings to predict user-perceived quality.
  • Research and apply state-of-the-art techniques including scoring frameworks and model-based judging.
  • Build reusable tools, scorer libraries, and analysis frameworks to scale across the team's benchmark portfolio.
  • Execute rigorous experiments comparing model capabilities and perform gap analyses to guide development priorities.
  • Translate evaluation findings into actionable insights for training strategies and data decisions.
  • Manage third-party vendors for benchmarking and coordinate with annotation teams to assess model quality.

What we're looking for

  • MS or PhD in Computer Science, Machine Learning, Natural Language Processing, or a related technical field (equivalent practical experience considered).
  • 3+ years of experience in AI model evaluation, NLP, or related areas like natural language generation, information retrieval, or conversational AI.
  • Strong fundamentals in machine learning, natural language processing, and statistical analysis.
  • Proficiency in Python and experience with ML frameworks such as PyTorch or JAX.
  • Demonstrated ability to translate research insights into practical implementations.
  • Strong experimental design skills to conduct rigorous comparisons and draw valid conclusions.
  • Ability to communicate technical results clearly to provide actionable recommendations for cross-functional partners.
  • Experience evaluating large language models, including benchmark design, model-based judging, and human evaluation methodologies.

More like this

Similar roles

Distinguished Engineer, AIML Foundation Models

Apple Inc

Cupertino, CA 143 days ago $311,100$496,900
Python PyTorch JAX TensorFlow Deep Learning Natural Language Processing Multi-modal Understanding Information Retrieval Reward Modeling Foundation Models Machine Learning On-device Intelligence

ML Engineer, Foundation Models

Apple Inc

Cupertino, CA 70 days ago $150,400$277,600
LLM Multi-modal LLM Python PyTorch JAX TensorFlow Deep Learning Synthetic Data Reward Modeling Preference Learning Data Pipelines Agentic Systems Data-centric AI Pre-training Post-training

Machine Learning Researcher, Foundation Models

Apple Inc

Cupertino, CA 79 days ago $150,400$277,600
Python PyTorch JAX TensorFlow Deep Learning Foundation Models Large Language Models Reinforcement Learning Vision-Language Modeling Video Generation Data Pipelines Reward Modeling On-policy Distillation Multimodal Models

Machine Learning Researcher, Foundation Models

Apple Inc

New York, NY 79 days ago $150,400$277,600
Python PyTorch JAX TensorFlow Deep Learning Foundation Models Large Language Models Reinforcement Learning Vision-Language Modeling Video Generation Data Pipelines Reward Modeling Multimodal Models On-policy Distillation

Machine Learning Researcher, Foundation Models

Apple Inc

Cupertino, CA 8 days ago $262,500$394,000
Deep Learning Python PyTorch JAX TensorFlow Natural Language Processing Multi-modal Understanding Information Retrieval Foundation Models Machine Learning

AIML Site Lead & Lead Researcher, Foundation Models

Apple Inc

Seattle, WA 149 days ago $205,400$374,300
Foundation Models Deep Learning Python PyTorch JAX TensorFlow Natural Language Processing Multi-modal Understanding Information Retrieval On-device Intelligence Reward Modeling Software Engineering