AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Sunnyvale, CA
Salary
$216,200–$324,800 / yr
Posted
45 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $240k
This role $270k
$167k most similar roles pay here $342k

This role pays more than 76% of similar roles. Most pay $208,833–$270,500 — the shaded band above. At the midpoint, this role pays about $270k versus about $240k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

As part of the Multimodal Intelligence Team, the AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation will advance architectures, pre-training methods, and distillation techniques for multimodal models. The role involves investigating fundamental questions regarding model design, training objectives, data strategies, and teacher–student learning to transfer capabilities from large foundation models to smaller, efficient versions. You will develop and evaluate new architectures, including dense, recurrent, state-space, and mixture-of-experts models, while focusing on inference efficiency for Apple silicon. Key responsibilities include managing the full lifecycle of model development, from data mixtures and tokenization to scaling and evaluation across language, image, video, audio, and sensor-derived representations. Required skills include proficiency in PyTorch or JAX, distributed training systems, and experience with transformer-based architectures to solve challenges related to memory, latency, and energy constraints for on-device deployment.

What you'll do

  • Design and evaluate new multimodal foundation model architectures, including dense, recurrent, state-space, and mixture-of-experts models.
  • Develop pre-training objectives, data mixtures, and training curricula for large-scale multimodal models.
  • Implement teacher-student learning techniques to transfer capabilities from large models to smaller, efficient versions.
  • Optimize model architectures specifically for memory, latency, and energy constraints on Apple silicon.
  • Conduct large-scale experiments to improve reasoning, instruction following, and multimodal understanding in compact models.
  • Translate research findings into production-ready technologies for integration across the Apple ecosystem.
  • Publish research findings and open-source selected models, tools, and evaluation artifacts.

What we're looking for

  • Master's degree or equivalent practical experience in machine learning, computer science, or a related technical field.
  • Minimum of 6 years of relevant experience in the field.
  • Hands-on experience designing, implementing, and running large-scale pre-training experiments for large language models.
  • Experience with LLM pre-training topics including architecture, training objectives, data mixtures, tokenization, scaling, and optimization.
  • Proficiency with modern deep learning frameworks such as PyTorch or JAX and distributed training systems.
  • Experience evaluating pre-trained models across language understanding, reasoning, instruction following, or multimodal capabilities.
  • Strong understanding of transformer-based architectures and current approaches to efficient or scalable foundation-model training.
  • Preferred experience in knowledge distillation, multimodal models, inference efficiency, and research publications or influential open-source contributions.

More like this

Similar roles

Machine Learning Researcher, Foundation Models

Apple Inc

New York, NY 79 days ago $150,400$277,600
Python PyTorch JAX TensorFlow Deep Learning Foundation Models Large Language Models Reinforcement Learning Vision-Language Modeling Video Generation Data Pipelines Reward Modeling Multimodal Models On-policy Distillation

Machine Learning Researcher, Foundation Models

Apple Inc

Cupertino, CA 79 days ago $150,400$277,600
Python PyTorch JAX TensorFlow Deep Learning Foundation Models Large Language Models Reinforcement Learning Vision-Language Modeling Video Generation Data Pipelines Reward Modeling On-policy Distillation Multimodal Models

Applied AI Scientist, Multimodal Intelligence

Apple Inc

Sunnyvale, CA 49 days ago $150,400$277,600
Python PyTorch JAX Deep Learning Computer Vision Generative AI Multimodal Foundation Models Data Pipelines On-device Machine Learning Machine Learning