AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Seattle, WA
Salary
$205,400–$308,500 / yr
Posted
45 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $240k
This role $257k
$169k most similar roles pay here $323k

This role pays more than 72% of similar roles. Most pay $208,833–$270,500 — the shaded band above. At the midpoint, this role pays about $257k versus about $240k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

As an AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation, you will join the Multimodal Intelligence Team to advance foundation models for Apple experiences. You will investigate fundamental questions regarding how multimodal models should be designed, trained, and distilled, focusing on transferring capabilities from large models to smaller, efficient versions for on-device deployment. Your daily work involves developing new model architectures, pre-training objectives, data strategies, and teacher–student learning techniques while managing constraints like memory, latency, and energy. You will utilize PyTorch or JAX and distributed training systems to conduct experiments involving transformer-based architectures, mixture-of-experts, state-space models, and multimodal data including language, images, video, audio, and sensor-derived representations. The role addresses the technical challenge of creating high-performing, compact models that maintain reasoning and instruction following capabilities under strict hardware constraints on Apple silicon.

What you'll do

  • Develop and evaluate new multimodal model architectures, including dense, recurrent, state-space, and mixture-of-experts designs.
  • Design and implement pre-training objectives, data mixtures, and training curricula for large-scale foundation models.
  • Create teacher-student learning techniques to transfer capabilities from large models to smaller, efficient versions.
  • Optimize model architectures specifically for memory, latency, and energy constraints on Apple silicon.
  • Conduct large-scale experiments to improve multimodal understanding across language, image, video, and audio data.
  • Develop distillation methods that preserve reasoning and instruction following under constrained model capacities.
  • Translate research findings into production-ready technologies for integration into the Apple ecosystem.

What we're looking for

  • Master's degree or equivalent practical experience in machine learning, computer science, or a related technical field.
  • Minimum of 6 years of relevant experience in the field.
  • Hands-on experience designing, implementing, and running large-scale pre-training experiments for large language models.
  • Proficiency with modern deep learning frameworks such as PyTorch or JAX and distributed training systems.
  • Strong understanding of transformer-based architectures and current approaches to efficient or scalable foundation-model training.
  • Experience evaluating pre-trained models across language understanding, reasoning, instruction following, or multimodal capabilities.
  • Experience with knowledge distillation techniques, including offline, off-policy, on-policy, or sequence-level distillation.
  • Experience with multimodal models spanning language, vision, video, audio, or other sensor modalities.

More like this

Similar roles

Machine Learning Researcher, Foundation Models

Apple Inc

New York, NY 79 days ago $150,400$277,600
Python PyTorch JAX TensorFlow Deep Learning Foundation Models Large Language Models Reinforcement Learning Vision-Language Modeling Video Generation Data Pipelines Reward Modeling Multimodal Models On-policy Distillation

Machine Learning Researcher, Foundation Models

Apple Inc

Cupertino, CA 79 days ago $150,400$277,600
Python PyTorch JAX TensorFlow Deep Learning Foundation Models Large Language Models Reinforcement Learning Vision-Language Modeling Video Generation Data Pipelines Reward Modeling On-policy Distillation Multimodal Models