AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Seattle, WA
Salary
$142,300–$263,300 / yr
Posted
55 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $240k
This role $203k
$123k most similar roles pay here $322k

This role pays less than 78% of similar roles. Most pay $208,845–$270,500 — the shaded band above. At the midpoint, this role pays about $203k versus about $240k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation joins the Multimodal Intelligence Team to advance foundation models for Apple experiences. This role involves investigating how multimodal models should be designed, trained, and distilled to balance broad intelligence with memory, latency, energy, and privacy requirements for on-device deployment. You will develop and evaluate new model architectures, pre-training objectives, data strategies, optimization methods, and teacher–student learning techniques. The work focuses on co-developing frontier models and efficient models for Apple silicon by integrating hardware constraints into the research process from early experimentation through final evaluation. Key technical areas include transformer-based architectures, mixture-of-experts, state-space models, and multimodal pre-training across language, image, video, audio, and sensor data. Required skills include proficiency in PyTorch or JAX, distributed training systems, and experience with large-scale pre-training experiments for large language models.

What you'll do

  • Develop and evaluate new model architectures, pre-training objectives, and data strategies for multimodal foundation models.
  • Design teacher-student learning techniques to transfer capabilities from large models to smaller, efficient versions.
  • Conduct large-scale pre-training experiments using PyTorch or JAX across various modalities like text, image, video, and audio.
  • Optimize model architectures specifically for memory, latency, and energy constraints on Apple silicon.
  • Research and implement advanced distillation methods including offline, on-policy, and representation transfer techniques.
  • Conduct research on compute-optimal scaling, including data mixtures, curricula, and training objectives.
  • Translate successful research findings into production-ready technologies for integration across the Apple ecosystem.

What we're looking for

  • Master's degree or equivalent practical experience in machine learning, computer science, or a related technical field.
  • Hands-on experience designing, implementing, and running large-scale pre-training experiments for large language models.
  • Experience with LLM pre-training topics including architecture, training objectives, data mixtures, tokenization, scaling, and optimization.
  • Proficiency with modern deep learning frameworks such as PyTorch or JAX and distributed training systems.
  • Strong understanding of transformer-based architectures and current approaches to efficient or scalable foundation-model training.
  • Experience evaluating pre-trained models across language understanding, reasoning, instruction following, or multimodal capabilities.
  • Preferred experience contributing to major foundation-model pre-training efforts or leading architecture experiments for large training runs.
  • Preferred experience with knowledge distillation techniques and designing teacher-student training pipelines.

More like this

Similar roles

Applied AI Scientist, Multimodal Intelligence

Apple Inc

Sunnyvale, CA 49 days ago $150,400$277,600
Python PyTorch JAX Deep Learning Computer Vision Generative AI Multimodal Foundation Models Data Pipelines On-device Machine Learning Machine Learning

Applied AI Scientist, Multimodal Intelligence

Apple Inc

Seattle, WA 45 days ago $205,400$308,500
Generative AI Multimodal Foundation Models Deep Learning Python PyTorch JAX Computer Vision Data Pipelines On-device Machine Learning Data Curation
6+ yrs exp