AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Sunnyvale, CA
Salary
$150,400–$277,600 / yr
Posted
55 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $240k
This role $214k
$132k most similar roles pay here $321k

This role pays less than 65% of similar roles. Most pay $208,833–$270,500 — the shaded band above. At the midpoint, this role pays about $214k versus about $240k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

As an AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation, you will join the Multimodal Intelligence Team to advance architectures, pre-training methods, and distillation techniques for multimodal models. You will investigate fundamental questions regarding model design, develop new training objectives, and create data strategies to transfer capabilities from large foundation models to smaller, efficient versions. Your daily work involves experimenting with dense, recurrent, state-space, and mixture-of-experts architectures while managing constraints like memory, latency, and energy for on-device deployment. You will utilize PyTorch or JAX and distributed training systems to build models across language, image, video, audio, and sensor-derived representations. The role focuses on the technical challenge of co-developing frontier models with efficient versions specifically optimized for Apple silicon while ensuring high performance in reasoning, instruction following, and multimodal understanding.

What you'll do

  • Develop and evaluate new multimodal model architectures, including dense, recurrent, state-space, and mixture-of-experts designs.
  • Design and implement pre-training objectives, data mixtures, and training curricula for large-scale foundation models.
  • Create teacher-student learning techniques to transfer capabilities from large models to smaller, efficient versions.
  • Optimize model architectures specifically for memory, latency, and energy constraints on Apple silicon.
  • Conduct large-scale experiments to improve multimodal understanding across text, image, video, audio, and sensor data.
  • Develop distillation methods that preserve reasoning and instruction-following capabilities under constrained model capacities.
  • Translate research findings into production-ready technologies for integration into the Apple ecosystem.

What we're looking for

  • Master's degree or equivalent practical experience in machine learning, computer science, or a related technical field.
  • Hands-on experience designing, implementing, and running large-scale pre-training experiments for large language models.
  • Experience with LLM pre-training topics including architecture, training objectives, data mixtures, tokenization, scaling, and optimization.
  • Proficiency with modern deep learning frameworks such as PyTorch or JAX and distributed training systems.
  • Experience evaluating pre-trained models across language understanding, reasoning, instruction following, or multimodal capabilities.
  • Strong understanding of transformer-based architectures and current approaches to efficient or scalable foundation-model training.
  • Preferred experience contributing to major foundation-model pre-training efforts or leading architecture experiments for large training runs.
  • Preferred experience with knowledge distillation techniques and designing teacher-student training pipelines.

More like this

Similar roles

Applied AI Scientist, Multimodal Intelligence

Apple Inc

Sunnyvale, CA 49 days ago $150,400$277,600
Python PyTorch JAX Deep Learning Computer Vision Generative AI Multimodal Foundation Models Data Pipelines On-device Machine Learning Machine Learning

Applied AI Scientist, Multimodal Intelligence

Apple Inc

Sunnyvale, CA 45 days ago $216,200$324,800
Generative AI Multimodal Foundation Models Deep Learning Python PyTorch JAX Computer Vision Data Pipelines On-device Machine Learning Vision-Language Models Video-Language Models
6+ yrs exp