Multimodal AI Researcher

Apple Inc

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Sunnyvale, CA
Salary
$150,400–$277,600 / yr
Posted
30 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $222k
This role $214k
$135k most similar roles pay here $293k

This role pays less than 55% of similar roles. Most pay $189,462–$255,225 — the shaded band above. At the midpoint, this role pays about $214k versus about $222k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Multimodal AI Researcher

As a Multimodal AI Researcher within the Video Computer Vision organization, you will join an applied research team focused on developing breakthrough technologies for future products. You will tackle fundamental challenges in multimodal generative AI and agents by building foundation models that integrate real-time sensor data, such as video and audio, with other modalities like text. Your daily work involves conducting algorithm research, defining data requirements, and establishing validation strategies to move research into practical features. The role requires expertise in LLMs, VLMs, and training generative architectures including diffusion, reinforcement learning, flow matching, or normalizing flows at scale. You will utilize Python and PyTorch to develop systems for multimodal perception, speech understanding, and audio-to-audio modeling. This position addresses the technical challenge of creating human-centric solutions through advanced machine learning and computer vision research.

What you'll do

  • Develop foundation models for generative AI and multimodal systems integrating video, audio, and text.
  • Research and develop algorithms for interactive models and audio-to-audio modeling systems.
  • Solve open research problems in multimodal generative AI and intelligent agents.
  • Define data requirements, validation strategies, and key performance indicators for product features.
  • Train and tune large-scale foundation models using architectures like diffusion, reinforcement learning, or flow matching.
  • Develop real-time or streaming multimodal models for integration into consumer products.
  • Implement research findings through practical software engineering using Python and PyTorch.

What we're looking for

  • Bachelor's degree and at least 3 years of relevant industry experience.
  • Experience building models for multimodal perception systems.
  • Experience working with Large Language Models (LLMs) and Vision Language Models (VLMs).
  • Proficiency in Python, PyTorch, and software engineering skills.
  • MS or PhD in computer vision, graphics, machine learning, computer science, or related fields (preferred).
  • Experience developing, training, or tuning foundation models and multimodal LLMs (preferred).
  • Experience with generative architectures like diffusion, reinforcement learning, flow matching, or normalizing flow at scale (preferred).
  • Experience with real-time/streaming multimodal models, speech understanding, or applying RL to post-train foundation models (preferred).

More like this

Similar roles

Applied AI Scientist, Multimodal Intelligence

Apple Inc

Sunnyvale, CA 45 days ago $216,200$324,800
Generative AI Multimodal Foundation Models Deep Learning Python PyTorch JAX Computer Vision Data Pipelines On-device Machine Learning Vision-Language Models Video-Language Models
6+ yrs exp

Applied AI Scientist, Multimodal Intelligence

Apple Inc

Seattle, WA 45 days ago $205,400$308,500
Generative AI Multimodal Foundation Models Deep Learning Python PyTorch JAX Computer Vision Data Pipelines On-device Machine Learning Data Curation
6+ yrs exp

Applied AI Scientist, Multimodal Intelligence

Apple Inc

Sunnyvale, CA 49 days ago $150,400$277,600
Python PyTorch JAX Deep Learning Computer Vision Generative AI Multimodal Foundation Models Data Pipelines On-device Machine Learning Machine Learning