Applied Researcher: On-Device Multimodal Reasoning

Apple Inc

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Sunnyvale, CA
Salary
$150,400–$277,600 / yr
Posted
36 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $226k
This role $214k
$135k most similar roles pay here $293k

This role pays less than 54% of similar roles. Most pay $196,750–$256,162 — the shaded band above. At the midpoint, this role pays about $214k versus about $226k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Applied Researcher: On-Device Multimodal Reasoning

Applied Researcher: On-Device Multimodal Reasoning joins the Video Computer Vision organization to develop real-time, on-device multimodal AI systems. You will design and train compact vision-language models under 10B parameters that perform multi-step reasoning over images, video, and 3D scenes while operating under strict memory and latency constraints. Your daily work involves developing inference paths including speculative decoding, KV-cache compression, and quantization, alongside distilling reasoning from frontier teachers using methods like GRPO or STaR. You will also build systems where small models invoke specialized perception experts for geometry estimation and pose recovery. The role requires expertise in Python, PyTorch, and optimization toolchains like MLX or CoreML. This position solves the technical challenge of enabling sophisticated, multi-step reasoning on mobile hardware without relying on cloud connectivity by optimizing for fixed power and performance budgets.

What you'll do

  • Design and train compact vision-language models under 10B parameters for multi-step reasoning on mobile devices.
  • Implement efficient reasoning techniques such as chain-of-thought distillation, adaptive compute allocation, and test-time scaling.
  • Optimize the decoding stack using speculative decoding, KV-cache compression, and quantization for low-latency inference.
  • Integrate real-time perception experts like 3D scene reconstruction and object tracking into the multimodal reasoning pipeline.
  • Develop structured-output fusion to represent geometry and body parameters within limited token budgets.
  • Apply post-training methods including supervised distillation from frontier models and reinforcement learning (GRPO, RLVR).
  • Drive visual token efficiency through pruning, merging, and adaptive resolution policies.
  • Optimize model architectures specifically for Apple silicon hardware to balance performance with power and memory constraints.

What we're looking for

  • Must have a Master's degree in Computer Science, Machine Learning, AI, Computer Vision, or a related field (or equivalent experience).
  • Must possess a strong foundation in deep learning with specific experience in LLM or VLM training and inference optimization.
  • Must have demonstrated experience working with multimodal models in resource-constrained environments including model compression and reasoning quality.
  • Must be proficient in Python and modern deep learning frameworks like PyTorch.
  • Must be familiar with inference and optimization toolchains such as quantization, distillation, pruning, vLLM, SGLang, llama.cpp, MLX, or CoreML.
  • Preferred: PhD with research in efficient multimodal reasoning, model compression, or lightweight VLM architectures.
  • Preferred: Experience training/post-training vision-language models end-to-end including connector design and visual instruction tuning.
  • Preferred: Experience deploying LLM or multimodal models on mobile or edge hardware while managing ANE/GPU kernel and memory constraints.

More like this

Similar roles

Multimodal AI Researcher

Apple Inc

Sunnyvale, CA 30 days ago $150,400$277,600
Multimodal AI Generative AI LLMs VLMs Python PyTorch Computer Vision Machine Learning Reinforcement Learning Flow Matching Software Engineering Foundation Models
3+ yrs exp

Applied AI Scientist, Multimodal Intelligence

Apple Inc

Sunnyvale, CA 49 days ago $150,400$277,600
Python PyTorch JAX Deep Learning Computer Vision Generative AI Multimodal Foundation Models Data Pipelines On-device Machine Learning Machine Learning

Applied AI Scientist, Multimodal Intelligence

Apple Inc

Sunnyvale, CA 45 days ago $216,200$324,800
Generative AI Multimodal Foundation Models Deep Learning Python PyTorch JAX Computer Vision Data Pipelines On-device Machine Learning Vision-Language Models Video-Language Models
6+ yrs exp