AI Researcher, On-Device LLM Efficiency

Qualcomm

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
San Diego, CA
Salary
$138,800–$208,200 / yr
Posted
66 days ago
Freshness
Confirmed live 2 days ago
Closes
Jan 3, 2027

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $229k
This role $174k
$122k most similar roles pay here $298k

This role pays less than 91% of similar roles. Most pay $202,800–$254,750 — the shaded band above. At the midpoint, this role pays about $174k versus about $229k for comparable roles.

Based on 240 similar postings.

Employer

About Qualcomm

Qualcomm is a leading American semiconductor and telecommunications company based in San Diego, CA.

Qualcomm currently has 623 open roles on FindRole.

Listed pay typically runs $148,300–$222,500 across 603 roles with salary data.

Most-posted roles

View all roles at Qualcomm

At a glance

TL;DR · AI Researcher, On-Device LLM Efficiency

As an AI Researcher, On-Device LLM Efficiency within the Machine Learning Research team, you will conduct fundamental research to develop innovative machine learning methodologies that achieve beyond state-of-the-art performance. Your daily responsibilities involve researching and developing LLM inference efficiency algorithms, designing efficient model architectures, and conducting LLM training while addressing practical challenges specific to on-device constraints. You will implement and evaluate these solutions in both simulation and on-device environments. The role requires expertise in deep learning, Transformers, and Python and PyTorch programming. Key technical focus areas include LLM reasoning, inference acceleration, efficient attention, and KV cache compression. This position addresses the critical challenge of optimizing large language model performance for deployment across mobile, edge, auto, and IOT products to enable next-generation experiences on localized hardware.

What you'll do

  • Conduct fundamental research to develop innovative machine learning methodologies exceeding state-of-the-art performance.
  • Research and develop LLM inference efficiency algorithms and efficient model architecture designs.
  • Develop creative solutions for machine learning that address practical challenges on mobile, edge, and IoT devices.
  • Implement and evaluate potential solutions in both simulation and on-device environments.
  • Conduct research specifically focused on LLM reasoning or inference acceleration.
  • Research techniques for efficient attention, inference acceleration, and KV cache compression.
  • Develop and optimize machine learning models, systems, platforms, and methods.

What we're looking for

  • Master's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • PhD in Computer Science, Electrical Engineering, or a related field is preferred.
  • 6+ months of academic and/or work experience developing or optimizing machine learning models, systems, platforms, or methods.
  • 4+ years of AI research experience.
  • Strong background in deep learning and Transformers.
  • Strong programming skills in Python and PyTorch.
  • Experience in LLM reasoning or inference acceleration research.
  • Preferred experience in on-device AI deployment and publishing papers at top-tier AI/ML conferences.

More like this

Similar roles

Senior AI Research Quantization Engineer

Qualcomm

San Diego, CA 149 days ago $140,800$211,200
Generative AI LLM LVM Multi-modal VLA Python PyTorch Quantization Model Compression Deep Learning Machine Learning Systems Engineering Hardware Engineering
2+ yrs exp

Senior Staff LLM Serving Engineer, Cloud AI Engineering

Qualcomm

San Diego, CA +1 155 days ago $158,400$237,600
LLM PyTorch Python Triton-Inference Server vLLM SGLang CUDA Triton torch.compile torchDynamo Distributed Systems Kernel Design Deep Learning KV-Cache Management Model Optimization
4+ yrs exp

Senior Staff AI Performance Engineer

Qualcomm

San Diego, CA +1 155 days ago $178,400$267,600
PyTorch ONNX Python Triton LLM VLM Diffusion Models Transformer Architectures Machine Learning Compilers torch.compile torchDynamo Distributed Systems Computer Architecture Inference Acceleration Linear Algebra
6+ yrs exp

AI Model Optimization Architect

Qualcomm

San Diego, CA 179 days ago $158,400$237,600
PyTorch Python Triton ONNX torch.compile TorchDynamo LLM VLM Transformer Distributed Systems Kernel Fusion Continuous Batching KVcache Machine Learning Computer Architecture
4+ yrs exp

Applied Researcher I

Capital One Financial

McLean, VA +3 162 days ago $218,700$249,600
PyTorch AWS Huggingface Lightning VectorDBs LLM RLHF Supervised Finetuning Instruction-Tuning Self-Supervised Learning NLP Deep Learning Cloud Computing
2+ yrs exp

Applied Researcher I

Capital One Financial

New York, NY +4 25 days ago $218,700$249,600
Pytorch AWS Huggingface Lightning VectorDBs LLM NLP Deep Learning RLHF Self-Supervised Learning Quantization Model Sparsification Training Optimization Compiler Design Instruction-Tuning
2+ yrs exp