Senior AI Researcher, On-Device LLM Efficiency

Qualcomm

Confirmed live today High trust

Quick summary

Work type
On-site
Location
San Diego, CA
Salary
$159,100–$238,700 / yr
Posted
11 days ago
Freshness
Confirmed live today
Closes
Mar 29, 2027

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $226k
This role $199k
$145k $292k
below market most similar roles pay here above market

This role pays less than 73% of similar roles. Most pay $196,750–$254,750 — the blue band above. At the midpoint, this role pays about $199k versus about $226k for comparable roles.

Based on 240 similar postings.

Employer

About Qualcomm

Qualcomm is a leading American semiconductor and telecommunications company based in San Diego, CA.

Qualcomm currently has 558 open roles on FindRole.

Listed pay typically runs $152,130–$222,500 across 534 roles with salary data.

Most-posted roles

View all roles at Qualcomm

At a glance

TL;DR · Senior AI Researcher, On-Device LLM Efficiency

The Senior AI Researcher, On-Device LLM Efficiency joins the Machine Learning Researcher team to conduct fundamental research creating innovative machine learning methodology. The role focuses on research and development regarding LLM inference efficiency algorithms, efficient model architecture design, and LLM training. The successful candidate will develop creative solutions that address practical challenges on devices, implementing and evaluating these solutions in both simulation and on-device environments. Key technical requirements include a strong background in deep learning and Transformers, along with proficiency in Python and PyTorch. The work specifically addresses the technical problem of LLM reasoning and inference acceleration, including efficient attention, KV cache compression, and on-device AI deployment for mobile, edge, and IOT products.

What you'll do

  • Research and develop LLM inference efficiency algorithms and efficient model architectures.
  • Conduct research on LLM training methodologies to achieve beyond state-of-the-art performance.
  • Develop creative solutions that address practical challenges of running models on devices.
  • Implement and evaluate potential solutions in both simulation and on-device environments.
  • Optimize machine learning models, systems, platforms, and methods for mobile and edge products.
  • Research inference acceleration techniques such as efficient attention and KV cache compression.

What we're looking for

  • Master's degree in Computer Engineering, Computer Science, Electrical Engineering, or related field and 2+ years of engineering work experience.
  • PhD in Computer Engineering, Computer Science, Electrical Engineering, or related field.
  • 6+ months of academic and/or work experience developing and/or optimizing machine learning models, systems, platforms, or methods.
  • 4+ years of AI research experience.
  • Strong background in deep learning and Transformers.
  • Strong programming skills in Python and PyTorch.
  • Experience in LLM reasoning or inference acceleration research.
  • PhD in Computer Science, Electrical Engineering, or related field (preferred).
  • Experience in LLM efficiency research such as efficient attention, inference acceleration, or KV cache compression (preferred).
  • Experience in on-device AI deployment on mobile or edge devices (preferred).
  • Publishing research papers at top-tier AI/ML conferences as a lead author (preferred).

More like this

Similar roles

Senior Staff LLM Serving Engineer

Qualcomm

San Diego, CA +1 24 days ago $162,000–$243,000
LLM Serving PyTorch Python Triton-Inference Server vLLM SGLang CUDA Triton torch.compile torchDynamo Distributed Systems Deep Learning Model Optimization KV-Cache Management Inference Acceleration KServe Ollama LMCache MoonCake
4+ yrs exp

Senior Staff AI Performance Engineer

Qualcomm

San Diego, CA +1 24 days ago $194,400–$291,600
PyTorch ONNX Python Triton LLM VLM Diffusion Models Transformer Architectures Distributed Systems torch.compile torchDynamo Computer Architecture Inference Acceleration Linear Algebra
6+ yrs exp

Senior Applied AI ML Engineer

JPMorgan Chase

Jersey City, NJ +1 39 days ago
Python Java PyTorch TensorFlow Huggingface Transformers NeMo Databricks AWS SQL NoSQL MLOps NLP LLM s Microservices Distributed Systems Vector Databases Information Retrieval
8+ yrs exp

Applied Researcher I

Capital One Financial

San Jose, CA +3 30 days ago
Pytorch AWS Huggingface Lightning VectorDBs LLM RLHF NLP Deep Learning Supervised Finetuning Instruction-Tuning Transfer Learning Tokenization
2+ yrs exp

Senior Engineering Program Manager, On-Device ML

Apple Inc

Cupertino, CA 25 days ago $144,600–$263,800
Machine Learning ML Training Data Processing Pipelines Data Analytics Backend Services Research to Production Technical Roadmaps Apple Silicon Software Platforms Metrics Data Processing
6+ yrs exp

Senior ML Engineer

Salesforce

Bellevue, WA +2 41 days ago
Python MLOps Apache Kafka Flink Ray Spark Pyspark Docker Kubernetes Apache Airflow CI/CD Graph Analytics Supervised Learning Anomaly Detection Clustering Feature Stores MITRE ATT&CK OCSF
3+ yrs exp Hybrid