LLM Serving Engineer

Qualcomm

Confirmed live today High trust

Quick summary

Work type
On-site
Location
San Diego, CA
Salary
$122,800–$184,200 / yr
Posted
3 days ago
Freshness
Confirmed live today
Closes
Mar 10, 2027

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $204k
This role $154k
$107k most similar roles pay here $275k

This role pays less than 79% of similar roles. Most pay $159,750–$248,853 — the shaded band above. At the midpoint, this role pays about $154k versus about $204k for comparable roles.

Based on 240 similar postings.

Employer

About Qualcomm

Qualcomm is a leading American semiconductor and telecommunications company based in San Diego, CA.

Qualcomm currently has 639 open roles on FindRole.

Listed pay typically runs $146,400–$221,900 across 619 roles with salary data.

Most-posted roles

View all roles at Qualcomm

At a glance

TL;DR · LLM Serving Engineer

The LLM Serving Engineer joins the Machine Learning Engineering team to create and implement machine learning techniques, frameworks, and tools for various technology verticals. This role involves modeling, architecting, and developing machine learning hardware co-designed with software for inference or training solutions. The engineer will develop optimized software such as machine learning kernels, compiler tools, and model efficiency tools to enable AI models on specific hardware features. Day-to-day responsibilities include prototyping complex algorithms, conducting experiments to train and evaluate models, and extending runtime frameworks with new optimizations. Required skills include proficiency in Python, R, C, or C++, experience with frameworks like TensorFlow, Keras, or PyTorch, and knowledge of statistics and probability. The role focuses on solving technical challenges in mobile, edge, auto, and IOT products through hardware and software integration.

What you'll do

  • Apply machine learning knowledge to extend training or runtime frameworks with new features and optimizations.
  • Model, architect, and develop machine learning hardware co-designed with software for inference or training solutions.
  • Develop optimized software including kernels and compiler tools to enable AI models on specific hardware.
  • Integrate machine learning techniques into products and AI solutions to enhance customer offerings.
  • Prototype complex machine learning algorithms, models, and frameworks aligned with product roadmaps.
  • Conduct complex experiments to independently train and evaluate machine learning models and software.

What we're looking for

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 2+ years of relevant engineering experience.
  • Master's degree in Computer Science, Engineering, Information Systems, or related field and 1+ year of relevant engineering experience.
  • PhD in Computer Science, Engineering, Information Systems, or related field.
  • Master's degree in Computer Science, Engineering, Information Systems, or related field (preferred).
  • 2+ years of experience with Machine Learning frameworks such as TensorFlow, Caffe, PyTorch, or Keras (preferred).
  • 2+ years of experience in embedded system development and optimization for ML domains like NLP or multi-media (preferred).
  • 2+ years of experience with programming languages suitable for machine learning, such as Python, R, C, or C++ (preferred).
  • 2+ years of experience using statistics and probability, including conditional probability and Bayes rule (preferred).

More like this

Similar roles

Principal Engineer

Qualcomm

San Diego, CA 139 days ago $200,800$301,200
Computer Vision Deep Learning Python C++ PyTorch TensorFlow Keras Caffe Linux Android QNX Quantization Pruning Distillation CPU GPU DSP NPU Containerization
8+ yrs exp

Modem Systems Engineer

Qualcomm

San Diego, CA 5 days ago $104,000$156,000
Machine Learning C++ Python Matlab TensorFlow PyTorch Keras Signal Processing Communication Theory Neural Networks Convex Optimization Numerical Analysis Information Theory Channel Coding Systems Engineering

Senior Staff LLM Serving Engineer, Cloud AI Engineering

Qualcomm

San Diego, CA +1 158 days ago $158,400$237,600
LLM PyTorch Python Triton-Inference Server vLLM SGLang CUDA Triton torch.compile torchDynamo Distributed Systems Kernel Design Deep Learning KV-Cache Management Model Optimization
4+ yrs exp

Engineering Manager, LLM Inference & Deployment at Scale

Nvidia

Santa Clara, CA 17 days ago $224,000$356,500
LLMs VLMs TensorRT TensorRT-LLM vLLM SGLang Quantization Speculative Decoding Continuous Batching Prefix Caching KV-cache Optimization Distributed Computing GPU Cluster Orchestration Model Serving Inference Optimization Deep Learning
8+ yrs exp Hybrid