Senior Staff LLM Serving Engineer, Cloud AI Engineering

Qualcomm

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
San Diego, CAMarkham, ON, Canada
Salary
$158,400–$237,600 / yr
Posted
155 days ago
Freshness
Confirmed live yesterday
Closes
Oct 6, 2026

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $213k
This role $198k
$146k most similar roles pay here $274k

This role pays less than 64% of similar roles. Most pay $171,500–$254,562 — the shaded band above. At the midpoint, this role pays about $198k versus about $213k for comparable roles.

Based on 240 similar postings.

Employer

About Qualcomm

Qualcomm is a leading American semiconductor and telecommunications company based in San Diego, CA.

Qualcomm currently has 623 open roles on FindRole.

Listed pay typically runs $148,300–$222,500 across 603 roles with salary data.

Most-posted roles

View all roles at Qualcomm

At a glance

TL;DR · Senior Staff LLM Serving Engineer, Cloud AI Engineering

LLM Serving Engineer (Cloud AI Engineering), Senior / Staff Engineer joins the Cloud AI team to develop hardware and software solutions for inference acceleration. This role involves building a scalable LLM inference platform using techniques such as disaggregated serving, KV-Cache management, advanced parallelism, speculative algorithms, and specialized kernels. The engineer will contribute to serving packages like vLLM, SGLang, TGI, Triton-Inference server, Dynamo, and LLM-d while driving efficient serving through autoscaling, load balancing, and routing. Candidates must possess expertise in PyTorch, Python, and distributed systems, alongside a deep understanding of transformer-based architectures, MoEs, and attention mechanisms. The role requires proficiency in analyzing and optimizing deep learning workloads using tools like CUDA or Triton. This position addresses the technical challenge of optimizing large-scale inference for generative AI models within a complex cloud infrastructure environment.

What you'll do

  • Build a scalable LLM inference platform using techniques like disaggregated serving, KV-Cache management, and advanced parallelism.
  • Contribute to the development of LLM serving packages such as vLLM, SGLang, TGI, and Triton-Inference server.
  • Identify new optimization opportunities by analyzing advanced algorithms like attention mechanisms and MoEs.
  • Drive efficient serving through the implementation of smart autoscaling, load balancing, and routing.
  • Engage with open-source communities to evolve and improve inference frameworks.
  • Analyze, profile, and optimize deep learning workloads for production environments.
  • Develop high-performance kernels using PyTorch, CUDA, or Triton to accelerate model execution.

What we're looking for

  • Bachelor's degree with 4+ years of experience in engineering or a related field.
  • Master's degree with 3+ years of experience in engineering or a related field.
  • PhD with 2+ years of experience in engineering or a related field.
  • Hands-on experience with LLM serving or orchestration packages like vLLM, SGLang, TGI, or Triton-Inference Server.
  • Deep understanding of transformer-based architectures and foundational models including VLMs and SLMs.
  • Strong experience developing language models using PyTorch.
  • Proficiency in Python for large-scale projects and strong computer science fundamentals in algorithms and distributed programming.
  • Experience analyzing, profiling, and optimizing deep learning workloads and inference techniques.

More like this

Similar roles

Engineering Manager, LLM Inference & Deployment at Scale

Nvidia

Santa Clara, CA 14 days ago $224,000$356,500
LLMs VLMs TensorRT TensorRT-LLM vLLM SGLang Quantization Speculative Decoding Continuous Batching Prefix Caching KV-cache Optimization Distributed Computing GPU Cluster Orchestration Model Serving Inference Optimization Deep Learning
8+ yrs exp Hybrid

Senior Staff AI Performance Engineer

Qualcomm

San Diego, CA +1 155 days ago $178,400$267,600
PyTorch ONNX Python Triton LLM VLM Diffusion Models Transformer Architectures Machine Learning Compilers torch.compile torchDynamo Distributed Systems Computer Architecture Inference Acceleration Linear Algebra
6+ yrs exp

AI Model Optimization Architect

Qualcomm

San Diego, CA 179 days ago $158,400$237,600
PyTorch Python Triton ONNX torch.compile TorchDynamo LLM VLM Transformer Distributed Systems Kernel Fusion Continuous Batching KVcache Machine Learning Computer Architecture
4+ yrs exp

AI Researcher, On-Device LLM Efficiency

Qualcomm

San Diego, CA 66 days ago $138,800$208,200
LLM PyTorch Python Deep Learning Transformers Inference Acceleration Model Architecture KV Cache Compression Machine Learning Research
4+ yrs exp

Senior Lead AI Engineer, LLM Gateway, FM Hosting

Capital One Financial

San Jose, CA +3 8 days ago $229,900$262,400
LLM Python Go Scala Java C++ C# AWS Huggingface PyTorch VectorDBs Nemo Guardrails Machine Learning LLM Inference Similarity Search Model Evaluation Observability
6+ yrs exp

Senior Staff ML Engineer

GEICO

Remote (Palo Alto, CA) 45 days ago $150,000$300,000
Generative AI LLM Agentic Workflows Python Java RAG LangSmith LangGraph Kubernetes CI/CD AWS Azure Prompt Engineering
8+ yrs exp Remote