Engineering Manager, LLM Inference & Deployment at Scale
Nvidia
Quick summary
Market check
How this pay compares to similar roles
This role pays less than 64% of similar roles. Most pay $171,500–$254,562 — the shaded band above. At the midpoint, this role pays about $198k versus about $213k for comparable roles.
Based on 240 similar postings.
Employer
Qualcomm is a leading American semiconductor and telecommunications company based in San Diego, CA.
Qualcomm currently has 623 open roles on FindRole.
Listed pay typically runs $148,300–$222,500 across 603 roles with salary data.
Most-posted roles
At a glance
LLM Serving Engineer (Cloud AI Engineering), Senior / Staff Engineer joins the Cloud AI team to develop hardware and software solutions for inference acceleration. This role involves building a scalable LLM inference platform using techniques such as disaggregated serving, KV-Cache management, advanced parallelism, speculative algorithms, and specialized kernels. The engineer will contribute to serving packages like vLLM, SGLang, TGI, Triton-Inference server, Dynamo, and LLM-d while driving efficient serving through autoscaling, load balancing, and routing. Candidates must possess expertise in PyTorch, Python, and distributed systems, alongside a deep understanding of transformer-based architectures, MoEs, and attention mechanisms. The role requires proficiency in analyzing and optimizing deep learning workloads using tools like CUDA or Triton. This position addresses the technical challenge of optimizing large-scale inference for generative AI models within a complex cloud infrastructure environment.
Skills
What you'll do
What we're looking for
More like this
Nvidia
Qualcomm
Qualcomm
Qualcomm
Capital One Financial
GEICO