Senior Staff AI Performance Engineer
Qualcomm
Quick summary
Market check
How this pay compares to similar roles
This role pays more than 52% of similar roles. Most pay $162,000–$235,750 — the shaded band above. At the midpoint, this role pays about $198k versus about $199k for comparable roles.
Based on 240 similar postings.
Employer
Qualcomm is a leading American semiconductor and telecommunications company based in San Diego, CA.
Qualcomm currently has 623 open roles on FindRole.
Listed pay typically runs $148,300–$222,500 across 603 roles with salary data.
Most-posted roles
At a glance
As an AI Model Optimization Architect on the Machine Learning Engineering team, you will lead end-to-end model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models on inference accelerators. You will be responsible for developing optimization strategies to transform PyTorch models into efficient executions while balancing throughput, latency, memory, and quality. Your daily work involves driving graph capture using PyTorch, ONNX, and torch.compile, designing fusion kernels via Triton, and managing KVcache and continuous batching systems. You will collaborate with compiler and performance teams to solve complex technical challenges regarding distributed inference and hardware-specific scaling. The role requires expert proficiency in PyTorch, Python, and transformer architectures, alongside a deep understanding of computer architecture, ML accelerators, and distributed systems to ensure production-ready solutions for large-scale foundation models across various hardware platforms.
Skills
What you'll do
What we're looking for
More like this
Qualcomm
Qualcomm
Qualcomm
Qualcomm
Qualcomm
Qualcomm