Senior Staff AI Performance Engineer

Qualcomm

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
San Diego, CAMarkham, Ontario, Canada
Salary
$178,400–$267,600 / yr
Posted
155 days ago
Freshness
Confirmed live yesterday
Closes
Oct 6, 2026

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $200k
This role $223k
$136k most similar roles pay here $282k

This role pays more than 65% of similar roles. Most pay $159,250–$241,750 — the shaded band above. At the midpoint, this role pays about $223k versus about $200k for comparable roles.

Based on 240 similar postings.

Employer

About Qualcomm

Qualcomm is a leading American semiconductor and telecommunications company based in San Diego, CA.

Qualcomm currently has 623 open roles on FindRole.

Listed pay typically runs $148,300–$222,500 across 603 roles with salary data.

Most-posted roles

View all roles at Qualcomm

At a glance

TL;DR · Senior Staff AI Performance Engineer

AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff joins the Cloud AI team to develop hardware and software solutions for inference acceleration. This role involves converting, optimizing, and deploying models using PyTorch and ONNX while addressing performance analysis for LLM, VLM, and diffusion models to meet throughput and latency constraints. The engineer will design high-level kernels in Triton, map next-generation workloads onto hardware designs, and collaborate with internal compiler, firmware, and platform teams to resolve complex stability issues. Key technical requirements include proficiency in Python, deep knowledge of transformer architectures, attention mechanisms, and MoEs, alongside experience with sharding and parallelisms. The role focuses on the technical challenge of optimizing inference for advanced algorithms and ensuring efficient execution across various hardware designs while managing the full product lifecycle from research to deployment.

What you'll do

  • Convert, optimize, and deploy AI models for efficient inference using PyTorch and ONNX.
  • Analyze and optimize LLM, VLM, and diffusion models to meet throughput and latency constraints.
  • Map next-generation AI workloads onto current and future hardware designs.
  • Design and implement high-level kernels in Triton to generate efficient low-level code.
  • Identify optimization opportunities by analyzing advanced GenAI algorithms like attention mechanisms and MoEs.
  • Investigate and resolve complex performance or stability issues to identify root causes.
  • Create engineering solutions to provide continuous insights into the performance of AI workloads.

What we're looking for

  • Bachelor's degree with 6+ years of experience in Hardware, Software, or Systems Engineering.
  • Master's degree with 5+ years of experience in Hardware, Software, or Systems Engineering.
  • PhD with 4+ years of experience in Hardware, Software, or Systems Engineering.
  • Hands-on experience building and optimizing language models using PyTorch and ONNX.
  • Deep understanding of transformer architectures, attention mechanisms, and performance trade-offs.
  • Proficiency in Python programming and designing high-level kernels in Triton.
  • Knowledge of computer architecture, ML accelerators, in-memory processing, and distributed systems.
  • Experience with workload mapping strategies involving sharding or various parallelisms.

More like this

Similar roles

AI Model Optimization Architect

Qualcomm

San Diego, CA 179 days ago $158,400$237,600
PyTorch Python Triton ONNX torch.compile TorchDynamo LLM VLM Transformer Distributed Systems Kernel Fusion Continuous Batching KVcache Machine Learning Computer Architecture
4+ yrs exp

Senior Staff LLM Serving Engineer, Cloud AI Engineering

Qualcomm

San Diego, CA +1 155 days ago $158,400$237,600
LLM PyTorch Python Triton-Inference Server vLLM SGLang CUDA Triton torch.compile torchDynamo Distributed Systems Kernel Design Deep Learning KV-Cache Management Model Optimization
4+ yrs exp

Senior AI/ML Performance Engineer

General Motors (GM)

Sunnyvale, CA +1 25 days ago $144,700$261,300
Python PyTorch Kubernetes Nvidia DCGM nvidia-smi Grafana AWS GCP Azure Hugging Face BigQuery Nvidia Nsight high-performance computing Distributed Systems ML Infrastructure GPU Architecture
5+ yrs exp Hybrid

AI Systems Performance Engineer

Broadcom

San Jose, CA 144 days ago $141,300$226,000
Ethernet MLPerf NCCL Python C++ Linux PyTorch RDMA RoCEv2 Docker Kubernetes CI/CD Performance Benchmarking Distributed Systems
10+ yrs exp

Senior AI Engineer, Platform Engineering

The Hartford

Hartford, CT +3 42 days ago $127,600$191,400
Python TypeScript LangChain LangGraph RAG GraphRAG GCP Vertex AI AlloyDB PostgreSQL Terraform CI/CD MCP Google ADK GitHub Spec-Kit OpenSpec BMAD-METHOD Cloud Run Vector Search
6+ yrs exp Hybrid