AI Model Optimization Architect

Qualcomm

Confirmed live yesterday High trust
Closes tomorrow

Quick summary

Work type
On-site
Location
San Diego, CA
Salary
$158,400–$237,600 / yr
Posted
179 days ago
Freshness
Confirmed live yesterday
Closes
Sep 12, 2026 (soon)

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $199k
This role $198k
$139k most similar roles pay here $252k

This role pays more than 52% of similar roles. Most pay $162,000–$235,750 — the shaded band above. At the midpoint, this role pays about $198k versus about $199k for comparable roles.

Based on 240 similar postings.

Employer

About Qualcomm

Qualcomm is a leading American semiconductor and telecommunications company based in San Diego, CA.

Qualcomm currently has 623 open roles on FindRole.

Listed pay typically runs $148,300–$222,500 across 603 roles with salary data.

Most-posted roles

View all roles at Qualcomm

At a glance

TL;DR · AI Model Optimization Architect

As an AI Model Optimization Architect on the Machine Learning Engineering team, you will lead end-to-end model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models on inference accelerators. You will be responsible for developing optimization strategies to transform PyTorch models into efficient executions while balancing throughput, latency, memory, and quality. Your daily work involves driving graph capture using PyTorch, ONNX, and torch.compile, designing fusion kernels via Triton, and managing KVcache and continuous batching systems. You will collaborate with compiler and performance teams to solve complex technical challenges regarding distributed inference and hardware-specific scaling. The role requires expert proficiency in PyTorch, Python, and transformer architectures, alongside a deep understanding of computer architecture, ML accelerators, and distributed systems to ensure production-ready solutions for large-scale foundation models across various hardware platforms.

What you'll do

  • Architect and deliver model optimization strategies to transform PyTorch models for efficient inference on Qualcomm accelerators.
  • Drive graph capture and deployment using PyTorch, ONNX, and torch.compile through model rewrites and transformations.
  • Design and implement fusion kernels using DSL-based approaches like Triton for performance-critical algorithmic rewrites.
  • Profile and optimize LLM, VLM, and diffusion inference for throughput and latency across various serving modes.
  • Manage transformer-specific optimizations including KVcache management, decoding behavior, and long context performance.
  • Enable and optimize continuous batching systems to manage memory, scheduling, and tail latency.
  • Architect and scale distributed inference strategies such as sharding and parallelism across multi-core and multi-device systems.
  • Establish reusable optimization patterns and tooling to scale model optimizations across new hardware architectures.

What we're looking for

  • Expert level expertise in PyTorch and inference focused model optimization with strong Python engineering skills.
  • Hands-on experience with torch.compile, TorchDynamo, or related graph capture and compilation workflows.
  • Deep understanding of transformer architectures, attention mechanisms, MoEs, and performance trade-offs.
  • Practical experience with KVcache behavior, serving time optimizations, and memory/performance tradeoffs.
  • Strong foundation in computer architecture, ML accelerators, and distributed systems.
  • Proven ability to lead cross-functional technical efforts and influence design decisions.
  • Bachelor's degree and 4+ years of relevant engineering experience, or a Master's degree and 3+ years of experience, or a PhD and 2+ years of experience.

More like this

Similar roles

Senior Staff AI Performance Engineer

Qualcomm

San Diego, CA +1 155 days ago $178,400$267,600
PyTorch ONNX Python Triton LLM VLM Diffusion Models Transformer Architectures Machine Learning Compilers torch.compile torchDynamo Distributed Systems Computer Architecture Inference Acceleration Linear Algebra
6+ yrs exp

Senior Staff LLM Serving Engineer, Cloud AI Engineering

Qualcomm

San Diego, CA +1 155 days ago $158,400$237,600
LLM PyTorch Python Triton-Inference Server vLLM SGLang CUDA Triton torch.compile torchDynamo Distributed Systems Kernel Design Deep Learning KV-Cache Management Model Optimization
4+ yrs exp

Senior Engineer, Machine Learning

Qualcomm

San Diego, CA 72 days ago $140,800$211,200
Python C++ Rust Go PyTorch TensorFlow LLM Transformer vLLM ONNX Kubernetes Docker CI/CD RAG Vector Databases OpenSearch Qdrant Quantization Distributed Computing Microservices
2+ yrs exp

Senior Engineer, Machine Learning

Qualcomm

San Diego, CA 72 days ago $140,800$211,200
Python C++ Rust Go PyTorch TensorFlow LLM Transformer vLLM ONNX Kubernetes Docker CI/CD RAG Vector Databases OpenSearch Qdrant Quantization Distributed Computing Microservices
2+ yrs exp

Senior AI Engineer

Qualcomm

Santa Clara, CA 72 days ago $129,300$193,900
Machine Learning Deep Learning Python PyTorch Hugging Face Transformers onnxruntime scikit-learn NumPy LLMs Vision Transformers MLOps CI/CD Containerization C++ Java Distributed Training
2+ yrs exp

AI Researcher, On-Device LLM Efficiency

Qualcomm

San Diego, CA 66 days ago $138,800$208,200
LLM PyTorch Python Deep Learning Transformers Inference Acceleration Model Architecture KV Cache Compression Machine Learning Research
4+ yrs exp