Principal GenAI Inference Optimization Engineer

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Jose, CA
Salary
$240,000–$360,000 / yr
Posted
167 days ago
Freshness
Confirmed live yesterday
Closes
Mar 27, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $220k
This role $300k
$153k most similar roles pay here $382k

This role pays more than 91% of similar roles. Most pay $185,900–$254,750 — the shaded band above. At the midpoint, this role pays about $300k versus about $220k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Principal GenAI Inference Optimization Engineer

The Principal GenAI Inference Optimization Engineer joins the Models and Applications team to enhance the performance, efficiency, and scalability of generative AI inference workloads on AMD GPU platforms. This role involves optimizing latency, throughput, and cost efficiency for large-scale LLM and multimodal model serving across single-node and distributed environments. The engineer will address bottlenecks in compute, memory, and communication by implementing techniques like quantization, batching strategies, prefix caching, and speculative decoding. Key responsibilities include developing profiling tools and collaborating with hardware and compiler teams to improve the software-hardware stack. Required skills include proficiency in Python and systems languages like C++, CUDA, or HIP, alongside experience with frameworks such as PyTorch, JAX, TensorFlow, vLLM, SGLang, Triton, and TensorRT-LLM. The role focuses on solving technical challenges related to GPU architecture, memory bandwidth, and distributed serving systems.

What you'll do

  • Optimize performance of GenAI inference workloads on AMD GPU platforms in single-node and distributed environments.
  • Improve latency, throughput, and cost efficiency for LLM and multimodal model serving in production.
  • Identify and resolve bottlenecks across compute, memory, and communication layers including kernel efficiency and KV-cache usage.
  • Implement optimization techniques such as batching strategies, quantization, prefix caching, and speculative decoding.
  • Develop and optimize scalable serving systems including request scheduling and resource utilization.
  • Create and utilize profiling, benchmarking, and performance analysis tools for inference workloads.
  • Contribute to cross-stack optimizations across kernels, runtimes, communication libraries, and inference frameworks.
  • Document best practices and establish performance guidelines for GenAI deployment.

What we're looking for

  • Proficiency in Python and at least one systems language such as C++, CUDA, or HIP.
  • Experience with GenAI inference optimization techniques including quantization, KV-cache optimization, batching, and speculative decoding.
  • Hands-on experience with inference/serving frameworks like vLLM, SGLang, Triton, or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchy, and interconnects such as PCIe, Infinity Fabric, or RDMA.
  • Experience working with ML frameworks including PyTorch, JAX, or TensorFlow for inference workloads.
  • Experience with profiling, debugging, and performance tuning tools to analyze compute, memory, and communication bottlenecks.
  • Familiarity with distributed systems and serving architectures for large-scale model deployment.
  • A B.S., M.S., or Ph.D. in Computer Science, Computer Engineering, or a related field is preferred.

More like this

Similar roles

AI Inference Engineer

F5 Inc

San Jose, CA +1 15 days ago $176,600$265,000
vLLM TGI NVIDIA Triton Python C++ Rust TensorRT Llama.cpp Ollama Kubernetes Docker AWS GCP Azure CUDA TPUs MLOps
Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid