Senior Inference Engineer, GPU Kernel Optimization
Nvidia
Quick summary
Market check
How this pay compares to similar roles
This role pays more than 91% of similar roles. Most pay $185,900–$254,750 — the shaded band above. At the midpoint, this role pays about $300k versus about $220k for comparable roles.
Based on 240 similar postings.
Employer
AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors
Amd currently has 367 open roles on FindRole.
Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.
Most-posted roles
At a glance
The Principal GenAI Inference Optimization Engineer joins the Models and Applications team to enhance the performance, efficiency, and scalability of generative AI inference workloads on AMD GPU platforms. This role involves optimizing latency, throughput, and cost efficiency for large-scale LLM and multimodal model serving across single-node and distributed environments. The engineer will address bottlenecks in compute, memory, and communication by implementing techniques like quantization, batching strategies, prefix caching, and speculative decoding. Key responsibilities include developing profiling tools and collaborating with hardware and compiler teams to improve the software-hardware stack. Required skills include proficiency in Python and systems languages like C++, CUDA, or HIP, alongside experience with frameworks such as PyTorch, JAX, TensorFlow, vLLM, SGLang, Triton, and TensorRT-LLM. The role focuses on solving technical challenges related to GPU architecture, memory bandwidth, and distributed serving systems.
Skills
What you'll do
What we're looking for
More like this
Nvidia
F5 Inc
Amd
Nvidia
Nvidia
Nvidia