Senior Inference Engineer, GPU Kernel Optimization
Nvidia
Quick summary
Market check
How this pay compares to similar roles
This role pays less than 56% of similar roles. Most pay $188,562–$241,750 — the shaded band above. At the midpoint, this role pays about $205k versus about $215k for comparable roles.
Based on 240 similar postings.
Employer
AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors
Amd currently has 367 open roles on FindRole.
Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.
Most-posted roles
At a glance
Senior GPU Inference Performance Engineer will own end-to-end performance analysis of GPU-accelerated AI inference workloads. Working at the intersection of hardware, systems software, networking, and AI infrastructure, this role involves profiling, diagnosing, and explaining performance across the full stack from silicon to software runtimes. The engineer will identify bottlenecks in HBM bandwidth, kernel scheduling, and memory allocation while conducting competitive head-to-head benchmarks against other accelerator vendors. Key responsibilities include analyzing multi-server inference networking using RDMA/RoCE, NCCL/RCCL, and GPUDirect RDMA, as well as profiling Kubernetes and container runtime overheads. The role requires proficiency in Python, C/C++, and GPU toolchains like ROCm, Nsight Systems, and CUDA. Candidates must analyze complex distributed systems to solve performance issues related to large language model workloads across multi-node clusters and high-performance computing environments.
What you'll do
What we're looking for
More like this
Nvidia
Nvidia
Nvidia
Nvidia
Samsung Electronics
Samsung Electronics