Senior GPU Inference Performance Engineer

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$164,000–$246,000 / yr
Posted
63 days ago
Freshness
Confirmed live yesterday
Closes
Jul 9, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $215k
This role $205k
$152k most similar roles pay here $276k

This role pays less than 56% of similar roles. Most pay $188,562–$241,750 — the shaded band above. At the midpoint, this role pays about $205k versus about $215k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Senior GPU Inference Performance Engineer

Senior GPU Inference Performance Engineer will own end-to-end performance analysis of GPU-accelerated AI inference workloads. Working at the intersection of hardware, systems software, networking, and AI infrastructure, this role involves profiling, diagnosing, and explaining performance across the full stack from silicon to software runtimes. The engineer will identify bottlenecks in HBM bandwidth, kernel scheduling, and memory allocation while conducting competitive head-to-head benchmarks against other accelerator vendors. Key responsibilities include analyzing multi-server inference networking using RDMA/RoCE, NCCL/RCCL, and GPUDirect RDMA, as well as profiling Kubernetes and container runtime overheads. The role requires proficiency in Python, C/C++, and GPU toolchains like ROCm, Nsight Systems, and CUDA. Candidates must analyze complex distributed systems to solve performance issues related to large language model workloads across multi-node clusters and high-performance computing environments.

What you'll do

  • Perform full-stack profiling of GPU inference workloads across both AMD and NVIDIA hardware platforms.
  • Identify performance bottlenecks in HBM bandwidth, compute utilization, kernel scheduling, and memory allocation.
  • Analyze system-level interactions between GPU runtimes, Linux operating systems, device drivers, and container environments.
  • Execute head-to-head competitive benchmarks against other vendors to provide data-backed evidence of performance differences.
  • Profile and optimize multi-node inference topologies including pipeline parallelism and tensor parallelism across distributed clusters.
  • Analyze network-level bottlenecks using RDMA/RoCE traces and communication libraries like NCCL or RCCL.
  • Evaluate the impact of Kubernetes scheduling, GPU operators, and container runtimes on inference latency.
  • Develop automated benchmarking harnesses, profiling scripts, and performance regression dashboards for continuous validation.

What we're looking for

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field preferred; advanced degree desired.
  • Proficiency with AMD (ROCm, rocProfiler, Omniperf) or NVIDIA (CUDA, Nsight Systems/Compute) profiling toolchains and GPU architecture.
  • Experience analyzing GPU communication and networking performance including NCCL/RCCL, RDMA/RoCE, GPUDirect RDMA, UCX, MPI, ConnectX, or Pensando.
  • Experience with multi-GPU and multi-node inference, training, or HPC environments involving tensor parallelism and pipeline parallelism.
  • Experience with Linux systems performance analysis, device drivers, virtualization, container runtimes, or low-level systems software development.
  • Strong Python and C/C++ skills with the ability to read GPU kernel code (HIP/CUDA) and runtime code.
  • Ability to produce data-backed reports and presentations explaining performance differences based on hardware and software evidence.
  • Experience with Kubernetes GPU scheduling, MIG, GPU operator performance, or related open-source infrastructure projects.

More like this

Similar roles

Senior System Software Engineer, GPU Performance

Nvidia

Remote 6 days ago $152,000$241,500
HPC C++ Python CUDA NCCL UCX NVSHMEM MPI Infiniband Ethernet RDMA PyTorch TensorFlow Kubernetes Docker SLURM Ansible Performance Engineering Parallel Programming
Remote

Principal Developer, AI Networking

Nvidia

Remote (Santa Clara, CA) +3 91 days ago $272,000$431,250
NCCL RDMA CUDA PyTorch TensorFlow C++ Python Bash RoCE MPI SHARP Distributed Systems LLM Training High-Performance Networking Profiling Benchmarking
10+ yrs exp Remote

Senior Staff Engineer, GPU Architect, Machine Learning

Samsung Electronics

Remote (San Jose, CA) 43 days ago $198,200$297,200
GPU Architecture Machine Learning C++ Python TensorFlow PyTorch LLMs LVMs Computer Vision Ray Tracing Graphics Pipeline RTL Workload Analysis Performance Optimization PPA Hardware-Software Integration
10+ yrs exp Remote