Fellow GPU Performance Optimization Engineer

Amd

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Jose, CA
Salary
$268,000–$402,000 / yr
Posted
167 days ago
Freshness
Confirmed live 2 days ago
Closes
Mar 27, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $213k
This role $335k
$136k most similar roles pay here $430k

This role pays more than 97% of similar roles. Most pay $182,125–$244,000 — the shaded band above. At the midpoint, this role pays about $335k versus about $213k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Fellow GPU Performance Optimization Engineer

The Fellow GPU Performance Optimization Engineer joins the Models and Applications team to maximize performance and efficiency of large-scale AI training workloads on AMD GPU platforms. This role involves driving innovations across the full software-hardware stack, focusing on distributed training at scale to improve system throughput, scalability, and utilization for generative AI workloads. The engineer will identify and eliminate bottlenecks in compute, memory, and communication while optimizing strategies like Data, Tensor, Pipeline Parallelism, and ZeRO. Key technical requirements include expertise in GPU architecture, interconnects such as PCIe/Infinity Fabric/RDMA, and performance profiling tools like ROCm. Candidates must be proficient in Python and systems languages including C++, CUDA, or HIP. The role involves optimizing ML frameworks like PyTorch, JAX, and TensorFlow while collaborating on compiler stacks, kernel optimizations, and communication libraries like NCCL or RCCL.

What you'll do

  • Optimize large-scale AI training workloads on AMD GPU platforms across single and multi-node environments.
  • Identify and eliminate system bottlenecks related to compute, memory bandwidth, and network utilization.
  • Optimize distributed training strategies including Data, Tensor, Pipeline Parallelism, and ZeRO for scalability.
  • Drive cross-stack optimizations across kernels, compilers, runtimes, communication libraries, and ML frameworks.
  • Develop and apply advanced profiling, benchmarking, and performance modeling methodologies.
  • Influence next-generation GPU architecture and software stack design through collaboration with hardware and compiler teams.
  • Lead open-source efforts to improve ecosystem performance on AMD platforms.
  • Define best practices and provide guidance for performance tuning of large-scale training workloads.

What we're looking for

  • A Ph.D. in Computer Science, Computer Engineering, or a related field is preferred, or equivalent experience with significant technical impact.
  • Expertise in GPU architecture and performance characteristics including compute units, memory hierarchy, and interconnects like PCIe/Infinity Fabric/RDMA.
  • Experience optimizing large-scale distributed training workloads across thousands of GPUs.
  • Proficiency in Python and at least one systems language such as C++, CUDA, or HIP.
  • Experience with machine learning frameworks including PyTorch, JAX, or TensorFlow with a focus on performance tuning.
  • Expertise in communication libraries and patterns such as NCCL/RCCL and collective operations.
  • Experience with distributed training frameworks like Megatron-LM, Torchtitan, MaxText, or equivalent.
  • Demonstrated technical leadership and the ability to influence cross-functional teams.

More like this

Similar roles

Software Engineer, GPU AI ML

Amd

Santa Clara, CA 77 days ago $204,000$306,000
C++ HIP CUDA ROCm PyTorch TensorFlow JAX GPU Architecture Kernel Optimization Distributed Systems LLMs SFT RLHF GRPO Quantization Verilog SystemVerilog RTL Design
Hybrid

Principal GenAI Inference Optimization Engineer

Amd

San Jose, CA 167 days ago $240,000$360,000
GenAI LLM GPU Architecture vLLM SGLang Triton TensorRT-LLM PyTorch JAX TensorFlow Python C++ CUDA HIP Quantization Distributed Systems RDMA Profiling Benchmarking
Hybrid