Principal Senior GPU SW Performance Engineer, Post-Training

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Jose, CA
Salary
$240,000–$360,000 / yr
Posted
170 days ago
Freshness
Confirmed live yesterday
Closes
Mar 25, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $216k
This role $300k
$145k most similar roles pay here $383k

This role pays more than 92% of similar roles. Most pay $187,700–$244,000 — the shaded band above. At the midpoint, this role pays about $300k versus about $216k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Principal Senior GPU SW Performance Engineer, Post-Training

As a Principal / Senior GPU SW Performance Engineer — Post‑Training, you will join the team to drive performance for post-training workloads on AMD Instinct GPUs. You will be responsible for delivering fast, stable, and reproducible training pipelines on ROCm by addressing cross-stack issues involving data loaders, kernels, distributed training, and compilers. Your daily work involves optimizing throughput, memory efficiency, and stability across data, model, and optimizer steps while improving multi-GPU/multi-node communication patterns. You will develop efficient kernels, perform graph-level optimizations, and resolve bottlenecks using standard profiling tools to prevent regressions in CI. The role requires expertise in PyTorch, Python, C++, Triton, and HIP for deep learning tasks like SFT, LoRA, and RL-based training at scale. You will collaborate with framework, compiler, and model teams to solve complex performance challenges within the AI domain.

What you'll do

  • Drive performance for fine-tuning and RL training solutions on AMD GPUs.
  • Improve throughput, memory efficiency, and stability across data, model, and optimizer steps.
  • Optimize multi-GPU and multi-node training and communication patterns.
  • Develop efficient kernels, operations, and targeted graph-level optimizations.
  • Profile, diagnose, and resolve performance bottlenecks using standard tooling.
  • Prevent regressions in CI and ensure stable training pipelines on ROCm.
  • Deliver reproducible pipelines and documentation for internal and external developers.
  • Coordinate with framework, compiler, and model teams to implement durable improvements.

What we're looking for

  • Bachelor's, Master's, or Doctorate degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • Proven experience in GPU performance engineering for deep learning using ROCm, HIP, Triton, or similar technologies.
  • Hands-on experience with SFT, LoRA, and RL-based training at scale.
  • Strong proficiency in PyTorch, including torch.distributed, FSDP, ZeRO, or equivalent frameworks.
  • Proficiency in Python and C++ with the ability to read and write kernels.
  • Experience with distributed systems and collective communication libraries.
  • Ability to profile, diagnose, and resolve performance bottlenecks while ensuring reproducibility.
  • Track record of turning profiles into fixes, upstreaming changes, and documenting results.

More like this

Similar roles

Fellow GPU Performance Optimization Engineer

Amd

San Jose, CA 167 days ago $268,000$402,000
GPU Distributed Training PyTorch JAX TensorFlow CUDA HIP ROCm Python C++ RDMA NCCL RCCL Megatron-LM Torchtitan MaxText Kernel Optimization Compiler Stack Performance Profiling
Hybrid

Software Engineer, GPU AI ML

Amd

Santa Clara, CA 77 days ago $204,000$306,000
C++ HIP CUDA ROCm PyTorch TensorFlow JAX GPU Architecture Kernel Optimization Distributed Systems LLMs SFT RLHF GRPO Quantization Verilog SystemVerilog RTL Design
Hybrid

Senior AI/ML Platform Engineer

Amd

Santa Clara, CA 51 days ago $204,000$306,000
Python C++ Go Rust Kubernetes Ray Slurm vLLM SGLang Triton PyTorch JAX ROCm HIP CUDA MLflow Weights & Biases CI/CD Distributed Systems GPU Infrastructure
Hybrid