Principal Senior GPU SW Performance Engineer, Post-Training

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Jose, CA
Salary
$204,000–$306,000 / yr
Posted
133 days ago
Freshness
Confirmed live yesterday
Closes
May 1, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $216k
This role $255k
$151k most similar roles pay here $323k

This role pays more than 84% of similar roles. Most pay $187,700–$244,000 — the shaded band above. At the midpoint, this role pays about $255k versus about $216k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Principal Senior GPU SW Performance Engineer, Post-Training

As a Principal / Senior GPU SW Performance Engineer — Post‑Training, you will drive the performance of post-training workloads on AMD Instinct GPUs. You will be responsible for delivering fast, stable, and reproducible fine-tuning and RL training pipelines on ROCm by optimizing kernels, distributed training, compiler integrations, and framework workflows. Your daily work involves improving throughput and memory efficiency across data, model, optimizer, and runtime steps while resolving complex cross-stack issues involving multi-GPU and multi-node scaling. You will utilize Python, C++, PyTorch, Triton, and collective communication libraries to identify bottlenecks and implement optimizations. Additionally, you will build agentic and AI-assisted workflows to automate profiling analysis, experiment orchestration, and regression triage. This role focuses on the technical challenge of optimizing large-scale deep learning training, specifically addressing SFT, LoRA, and RL-based workloads within a high-performance computing environment.

What you'll do

  • Drive performance for fine-tuning and RL training workloads on AMD Instinct GPUs.
  • Improve throughput, memory efficiency, and stability across data, model, optimizer, and runtime steps.
  • Optimize multi-GPU and multi-node training performance including communication and scaling behavior.
  • Develop efficient kernels and perform graph-level or runtime-level optimizations for high impact.
  • Profile, diagnose, and resolve bottlenecks while preventing regressions in CI and benchmarking pipelines.
  • Build agentic and AI-assisted workflows to accelerate profiling analysis and regression triage.
  • Create scalable tooling and automation to improve reproducibility and performance reporting across teams.
  • Deliver reproducible training pipelines and documentation for internal and external developers.

What we're looking for

  • Bachelor's, Master's, or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • Proven experience in GPU performance engineering for deep learning workloads using ROCm, HIP, Triton, or similar technologies.
  • Hands-on experience with SFT, LoRA, and RL-based training at scale.
  • Strong PyTorch experience including torch.distributed, FSDP/ZeRO, or equivalent distributed training approaches.
  • Proficiency in Python and C++ including the ability to read and write kernels.
  • Experience with distributed systems and collective communication libraries.
  • Ability to use AI-assisted workflows and agentic tooling for performance analysis and engineering productivity.
  • Track record of profiling, diagnosing bottlenecks, and implementing durable performance improvements across cross-stack components.

More like this

Similar roles

Fellow GPU Performance Optimization Engineer

Amd

San Jose, CA 167 days ago $268,000$402,000
GPU Distributed Training PyTorch JAX TensorFlow CUDA HIP ROCm Python C++ RDMA NCCL RCCL Megatron-LM Torchtitan MaxText Kernel Optimization Compiler Stack Performance Profiling
Hybrid

Software Engineer, GPU AI ML

Amd

Santa Clara, CA 77 days ago $204,000$306,000
C++ HIP CUDA ROCm PyTorch TensorFlow JAX GPU Architecture Kernel Optimization Distributed Systems LLMs SFT RLHF GRPO Quantization Verilog SystemVerilog RTL Design
Hybrid

AI Engineer, Recursive Self-Improvement for Compute

Amd

Santa Clara, CA 45 days ago $204,000$306,000
Python C++ CUDA HIP ROCm Triton PyTorch JAX TensorFlow Reinforcement Learning GPU Kernels Compiler Optimization Performance Engineering Profiling Distributed Training Agentic Workflows Hardware-aware Optimization
Hybrid