Fellow, AI Workload Optimization

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Bellevue, WA
Salary
$256,000–$384,000 / yr
Posted
14 days ago
Freshness
Confirmed live yesterday
Closes
Aug 28, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $194k
This role $320k
$120k most similar roles pay here $412k

This role pays more than 97% of similar roles. Most pay $149,375–$237,946 — the shaded band above. At the midpoint, this role pays about $320k versus about $194k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Fellow, AI Workload Optimization

Fellow, AI Workload Optimization joins the AI Software group to define and drive the end-to-end software optimization strategy for top-tier customers. This role sits at the intersection of architecture, customer engagement, and software engineering, focusing on tuning the ROCm stack, compilers, and high-level AI frameworks to maximize performance for demanding workloads. The individual will lead the profiling, analysis, and tuning of large-scale models including LLMs, Diffusion, Multimodal, and MoE. Key responsibilities include hardware-software co-design, developing tools for performance estimation, and engaging with hyperscalers to solve critical performance needs. Required expertise includes PyTorch, JAX, vLLM, SGLang, and the ROCm stack, alongside proficiency in TorchProfiler, ROCm Profiler, and Nsight. The role addresses the technical challenge of optimizing distributed inference and training across multi-node environments using techniques like quantization, speculative decoding, and FlashAttention.

What you'll do

  • Define and drive the end-to-end software optimization strategy for high-performance AI workloads.
  • Lead the profiling, analysis, and tuning of large-scale models like LLMs and Diffusion models.
  • Partner with top customers and hyperscalers to deliver tailored architectural wins and optimizations.
  • Influence future silicon features by collaborating across hardware architecture, compiler, and framework teams.
  • Drive the development of advanced tools for performance estimation, modeling, and automated reporting.
  • Act as a technical ambassador in industry forums and open-source communities.
  • Mentor and inspire the next generation of AMD engineers and technical leaders.

What we're looking for

  • Experience in a high-level technical leadership role such as Fellow or equivalent.
  • Expertise in AI Frameworks including PyTorch, JAX, vLLM, and SGLang.
  • Proficiency with the ROCm software stack.
  • Proven experience optimizing distributed inference and training across multi-node/multi-GPU environments.
  • Mastery of performance profiling tools such as TorchProfiler, ROCm Profiler, and Nsight.
  • Deep understanding of modern model architectures and optimization techniques like quantization and FlashAttention.
  • A PhD or Master’s degree in Computer Science, Electrical Engineering, or a related field.
  • Demonstrated research or applied experience in AI/ML fields including deep learning and large language models.

More like this

Similar roles

Fellow GPU Performance Optimization Engineer

Amd

San Jose, CA 167 days ago $268,000$402,000
GPU Distributed Training PyTorch JAX TensorFlow CUDA HIP ROCm Python C++ RDMA NCCL RCCL Megatron-LM Torchtitan MaxText Kernel Optimization Compiler Stack Performance Profiling
Hybrid

Software Engineer, GPU AI ML

Amd

Santa Clara, CA 77 days ago $204,000$306,000
C++ HIP CUDA ROCm PyTorch TensorFlow JAX GPU Architecture Kernel Optimization Distributed Systems LLMs SFT RLHF GRPO Quantization Verilog SystemVerilog RTL Design
Hybrid

HPC Systems Engineer, AI Workloads

Amd

San Jose, CA 2 days ago $173,600$260,400
GPU HPC Kubernetes Python SLURM RoCEv2 KVM Ubuntu Shell Ansible Saltstack Terraform Prometheus Grafana Distributed ML LLMs 400G Networking

Senior Staff Engineer, AI Workloads & Storage

Samsung Semiconductor

San Jose, CA 15 days ago $189,000$301,000
NVMe NAND Flash SSD Firmware Python C++ Rust Go eBPF fio perf ftrace blktrace vLLM SGLang Triton TensorRT-LLM RDMA NVMe-oF SPDK SystemC SimPy CXL PCIe Gen5
10+ yrs exp

Senior AI Performance and Efficiency Engineer

Nvidia

Remote (Santa Clara, CA) +2 176 days ago $152,000$241,500
Python Go Bash CUDA NCCL PyTorch TensorFlow NSight Systems NSight Compute AWS GCP Azure InfiniBand RDMA Lustre GPFS MLPerf Distributed Training Parallel Computing Machine Learning
5+ yrs exp Remote

AI Model Optimization Architect

Qualcomm

San Diego, CA 179 days ago $158,400$237,600
PyTorch Python Triton ONNX torch.compile TorchDynamo LLM VLM Transformer Distributed Systems Kernel Fusion Continuous Batching KVcache Machine Learning Computer Architecture
4+ yrs exp