Fellow Software Engineer, AI Performance & Reliability

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Jose, CABellevue, WA
Salary
$268,000–$402,000 / yr
Posted
42 days ago
Freshness
Confirmed live yesterday
Closes
Jul 30, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $201k
This role $335k
$122k most similar roles pay here $432k

This role pays more than 98% of similar roles. Most pay $165,375–$235,750 — the shaded band above. At the midpoint, this role pays about $335k versus about $201k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Fellow Software Engineer, AI Performance & Reliability

Fellow Software Engineer — AI Performance & Reliability will join the AI Infrastructure team to improve the performance, efficiency, and reliability of AI workloads across model training and inference. This role involves profiling and optimizing large language models, diffusion models, and recommendation systems while identifying bottlenecks across frameworks, compilers, runtimes, operating systems, and hardware. The engineer will develop performance tooling, benchmarks, and observability systems while collaborating with customers to resolve complex production issues and translate feedback into infrastructure improvements. Key technical requirements include proficiency in Python or C++, experience with PyTorch, TensorFlow, or JAX, and a strong foundation in computer architecture and distributed communication. Candidates should possess expertise in GPU environments and tools like ROCm, HIP, CUDA, Triton, XLA, MLIR, and NCCL to solve challenges related to throughput, latency, and memory efficiency at scale.

What does a Software Engineer earn in California?

Median $214000 from 775 postings across 63 companies.

See salary data

What you'll do

  • Profile and optimize performance, throughput, latency, and memory efficiency for AI training and inference workloads.
  • Identify and resolve technical bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.
  • Develop reusable performance tooling, benchmarks, automation, and observability systems for machine learning infrastructure.
  • Investigate and resolve complex production issues affecting large language models, diffusion models, and recommendation systems.
  • Partner directly with customers to understand technical requirements, reproduce issues, and provide tailored solutions.
  • Translate customer feedback into actionable product improvements and infrastructure capabilities.
  • Document performance findings, technical recommendations, and best practices for internal and external stakeholders.

What we're looking for

  • A PhD or a master's degree in artificial intelligence, machine learning, computer science, or a related field.
  • Strong software engineering skills and experience building production-quality systems.
  • Experience working with AI infrastructure for model training, inference, or both.
  • Demonstrated experience profiling and optimizing machine learning models or AI workloads.
  • Strong foundations in computer architecture including processors, memory hierarchies, parallelism, and performance tradeoffs.
  • Proficiency in systems-oriented programming languages such as Python or C++.
  • Experience with machine learning frameworks like PyTorch, TensorFlow, or JAX.
  • Experience with GPU, accelerator, or distributed computing environments.

More like this

Similar roles

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp

Fellow GPU Performance Optimization Engineer

Amd

San Jose, CA 167 days ago $268,000$402,000
GPU Distributed Training PyTorch JAX TensorFlow CUDA HIP ROCm Python C++ RDMA NCCL RCCL Megatron-LM Torchtitan MaxText Kernel Optimization Compiler Stack Performance Profiling
Hybrid

Software Engineer, GPU AI ML

Amd

Santa Clara, CA 77 days ago $204,000$306,000
C++ HIP CUDA ROCm PyTorch TensorFlow JAX GPU Architecture Kernel Optimization Distributed Systems LLMs SFT RLHF GRPO Quantization Verilog SystemVerilog RTL Design
Hybrid

ML Systems Research Engineer, RL / Inference / Agent Systems

Amd

Santa Clara, CA 44 days ago $204,000$306,000
Python PyTorch JAX TensorFlow Reinforcement Learning RLHF LLM Agents Kubernetes Ray Slurm CUDA ROCm HIP Distributed Systems Data Pipelines Model Serving Compiler Optimization Profiling Hardware Engineering
Hybrid

Lead Software Engineer, AI Platform Reliability

JPMorgan Chase

Seattle, WA 37 days ago
Python Artificial Intelligence Machine Learning Generative AI System Design Distributed Systems Cloud Platforms Infrastructure Engineering SDKs Observability Logging Metrics Service Level Objectives Incident Response Automated Testing root-cause analysis
5+ yrs exp

Staff Software Engineer, Core AI Infrastructure

Coinbase

Remote 49 days ago $218,025$256,500
AWS Kubernetes Terraform Go Python Docker CI/CD Ansible Chef Puppet Salt Git Bash Ruby Distributed Systems Data Pipelines infrastructure-as-code
Remote