Fellow AI Performance Software Architect

Amd

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$268,000–$402,000 / yr
Posted
122 days ago
Freshness
Confirmed live yesterday
Closes
Jun 10, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $224k
This role $335k
$141k $430k
below market most similar roles pay here above market

This role pays more than 95% of similar roles. Most pay $195,000–$253,500 — the blue band above. At the midpoint, this role pays about $335k versus about $224k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 502 open roles on FindRole.

Listed pay typically runs $168,000–$252,000 across 502 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Fellow AI Performance Software Architect

Fellow, AI Performance Software Architect joins the AI Software Solutions Team to optimize the software ecosystem for next-generation GPU computational accelerators. This role involves enabling deep learning models, libraries, and applications for Instinct GPUs across on-prem and cloud environments. You will work with sophisticated clients to combine new hardware with the latest applications, libraries, frameworks, and SDKs to solve complex challenges. Key responsibilities include analyzing and optimizing AI software performance, understanding hardware bottlenecks, and harnessing performance to hit close to the roofline. Required skills include strong programming in Python and C++, along with development experience in major deep learning frameworks for inference, fine-tuning, and training. Preferred expertise includes profiling tools like Torchprofiler and Nsight, as well as implementing parallel methods using NCCL, RCCL, OpenMP, and MPI.

What you'll do

  • Enable and optimize deep learning models, libraries, and applications for Instinct GPUs.
  • Develop and optimize AI software for on-prem and cloud environments.
  • Analyze and optimize AI software performance to overcome hardware bottlenecks.
  • Implement and optimize parallel methods on GPU accelerators using NCCL/RCCL, OpenMP, or MPI.
  • Use profiling tools like Torchprofiler, RocM profiler, Vtune, and Nsight to analyze the AI software stack.
  • Develop custom AI software solutions for industry-leading customers to leverage hardware capabilities.
  • Root-cause and address complex performance issues across CPU and GPU systems.
  • Provide clear and timely project status communications to the leadership team.

What we're looking for

  • Strong programming skills in C++ and Python.
  • Strong development experience in at least one major DL framework for inference, fine-tuning, and/or training.
  • MS with years of related experience or PhD with years of related experience in Computer Science, Computer Engineering, or a related equivalent.
  • Experience developing software and system-level performance optimizations with a solid architecture understanding in GPUs (preferred).
  • Experience with open-source software development, including collaboration with community maintainers and submitting contributions (preferred).
  • Publications in reputed peer-reviewed ML conferences or journals (preferred).
  • Expertise in profiling tools across the AI SW Stack, such as Torchprofiler, RocM profiler, Vtune, and Nsight (preferred).
  • Experience implementing and optimizing parallel methods on GPU accelerators, such as NCCL/RCCL, OpenMP, and MPI (preferred).

More like this

Similar roles

Senior AI Performance Architect

Microsoft

4 days ago $119,800–$234,700
AI Systems Computer Architecture Performance Modeling High-Performance Computing Distributed Systems Python C++ CUDA Triton ROCm HIP Data Parallelism HBM PCIe Compilers
5+ yrs exp Hybrid

Distinguished Technologist, AI Model Performance Architect

HP Inc.

Spring, TX +1 123 days ago $190,000–$274,000
AI Model Performance HW/SW Co-design Hardware Architecture Electrical Engineering FPGA MATLAB Systems Design Simulations Schematic Capture Printed Circuit Board GPU NPU CPU Memory Architecture Debugging New Product Development
10+ yrs exp

Senior Solutions Architect, AI Performance Engineering

Nvidia

Santa Clara, CA 86 days ago $184,000–$287,500
CUDA C++ Python Linux GPU Kernel Optimization cuDNN cuBLAS CUTLASS NCCL NVSHMEM InfiniBand RoCE NVLINK DeepGEMM FlashAttention Nsight Compute High-Performance Computing Parallel Programming
8+ yrs exp

Principal AI Performance Engineer

Arm Holdings

San Jose, CA 18 days ago $262,700–$355,400
Arm Python C++ Triton CUDA DNN Parallel Computing Performance Optimization AI Frameworks
Hybrid

Principal AI Performance Engineer

Arm Holdings

San Jose, CA 18 days ago $262,700–$355,400
Arm Python C++ Triton CUDA DNN Parallel Computing Performance Optimization AI Frameworks Hardware Architecture
Hybrid

Principal AI Software & Systems Architect

Amd

Austin, TX +2 137 days ago $212,000–$318,000
AI Software Stacks Inference Runtimes Heterogeneous Computing Hardware-aware Software Architecture GPU Performance Engineering Compilers Drivers Operating Systems Model Serving Benchmarking Accelerator Libraries