Principal System Software Architect, AI/GPU Platforms

Amd

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Austin, TXSanta Clara, CA
Salary
$204,000–$306,000 / yr
Posted
95 days ago
Freshness
Confirmed live 2 days ago
Closes
Aug 19, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $213k
This role $255k
$147k most similar roles pay here $323k

This role pays more than 78% of similar roles. Most pay $172,326–$253,362 — the shaded band above. At the midpoint, this role pays about $255k versus about $213k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Principal System Software Architect, AI/GPU Platforms

As a Principal System Software Architect, AI/GPU Platforms, you will join the system software architecture team to design and oversee the software stack for MI400-class accelerators. You will own end-to-end architecture for subsystems including memory management, scheduling, RAS, and scale-up/scale-out interconnects across GPU, node, and fabric layers. Your daily work involves defining hardware-software interfaces during pre-silicon stages, developing kernel-mode drivers like amdgpu/KFD, and managing user-mode runtimes such as ROCr and HSA. You will solve complex problems regarding multi-GPU topologies, memory coherence, and large-scale cluster reliability for AI and HPC deployments. Key technical competencies include Linux Memory Management, HMM, RDMA, Infinity Fabric, and ROCm. You will also navigate technologies like UALink, RCCL, PyTorch, and TensorFlow to optimize performance for high-demand training and inference workloads across distributed systems.

What you'll do

  • Own the end-to-end system software architecture for MI400-class subsystems including memory management, scheduling, and serviceability.
  • Drive architecture across the full stack including kernel-mode drivers, user-mode runtimes, firmware interfaces, and the ROCm platform.
  • Partner with silicon and SoC architects during pre-silicon definition to shape hardware/software interfaces and programming models.
  • Define software strategies for multi-GPU and rack-scale topologies, including interconnects, memory coherence, and address translation.
  • Establish architecture for reliability, availability, and serviceability across large-scale clusters.
  • Identify performance bottlenecks in the launch path, memory subsystem, and communication paths to define corrective software mechanisms.
  • Produce technical specifications and reference designs to align firmware, driver, runtime, and framework teams on shared plans.
  • Act as a technical anchor to translate customer workload requirements into architectural direction for hyperscale and AI clients.

What we're looking for

  • Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, Computer Science, or a closely related field.
  • Experience with Linux Memory Management and Heterogeneous Memory Management (HMM) (preferred).
  • Experience with GPU/DRM driver development (preferred).
  • Knowledge of cache coherence and memory consistency protocols (preferred).
  • Experience with GPU networking technologies including scale up transport, NVLink, UALink, RDMA, and peer-direct (preferred).
  • Experience with scale-up/scale-out networking, communication collectives (RCCL/NCCL), MPI, or SHMEM (preferred).
  • Experience with data-center fabrics such as Infinity Fabric, UALink, Ultra Ethernet, InfiniBand, and RoCE (preferred).
  • Experience with ROCm stack, CUDA, or other GPU compute ecosystems; kernel/driver contributions; and deep learning frameworks like PyTorch or TensorFlow (preferred).

More like this

Similar roles

Software Engineer, GPU AI ML

Amd

Santa Clara, CA 79 days ago $204,000$306,000
C++ HIP CUDA ROCm PyTorch TensorFlow JAX GPU Architecture Kernel Optimization Distributed Systems LLMs SFT RLHF GRPO Quantization Verilog SystemVerilog RTL Design
Hybrid