Triton Compiler and Kernel Software Engineer

Amd

Confirmed live today High trust

Quick summary

Work type
On-site
Location
San Jose, CA
Salary
$204,000–$306,000 / yr
Posted
10 days ago
Freshness
Confirmed live today
Closes
Sep 29, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $206k
This role $255k
$134k $324k
below market most similar roles pay here above market

This role pays more than 94% of similar roles. Most pay $175,750–$235,750 — the blue band above. At the midpoint, this role pays about $255k versus about $206k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 502 open roles on FindRole.

Listed pay typically runs $168,000–$252,000 across 502 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Triton Compiler and Kernel Software Engineer

The Senior Triton Compiler and Kernel Engineer will join the team to advance Triton performance and capabilities on AMD GPUs. This role involves developing and optimizing Triton compiler support, improving compiler lowering, optimization, scheduling, and code generation. The engineer will create high-performance kernels for attention, GEMM, and MoE workloads while developing multi-GPU communication primitives. Key responsibilities include analyzing compute, memory, and occupancy, as well as contributing designs to upstream Triton and LLVM/MLIR. The position requires expertise in GPU architecture, compiler development, and distributed AI training. Candidates should possess skills in Triton, HIP, CUDA, or GPU assembly, along with knowledge of ROCm, RCCL, and NCCL. The work focuses on connecting AI frameworks to GPU hardware through high-performance software abstractions.

What you'll do

  • Develop and optimize Triton compiler support specifically for AMD GPUs.
  • Improve compiler lowering, optimization, scheduling, and code generation processes.
  • Create high-performance kernels for attention, GEMM, MoE, and other AI workloads.
  • Develop and optimize multi-GPU kernels and communication primitives.
  • Enable new AMD GPU and interconnect capabilities through effective Triton abstractions.
  • Analyze compute, memory, communication, occupancy, register usage, and generated code.
  • Resolve complex correctness and performance issues across single- and multi-GPU workloads.
  • Contribute designs and implementations to upstream Triton and LLVM/MLIR projects.

What we're looking for

  • Bachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.
  • Experience in GPU architecture and programming.
  • Experience in compiler development and optimization.
  • Experience in high-performance GPU kernels.
  • Experience in multi-GPU communication and collective operations.
  • Experience in distributed AI training and inference.
  • Experience in AI workload and framework performance.
  • Experience in low-level performance analysis.
  • Experience with AMD GPU architecture and ROCm (preferred).
  • Experience optimizing kernels with Triton, HIP, CUDA, or GPU assembly (preferred).
  • Experience with Triton, LLVM, MLIR, or another optimizing compiler (preferred).
  • Knowledge of GPU execution models, memory hierarchies, synchronization, and instruction pipelines (preferred).
  • Experience with collective communication, distributed programming, and libraries such as RCCL or NCCL (preferred).
  • Understanding of communication topologies, interconnects, synchronization, and communication–computation overlap (preferred).
  • Familiarity with distributed training, inference, tensor parallelism, expert parallelism, or pipeline parallelism (preferred).
  • Experience using AI coding tools and autonomous agents to accelerate software development (preferred).
  • Familiarity with AI primitives, reduced-precision formats, and performance profiling (preferred).
  • Contributions to complex or open-source software projects (preferred).
  • Strong analytical, debugging, communication, and collaboration skills (preferred).

More like this

Similar roles

Senior Software Engineer, AI Triton Kernels

Amd

San Jose, CA 63 days ago $145,600–$218,400
Triton PyTorch vLLM SGLang LLVM MLIR ROCm GPU Kernel Development SIMT AMDGPU MoE Performance Engineering Quantized Kernels Instruction Scheduling Hardware Counters Profiling Microbenchmarking

Senior Performance Compiler Engineer

Nvidia

Remote (Redmond, WA) +5 155 days ago $184,000–$287,500
Triton MLIR C++ CUDA PTX Python OpenCL OpenMP TVM High-Performance Computing Parallel Programming Computer Architecture Assembly BLAS
8+ yrs exp Remote

NPU Compiler Engineer

Amd

San Jose, CA 17 days ago $172,000–$258,000
C++ Python MLIR Compiler Architecture NPU Vitis-AI Ryzen-AI Machine Learning Frameworks SystemC RTL Simulators Graph Compilers Object-Oriented Programming

Senior Software Engineer, AI Triton Communication

Amd

San Jose, CA 35 days ago $164,000–$246,000
Triton ROCm HIP CUDA MLIR LLVM PyTorch vLLM SGLang RCCL NCCL NVSHMEM rocSHMEM MPI InfiniBand PCIe XGMI Distributed Systems Compiler Development GPU Architecture HPC Performance Engineering

Senior Software Development Engineer, GPU Kernel Development

Amd

Santa Clara, CA 102 days ago $240,000–$360,000
C++ Python PyTorch TensorFlow HIP CUDA ROCm LLVM Triton Assembly GPU Kernel Development Distributed Computing Compiler Optimization Linux High-Performance Computing CUTLASS Compute Kernel Graph Compilers