Senior Data Center GPU Validation and Debug Engineer

Amd

Confirmed live today High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Austin, TX
Salary
$174,400–$261,600 / yr
Posted
20 days ago
Freshness
Confirmed live today
Closes
Sep 10, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $208k
This role $218k
$139k most similar roles pay here $275k

This role pays more than 67% of similar roles. Most pay $177,250–$238,025 — the shaded band above. At the midpoint, this role pays about $218k versus about $208k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 374 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 374 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Senior Data Center GPU Validation and Debug Engineer

The Sr. Data Center GPU Validation and Debug Engineer joins the team to validate and debug data center GPU products across the complete software and hardware stack. This role involves identifying root causes for hangs, crashes, memory faults, and performance regressions while isolating issues across firmware, kernel, compiler, library, and framework layers. The engineer will characterize performance bottlenecks, build automated diagnostics and stress tests, and validate AI, HPC, and communication workloads. Key technical requirements include proficiency in Linux systems engineering, C/C++, Python, and shell automation. Candidates should possess expertise in GPU architecture, memory hierarchy, and multi-GPU topologies. Experience with ROCm/HIP, PyTorch, vLLM, and SGLang is preferred. The role addresses critical challenges in high-performance computing by ensuring platform health and performance across complex hardware environments involving PCIe interconnects, NUMA, and rack-scale systems.

What you'll do

  • Debug GPU failures including hangs, crashes, memory faults, and performance regressions across the full stack.
  • Reproduce system failures and reduce them to minimal, actionable test cases.
  • Triage complex interactions between GPUs, CPUs, memory, networking, power, and platform topology.
  • Validate AI, HPC, and communication workloads across GPU platforms and software releases.
  • Characterize performance and identify compute, memory, and communication bottlenecks.
  • Analyze multi-GPU behavior under various configurations including NUMA, power, and thermal conditions.
  • Establish benchmarks, baselines, and regression-detection methods for real workload behavior.
  • Build automated diagnostics, stress tests, and performance suites for validation infrastructure.

What we're looking for

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.
  • Strong Linux and systems-engineering experience with C/C++, Python, and shell automation.
  • Experience with data center GPUs, accelerators, or comparable high-performance systems.
  • Understanding of GPU architecture, memory hierarchy, parallel execution, communication, and synchronization.
  • Experience debugging across software, firmware, kernel, and hardware boundaries.
  • Ability to analyze logs, traces, hardware telemetry, and performance counters.
  • Knowledge of Linux kernel drivers, PCIe, firmware interaction, and memory management (preferred).
  • Familiarity with GPU programming models, AI frameworks, and multi-GPU topology (preferred).

More like this

Similar roles

Lead Systems Debug Engineer, Data Center GPU

Amd

Austin, TX 83 days ago $163,200–$244,800
GPU SoC PCIe HBM C C++ Python Shell Perl Git Agile Root-Cause Analysis Oscilloscopes Board-level Diagnostics Power Delivery Networking Cloud Infrastructure
8+ yrs exp Hybrid

Senior Data Center GPU Firmware Engineer

Amd

Santa Clara, CA +1 22 days ago $161,200–$241,800
Firmware GPU Drivers OS Kernel C C++ Python Bash Linux CUDA ROCm OpenCL Hardware Bring-up Server Architecture BIOS Performance Optimization
Hybrid

Senior GPU Validation Engineer

Nvidia

Santa Clara, CA 15 days ago $140,000–$224,250
Python C++ C Object-Oriented Programming AI Windows Linux RESTful APIs SQL Elasticsearch WinDBG gdb CUDA DLSS Frame Generation Reflex G-Sync SDKs