AI Cluster Technical Program Manager, Validation, Debug & Agentic AI

Amd

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Austin, TX
Salary
$162,640–$243,960 / yr
Posted
43 days ago
Freshness
Confirmed live yesterday
Closes
Jul 29, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $214k
This role $203k
$151k most similar roles pay here $273k

This role pays less than 61% of similar roles. Most pay $182,250–$246,700 — the shaded band above. At the midpoint, this role pays about $203k versus about $214k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · AI Cluster Technical Program Manager, Validation, Debug & Agentic AI

AI Cluster Technical Program Manager – Validation, Debug & Agentic AI serves as a technical leader within the engineering team to drive execution of AI cluster programs focused on GPU platforms and rack-level solutions. The role involves managing end-to-end delivery from server integration through rack bring-up, scale testing, and system debug closure for hyperscale and enterprise deployments. Responsibilities include defining program plans, tracking hardware and firmware issues, coordinating multi-node scale testing, and leading incident management to ensure fleet reliability. The candidate will utilize Jira, Confluence, and Excel while championing Agentic AI and AIOps solutions for automated triage and log analysis. This position addresses the complex technical challenges of GPU-based infrastructure, requiring expertise in hardware, firmware, networking, and power/cooling constraints to resolve system-level issues and ensure readiness for large-scale AI workloads across various production phases.

What you'll do

  • Manage end-to-end execution of AI cluster engineering programs including GPU platform integration and rack bring-up.
  • Develop and maintain program artifacts such as schedules, dependency maps, risk logs, and executive status reports.
  • Coordinate across hardware, firmware, and networking teams to ensure readiness for scale testing and customer workloads.
  • Lead multi-node and multi-rack scale testing including test strategy, coverage tracking, and infrastructure readiness gates.
  • Drive system-level debug activities to perform root-cause analysis on GPU, network, and firmware failures.
  • Manage critical deployment incidents by leading cross-functional war rooms and establishing operational metrics like MTTR.
  • Implement Agentic AI and AIOps solutions to automate incident triage, log analysis, and operational workflows.

What we're looking for

  • Bachelor's or master's degree in systems, EE, CS, or a related engineering discipline.
  • PMP, Scrum Master, or equivalent program management training.
  • Experience leading complex hardware or AI infrastructure programs through bring-up, validation, and deployment phases.
  • Strong technical understanding of GPU-based AI systems, rack architectures, and datacenter infrastructure.
  • Proven ability to manage ambiguity, drive debug execution, and lead cross-functional teams without direct authority.
  • Proficiency with program management tools such as Jira, Confluence, dashboards, Excel, and PowerPoint.
  • Strong written and verbal communication skills for executive-level status reporting.

More like this

Similar roles

Technical Program Manager, AI Cluster Validation

Amd

Austin, TX 30 days ago $185,600$278,400
GPU AI Infrastructure Rack Architecture Firmware BIOS BMC Networking power, cooling EVT DVT PVT Jira Confluence Excel PowerPoint Root-Cause Analysis System Debug Scale Testing datacenter infrastructure

Technical Program Manager, Cluster Build

Amd

TX 45 days ago $162,640$243,960
AI HPC Cluster Build Hardware Firmware Infrastructure Validation Agile Jira Confluence Microsoft Project Microsoft Office Suite Project Management Scrum KPIs

Technical Program Manager, Cluster Build

Amd

TX 9 days ago $185,600$278,400
AI HPC Cluster Build Hardware Firmware Infrastructure Validation Agile Jira Confluence Microsoft Project Microsoft Office Suite PMP Scrum Master Project Management

AI Validation and Test Engineer

Amd

Secaucus, NJ 48 days ago $136,320$204,480
AI Machine Learning GPU Python Linux Bash PowerShell Firmware BIOS BMC Networking Storage Data Center HPC Root Cause Analysis Automation Telemetry System-level Validation
8+ yrs exp