Technical Program Manager, AI Cluster Validation

Amd

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Austin, TX
Salary
$185,600–$278,400 / yr
Posted
31 days ago
Freshness
Confirmed live yesterday
Closes
Aug 11, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $206k
This role $232k
$144k most similar roles pay here $293k

This role pays more than 75% of similar roles. Most pay $180,000–$232,475 — the shaded band above. At the midpoint, this role pays about $232k versus about $206k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Technical Program Manager, AI Cluster Validation

Technical Program Manager- AI Cluster Validation joins the engineering team to lead execution of AI cluster programs with a focus on GPU platforms, rack-level solutions, and AI cluster validation. This role manages end-to-end delivery from GPU and server integration through rack bring-up, scale testing, failure analysis, and system debug closure for hyperscale and enterprise AI deployments. The candidate will develop program plans, manage risk logs, and coordinate across hardware, firmware, BIOS/BMC, and networking teams to ensure readiness for customer workloads. Key responsibilities include overseeing multi-node and multi-rack scale testing, tracking high-impact failures like HSIO or network issues, and managing infrastructure components such as cabling, power, and cooling. Required skills include proficiency in Jira, Confluence, Excel, and PowerPoint, alongside a technical understanding of GPU-based systems, datacenter infrastructure, and hardware/firmware debug cycles.

What does a Technical Program Manager earn?

Median $193000 from 169 postings across 37 companies.

See salary data

What you'll do

  • Drive end-to-end execution of AI cluster engineering programs including GPU platforms and rack-level solutions.
  • Manage program artifacts such as schedules, dependency maps, resource forecasts, and risk logs.
  • Coordinate across hardware, firmware, and networking teams to ensure readiness for scale testing and customer workloads.
  • Lead the delivery of multi-node and multi-rack scale testing including test strategy and coverage tracking.
  • Oversee rack bring-up processes involving compute trays, cabling, power, cooling, and management infrastructure.
  • Act as the execution lead for platform debug to coordinate triage and root-cause analysis for system-level issues.
  • Track high-impact failures across GPU, firmware, and network components to ensure timely resolution.
  • Provide data-driven status reports and risk assessments to senior leadership.

What we're looking for

  • Bachelor's or master's degree in systems, EE, CS, or a related engineering discipline.
  • Experience leading complex hardware or AI infrastructure programs through bring-up, validation, and deployment phases.
  • Strong technical understanding of GPU-based AI systems, rack architectures, and datacenter infrastructure.
  • Proven ability to manage ambiguity, drive debug execution, and lead cross-functional teams without direct authority.
  • Strong written and verbal communication skills for executive-level status reporting.
  • Proficiency with program management tools such as Jira, Confluence, and Excel/PowerPoint.
  • Hands-on experience with GPU cluster scale testing, system stress, or performance validation (preferred).
  • Familiarity with rack-level bring-up, power/cooling constraints, networking, and failure modes at scale (preferred).

More like this

Similar roles

Technical Program Manager, Cluster Build

Amd

TX 46 days ago $162,640$243,960
AI HPC Cluster Build Hardware Firmware Infrastructure Validation Agile Jira Confluence Microsoft Project Microsoft Office Suite Project Management Scrum KPIs

Technical Program Manager, Cluster Build

Amd

TX 10 days ago $185,600$278,400
AI HPC Cluster Build Hardware Firmware Infrastructure Validation Agile Jira Confluence Microsoft Project Microsoft Office Suite PMP Scrum Master Project Management

Senior Technical Program Manager Atlas Clusters

MongoDB

Remote (Dublin, Ireland) +1 101 days ago
MongoDB Python Google Apps Script Slack Jira Rally MS Project Cloud Infrastructure Service Oriented Architecture Cloud Storage Scripting
10+ yrs exp Remote Hybrid

Principal Technical Program Manager, AI

Motorola Solutions

Remote (Waltham, MA) +4 105 days ago $160,000$220,000
Artificial Intelligence Machine Learning MLOps Computer Vision Data Platforms Video Processing Inference Optimization KPIs OKRs Analytics
10+ yrs exp Remote