Failure Analysis Engineering Manager, GPU ASIC and PCBA Debug

Amd

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Secaucus, NJ
Salary
$147,200–$220,800 / yr
Posted
99 days ago
Freshness
Confirmed live 2 days ago
Closes
Jun 4, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $200k
This role $184k
$137k most similar roles pay here $240k

This role pays less than 60% of similar roles. Most pay $173,906–$226,025 — the shaded band above. At the midpoint, this role pays about $184k versus about $200k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Failure Analysis Engineering Manager, GPU ASIC and PCBA Debug

The Failure Analysis Engineering Manager, GPU ASIC and PCBA Debug joins the Quality Engineering team to lead and develop a high-performing team of failure analysis engineers. This role involves overseeing customer and factory investigations for GPU accelerators while driving failure reproduction, isolation, and root cause analysis in collaboration with cross-functional design, validation, firmware, and manufacturing teams. The manager provides technical leadership across power, ASIC, firmware, and thermal domains, utilizing schematics, layouts, and design documentation to guide board-level debug. Key responsibilities include developing triage automation, managing team growth, and presenting findings to stakeholders. Required skills include expertise in GPU ASIC debug, PCBA diagnostics, and hardware verification using oscilloscopes and logic analyzers. Technical proficiency includes Python, shell scripting, and experience with Windows and Linux environments while addressing complex issues in high-speed digital design and data center systems.

What you'll do

  • Lead a team of failure analysis engineers by setting priorities, providing technical guidance, and coaching them through complex investigations.
  • Provide technical leadership for the triage and debug of complex GPU and PCBA failures across power, ASIC, firmware, and thermals.
  • Define debug plans and direct investigation paths to identify root causes for customer and factory failures.
  • Drive the development of debug automation, diagnostic tools, and data analysis methods to improve triage efficiency.
  • Lead cross-functional triage with manufacturing partners and internal teams to align on failure hypotheses and corrective actions.
  • Guide board-level debug using schematics and design documentation while mentoring engineers through the process.
  • Document failure analysis results, root cause findings, and recovery plans for stakeholders and senior leadership.
  • Drive continuous improvement of failure analysis methods and best practices across all hardware and firmware domains.

What we're looking for

  • Bachelor's degree in Electrical Engineering, Computer Engineering, or a related field.
  • Proven experience as a people manager building, mentoring, and leading high-performing engineering teams.
  • Deep expertise in GPU ASIC debug, validation, and functional or stress test development.
  • Strong background in PCBA diagnostics, failure analysis, and board-level debug from NPI through production.
  • Experience leading triage across power, ASIC, firmware, and thermal failure domains.
  • Hands-on lab experience with oscilloscopes, logic analyzers, and custom debug tools.
  • Proficiency in Python and shell scripting within Windows and Linux environments.
  • Ability to read schematics, interpret datasheets, and identify components for board-level debug.

More like this

Similar roles

Failure Analysis & System Validation Engineer

Amd

Austin, TX 51 days ago $130,400$195,600
Failure Analysis Post Silicon Validation PCIe Memory Validation Python Linux Windows Firmware Power Management Root Cause Analysis NPI Oscilloscopes Logic Analyzers GPU Architecture Data Center Infrastructure AI Tools
Hybrid

Senior Failure Analysis Engineer, Test Development

Amd

Secaucus, NJ 98 days ago $138,000$207,000
GPU AI/ML Python Shell Scripting Linux Windows Firmware Hardware Automation Telemetry Data Center Infrastructure Failure Analysis Validation Diagnostics Test Development Scripting Orchestration

Lead Systems Debug Engineer, Data Center GPU

Amd

Austin, TX 64 days ago $163,200$244,800
GPU SoC PCIe HBM C C++ Python Shell Perl Git Agile Root-Cause Analysis Oscilloscopes Board-level Diagnostics Power Delivery Networking Cloud Infrastructure
8+ yrs exp Hybrid

Senior Failure Analysis Engineer

Nvidia

Santa Clara, CA 53 days ago $144,000$230,000
Physical Failure Analysis SEM TEM Focused Ion Beam FIB Electron Microscopy IC Fabrication Advanced Packaging De-processing Cross-sectioning Design of Experiments DOE Circuit Design
8+ yrs exp

Failure Analysis Engineer

Amd

Secaucus, NJ 66 days ago $96,880$145,320
Failure Analysis Root Cause Analysis Firmware BIOS BMC IPMI PCIe Linux System Architecture Signal Integrity Power Delivery Networking AI-assisted Engineering Automation Hardware Debugging

Senior Physical Failure Analysis Engineer

Nvidia

Santa Clara, CA 64 days ago $144,000$230,000
Physical Failure Analysis SEM TEM FIB IC Fabrication Advanced Packaging Cross-sectioning De-layering Plasma De-processing Electron Microscopy Design of Experiments Root-cause Analysis Silicon Architecture
8+ yrs exp