Failure Analysis Engineer, Server Systems Integration

Amd

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Secaucus, NJ
Salary
$154,800–$232,200 / yr
Posted
21 days ago
Freshness
Confirmed live today
Closes
Sep 11, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $178k
This role $194k
$132k most similar roles pay here $243k

This role pays more than 65% of similar roles. Most pay $142,300–$214,000 — the shaded band above. At the midpoint, this role pays about $194k versus about $178k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 474 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 474 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Failure Analysis Engineer, Server Systems Integration

The Failure Analysis Engineer - Server Systems Integration joins the team to own end-to-end failure analysis across hardware, firmware, silicon, and integration for Helios server platform bring-up. This role involves diagnosing system-level failures by determining if issues stem from firmware behavior, hardware design, component quality, power delivery, or cross-domain interactions. The engineer will perform structured debug using register dumps, BIOS/BMC logs, telemetry, and oscilloscope captures to identify root causes across PCIe, high-speed interconnects, and power sequencing. Key responsibilities include developing diagnostic strategies, automating data collection with Python and Linux tools, and leading cross-functional technical interfaces to drive corrective actions. The role requires expertise in server platforms, including CPLD, PMBus, and signal path dependencies, to resolve complex issues within the server and hyperscale computing domain while ensuring high platform quality through rigorous evidence-based analysis.

What you'll do

  • Lead end-to-end failure analysis and root cause ownership for Helios server platform bring-up and system-level failures.
  • Debug hardware, firmware, and silicon interactions across BIOS, BMC, CPLD, PMBus, and high-speed interconnects.
  • Determine if failures originate from firmware behavior, hardware design, component quality, silicon issues, or power delivery.
  • Execute structured debug plans using register dumps, log analysis, oscilloscope captures, and logic analyzer traces.
  • Validate firmware-controlled hardware behaviors across boot, initialization, and various failure states.
  • Automate data collection, log parsing, and failure correlation using Python and Linux scripting tools.
  • Manage technical interfaces with cross-functional teams to drive corrective actions and resolve complex system issues.
  • Create technical reports and 8D documentation to communicate root causes and recommended actions to stakeholders.

What we're looking for

  • Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, Computer Science, or a related discipline (or equivalent hands-on experience).
  • Experience in end-to-end failure analysis and root cause ownership for server platform bring-up and system-level failures.
  • Expertise in debugging BIOS, BMC, CPLD, PMBus, POST, PCIe, and high-speed interconnect initialization.
  • Ability to isolate issues across firmware behavior, hardware design, silicon behavior, power delivery, and component quality.
  • Proficiency with oscilloscopes, logic analyzers, protocol tools, and other lab validation equipment.
  • Experience analyzing register dumps, BIOS/BMC logs, telemetry, and schematic evidence to drive corrective actions.
  • Proficiency in Python, Linux tools, and scripting for automating data collection and log parsing.
  • Knowledge of PCIe, high-speed interfaces, retimers, link training, and signal path dependencies (preferred).

More like this

Similar roles

Failure Analysis Engineer

Amd

Secaucus, NJ 87 days ago $96,880–$145,320
Failure Analysis Root Cause Analysis Firmware BIOS BMC IPMI PCIe Linux System Architecture Signal Integrity Power Delivery Networking AI-assisted Engineering Automation Hardware Debugging

Senior Engineer, Server System Integration and Debug

Qualcomm

San Diego, CA 15 days ago $117,800–$176,600
Python Linux PCIe Ethernet JTAG Oscilloscopes Logic Analyzers Firmware Hardware Debug Root Cause Analysis Test Automation DevOps HPC Signal Integrity System Bring-up
2+ yrs exp

GPU Boards Failure Analysis Engineer

Amd

Secaucus, NJ 23 days ago $130,400–$195,600
GPU Architecture Failure Analysis Python Shell Scripting Windows Linux Hardware Validation Firmware High Performance Computing PCIe HBM GDDR Oscilloscopes Logic Analyzers Soldering IPC-A-610 MS Excel Data Center Infrastructure
Hybrid

Senior Failure Analysis Engineer

Nvidia

Santa Clara, CA 74 days ago $144,000–$230,000
Physical Failure Analysis SEM TEM Focused Ion Beam FIB Electron Microscopy IC Fabrication Advanced Packaging De-processing Cross-sectioning Design of Experiments DOE Circuit Design
8+ yrs exp