Failure Analysis Engineer

Amd

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Secaucus, NJ
Salary
$96,880–$145,320 / yr
Posted
66 days ago
Freshness
Confirmed live yesterday
Closes
Jul 7, 2027

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $162k
This role $121k
$85k most similar roles pay here $206k

This role pays less than 82% of similar roles. Most pay $130,500–$193,000 — the shaded band above. At the midpoint, this role pays about $121k versus about $162k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Failure Analysis Engineer

As a Senior Failure Analysis Engineer, you will join a high-visibility team focused on bringing up, debugging, and improving next-generation server and rack-scale platforms in data center environments. You will perform system-level and rack-level failure analysis to identify root causes for complex hardware and firmware interactions involving CPUs, GPUs, memory, PCIe, networking, power delivery, and thermal systems. Day-to-day responsibilities include analyzing platform telemetry, BIOS, BMC, and system logs to resolve issues while collaborating with design, manufacturing, and quality teams. You will utilize lab debugging equipment, Linux-based environments, and AI-assisted engineering workflows to improve reliability and documentation. The role addresses critical technical challenges in the data center space, ensuring manufacturing readiness and operational success for high-performance computing products by providing technical guidance and structured debug processes across the entire product lifecycle.

What you'll do

  • Perform system-level and rack-level failure analysis on server and data center platforms from symptom identification to resolution.
  • Debug complex hardware, firmware, and platform issues involving CPUs, GPUs, memory, PCIe, networking, and power systems.
  • Analyze system logs, telemetry, BIOS, and BMC data to identify failure patterns and validate corrective actions.
  • Reproduce failures in lab environments using diagnostic tools to improve overall platform reliability.
  • Provide technical guidance and triage support to manufacturing and ODM partners for complex platform issues.
  • Create technical documentation, troubleshooting guides, and SOPs to scale organizational knowledge and improve debug processes.
  • Utilize AI-assisted tools and automation to accelerate investigations and improve engineering efficiency.

What we're looking for

  • Bachelor's or master's degree in Electrical Engineering, Computer Engineering, Systems Engineering, Computer Science, or a related technical field.
  • Experience performing system-level, server-level, or rack-level troubleshooting and failure analysis.
  • Background in hardware debugging, root cause analysis, and complex issue resolution.
  • Familiarity with Linux-based environments, platform logs, and diagnostic workflows.
  • Experience supporting server, data center, networking, storage, or enterprise hardware platforms.
  • Exposure to BIOS, BMC, IPMI, firmware interactions, or platform management technologies.
  • Understanding of networking technologies, signal integrity concepts, power delivery, or high-speed interfaces.
  • Familiarity with lab debugging equipment and system diagnostic tools.

More like this

Similar roles

Failure Analysis & System Validation Engineer

Amd

Austin, TX 51 days ago $130,400$195,600
Failure Analysis Post Silicon Validation PCIe Memory Validation Python Linux Windows Firmware Power Management Root Cause Analysis NPI Oscilloscopes Logic Analyzers GPU Architecture Data Center Infrastructure AI Tools
Hybrid

Failure Analysis Engineering Manager, GPU ASIC and PCBA Debug

Amd

Secaucus, NJ 99 days ago $147,200$220,800
GPU ASIC PCBA Failure Analysis Root Cause Analysis Python Shell Scripting Linux Windows Oscilloscopes Logic Analyzers Firmware Hardware Verification High-speed Digital Design HBM GDDR PCIe Schematics Data Analysis

Senior Failure Analysis Engineer

Nvidia

Santa Clara, CA 53 days ago $144,000$230,000
Physical Failure Analysis SEM TEM Focused Ion Beam FIB Electron Microscopy IC Fabrication Advanced Packaging De-processing Cross-sectioning Design of Experiments DOE Circuit Design
8+ yrs exp

Senior Failure Analysis Engineer

Nvidia

Santa Clara, CA 91 days ago $144,000$230,000
Python Rust Shell Scripting Linux CI/CD Data Pipelines AI Machine Learning EDA Workflows CAD Navigation Systems Databases Version Control Observability Automation Frameworks Workflow Orchestration
8+ yrs exp Hybrid

Senior Physical Failure Analysis Engineer

Nvidia

Santa Clara, CA 64 days ago $144,000$230,000
Physical Failure Analysis SEM TEM FIB IC Fabrication Advanced Packaging Cross-sectioning De-layering Plasma De-processing Electron Microscopy Design of Experiments Root-cause Analysis Silicon Architecture
8+ yrs exp

Senior Failure Analysis Engineer, Test Development

Amd

Secaucus, NJ 98 days ago $138,000$207,000
GPU AI/ML Python Shell Scripting Linux Windows Firmware Hardware Automation Telemetry Data Center Infrastructure Failure Analysis Validation Diagnostics Test Development Scripting Orchestration