Senior Failure Analysis Engineer, Test Development

Amd

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Secaucus, NJ
Salary
$138,000–$207,000 / yr
Posted
98 days ago
Freshness
Confirmed live 2 days ago
Closes
Jun 5, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $172k
This role $172k
$130k most similar roles pay here $215k

This role pays more than 53% of similar roles. Most pay $146,500–$197,562 — the shaded band above. At the midpoint, this role pays about $172k versus about $172k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Senior Failure Analysis Engineer, Test Development

Senior Failure Analysis Engineer - Test Development joins the Quality Engineering team to develop advanced test methods that surface elusive failures in GPU accelerator platforms. The role involves designing custom execution flows, creating stress-based scenarios, and building test content to improve repeatability and shorten debug cycles for lab, factory, and customer-return cases. The engineer will build intelligent test systems using internal knowledge and live model inference to guide real-time decisions. Key responsibilities include architecting targeted methods for rack-scale environments, developing automation scripts in Python and shell, and interpreting telemetry logs across Windows and Linux platforms. The position requires expertise in GPU and server platform behavior, VPOD environments, AI/ML workloads, and firmware interactions. This role solves the challenge of converting vague symptoms into testable conditions to accelerate root cause identification for complex hardware and software systems.

What you'll do

  • Architect targeted test methods to identify hard-to-capture platform behaviors across GPU and server environments.
  • Create new workload patterns and stress combinations to reveal conditions not covered by standard diagnostics.
  • Build and maintain VPOD-based environments to support scalable experimentation and long-duration reproduction studies.
  • Use inference and training activities as system stimuli to probe platform limits and timing sensitivities.
  • Develop automation, scripting, and orchestration tools to manage workloads and analyze results across Windows and Linux.
  • Interpret telemetry and logs to refine experiments and isolate specific trigger conditions for root cause analysis.
  • Create AI-enabled execution flows that use internal knowledge and live inference to guide test branching.
  • Translate vague symptoms from field issues into repeatable test content in collaboration with cross-functional teams.

What we're looking for

  • Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field.
  • Proven track record of developing custom test methodologies for intermittent or difficult-to-observe failure modes.
  • Strong foundation in GPU and server platform behavior, including system stress interactions and stability characterization.
  • Demonstrated ability to build, run, and optimize VPOD environments for large-scale failure analysis or validation.
  • Hands-on familiarity with inference and training environments as controllable system stressors.
  • Proficiency in Python, shell scripting, and automation development for workload orchestration and telemetry capture.
  • Ability to interpret system data and debug artifacts to identify signals and guide experimental steps.
  • Experience building AI-enabled test systems that incorporate internal engineering knowledge and real-time inference.

More like this

Similar roles

Failure Analysis & System Validation Engineer

Amd

Austin, TX 50 days ago $130,400$195,600
Failure Analysis Post Silicon Validation PCIe Memory Validation Python Linux Windows Firmware Power Management Root Cause Analysis NPI Oscilloscopes Logic Analyzers GPU Architecture Data Center Infrastructure AI Tools
Hybrid

Senior Failure Analysis Engineer

Nvidia

Santa Clara, CA 91 days ago $144,000$230,000
Python Rust Shell Scripting Linux CI/CD Data Pipelines AI Machine Learning EDA Workflows CAD Navigation Systems Databases Version Control Observability Automation Frameworks Workflow Orchestration
8+ yrs exp Hybrid

Senior Failure Analysis Engineer

Nvidia

Santa Clara, CA 53 days ago $144,000$230,000
Physical Failure Analysis SEM TEM Focused Ion Beam FIB Electron Microscopy IC Fabrication Advanced Packaging De-processing Cross-sectioning Design of Experiments DOE Circuit Design
8+ yrs exp

Senior Physical Failure Analysis Engineer

Nvidia

Santa Clara, CA 64 days ago $144,000$230,000
Physical Failure Analysis SEM TEM FIB IC Fabrication Advanced Packaging Cross-sectioning De-layering Plasma De-processing Electron Microscopy Design of Experiments Root-cause Analysis Silicon Architecture
8+ yrs exp

Failure Analysis Engineer

Amd

Secaucus, NJ 66 days ago $96,880$145,320
Failure Analysis Root Cause Analysis Firmware BIOS BMC IPMI PCIe Linux System Architecture Signal Integrity Power Delivery Networking AI-assisted Engineering Automation Hardware Debugging

Senior SWQA Test Development Engineer

Nvidia

Santa Clara, CA 11 days ago $168,000$270,250
Python Docker Kubernetes CI/CD Linux Microservices LLM AI NeMo NIM Test Automation Distributed Systems Reliability Testing root‑cause analysis
8+ yrs exp