Senior Staff AI Infrastructure Validation Engineer

Amd

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Secaucus, NJ
Salary
$164,080–$246,120 / yr
Posted
53 days ago
Freshness
Confirmed live yesterday
Closes
Jul 20, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $203k
This role $205k
$148k most similar roles pay here $257k

This role pays more than 54% of similar roles. Most pay $162,750–$243,650 — the shaded band above. At the midpoint, this role pays about $205k versus about $203k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Senior Staff AI Infrastructure Validation Engineer

The Senior Staff AI Infrastructure Validation Engineer joins the team to lead system-level, rack-level, and cluster-scale validation for next-generation AI infrastructure and accelerated computing platforms. This role involves defining validation strategies, developing automated frameworks, and creating telemetry-driven analysis pipelines to qualify products from initial bring-up through deployment readiness. The candidate will manage complex hardware-software interactions across CPUs, GPUs, memory subsystems, PCIe fabrics, firmware, BIOS/UEFI, BMCs, networking, and storage components. Key responsibilities include performing root-cause analysis on infrastructure issues and developing tools for Redfish, IPMI, OpenBMC, and RoCE technologies. The role requires expertise in Python, Linux, and large-scale data analysis to ensure reliability across distributed computing environments. This position solves critical quality risks and ensures the scalability of high-performance computing systems by establishing robust validation methodologies and technical leadership across multiple engineering organizations.

What you'll do

  • Lead system-level, rack-level, and cluster-scale validation strategies for AI server and accelerated computing platforms.
  • Design and maintain automated validation frameworks, telemetry pipelines, and data-driven workflows to qualify AI products.
  • Develop validation methodologies covering CPUs, GPUs, memory, PCIe fabrics, firmware, networking, and storage components.
  • Execute large-scale validation workloads across lab and manufacturing environments to evaluate performance and reliability.
  • Lead root-cause analysis for complex hardware, firmware, software, and networking issues using telemetry-driven data.
  • Develop tools to collect and analyze health metrics from Redfish, IPMI, OpenBMC, and other management interfaces.
  • Establish validation readiness criteria, quality gates, and coverage metrics to ensure product deployment readiness.
  • Mentor engineers and provide technical leadership to improve debugging effectiveness and automation capabilities.

What we're looking for

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering, or a related technical discipline.
  • Experience validating large-scale AI, cloud, HPC, server, rack-scale, or data center platforms.
  • Expertise in hardware-software integration, platform validation, system bring-up, and infrastructure qualification.
  • Proficiency in developing software and automation using Python, Linux, and modern test orchestration frameworks.
  • Ability to investigate and resolve complex issues across hardware, firmware, operating systems, networking, storage, and distributed infrastructure.
  • Experience leveraging telemetry, observability, and large-scale data analysis for root-cause analysis and engineering decision-making.
  • Familiarity with server technologies including PCIe architecture, BIOS/UEFI, BMCs (Redfish, IPMI), and high-speed interconnects like InfiniBand or RoCE.
  • Demonstrated ability to lead cross-functional technical initiatives and influence engineering decisions across organizational boundaries.

More like this

Similar roles

AI Systems Validation Engineer

Amd

Secaucus, NJ 51 days ago $136,320$204,480
AI Machine Learning GPU HPC Python Linux Bash PowerShell Firmware BIOS BMC Networking Storage System-level Validation Root Cause Analysis Telemetry Automation
8+ yrs exp

AI Validation and Test Engineer

Amd

Secaucus, NJ 48 days ago $136,320$204,480
AI Machine Learning GPU Python Linux Bash PowerShell Firmware BIOS BMC Networking Storage Data Center HPC Root Cause Analysis Automation Telemetry System-level Validation
8+ yrs exp

AI Switch Systems Design Engineer

Amd

Secaucus, NJ +1 8 days ago $64,080$96,120
Ethernet SerDes Python Broadcom SDK Redfish API IPMI BMC OpenBMC FEC PAM4 LinkCAT Signal Integrity PRBS Test Automation Hardware Bring-up Firmware PCB Layout Oscilloscope Protocol Analyzer

Senior Validation Engineer

Cardinal Health

Indianapolis, IN 14 days ago $68,500$107,580
cGMP CSV IQ OQ PQ FAT SAT URS FRS FDA (21 CFR Parts 210, 211, 820 ISO Standards Statistical Analysis Project Management Manufacturing Engineering
2+ yrs exp

Senior Server Validation Engineer

Amd

Austin, TX 57 days ago $163,200$244,800
Post-Silicon Validation System Architecture Firmware SoC C++ Python Perl Ruby Automation Frameworks Logic Analyzers Oscilloscopes Scripting Validation Strategy
Hybrid