Principal AI Cluster Performance Validation Engineer

Amd

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Austin, TXSeattle, WASanta Clara, CASecaucus, NJ
Salary
$200,000–$300,000 / yr
Posted
4 days ago
Freshness
Confirmed live today
Closes
Sep 9, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $218k
This role $250k
$150k most similar roles pay here $316k

This role pays more than 73% of similar roles. Most pay $180,462–$254,750 — the shaded band above. At the midpoint, this role pays about $250k versus about $218k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 379 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 379 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Principal AI Cluster Performance Validation Engineer

As a Principal AI Cluster Performance Validation Engineer, you will join a dynamic team focused on optimizing and achieving peak performance for GPU clusters. You will be responsible for evaluating scalability across various workloads, developing benchmarking strategies to identify bottlenecks, and performing performance profiling to provide actionable insights. Your daily work involves optimizing RDMA networks, RoCE v2 network congestion, latency, and collective communications while collaborating with cross-functional teams to integrate improvements into the architecture. The role requires expertise in GPU architectures, parallel computing, Linux kernel networking, and machine learning or HPC system design. You will utilize Python and Bash for automation and performance analysis to solve complex hardware and firmware issues. This position addresses critical challenges in data flows between GPUs, NICs, and cluster networks to ensure high-performance computing environments.

What you'll do

  • Evaluate GPU cluster scalability across various workloads and networking technologies like RoCE.
  • Develop and execute benchmarking strategies to identify performance bottlenecks in GPU environments.
  • Optimize RDMA throughput, RoCE v2 network congestion, and collective communications within the cluster.
  • Use profiling tools to analyze system performance and provide actionable insights for improvement.
  • Implement optimization strategies including protocol enhancements and load balancing techniques.
  • Create detailed documentation of performance analysis and tuning results for internal stakeholders.
  • Debug complex hardware, firmware, and clustered configurations to ensure peak performance.

What we're looking for

  • Bachelor's or Master's degree in computer science or electrical engineering.
  • Experience debugging complex hardware, firmware, and clustered configurations.
  • Expertise in RDMA networks, RoCE v2 network congestion, and collective communications.
  • Proficiency in scripting languages like Python or Bash for automation and performance analysis.
  • Strong understanding of GPU architectures, parallel computing concepts, and network protocols (preferred).
  • Experience with system level performance analysis tools and methodologies for GPU clusters (preferred).
  • Linux kernel networking expertise (preferred).
  • Machine learning and/or HPC system design experience (preferred).

More like this

Similar roles

AI Systems Validation Engineer

Amd

Secaucus, NJ 53 days ago $136,320$204,480
AI Machine Learning GPU HPC Python Linux Bash PowerShell Firmware BIOS BMC Networking Storage System-level Validation Root Cause Analysis Telemetry Automation
8+ yrs exp

AI Validation and Test Engineer

Amd

Secaucus, NJ 51 days ago $136,320$204,480
AI Machine Learning GPU Python Linux Bash PowerShell Firmware BIOS BMC Networking Storage Data Center HPC Root Cause Analysis Automation Telemetry System-level Validation
8+ yrs exp