AI/HPC Cluster Design Engineer

Amd

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Austin, TX
Salary
$133,200–$199,800 / yr
Posted
53 days ago
Freshness
Confirmed live yesterday
Closes
Jul 20, 2027

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $199k
This role $166k
$121k most similar roles pay here $248k

This role pays less than 72% of similar roles. Most pay $162,000–$235,750 — the shaded band above. At the midpoint, this role pays about $166k versus about $199k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · AI/HPC Cluster Design Engineer

As an AI/HPC Cluster Design Engineer, you will join a team focused on designing scalable AI and high-performance computing clusters that meet specific customer and design requirements. Your daily responsibilities involve reviewing and selecting critical compute, storage, networking, and power delivery components to optimize performance and reliability across global deployments. You will collaborate with cross-functional teams including hardware, software, network, data center, and operations experts to deliver infrastructure for AI workloads. The role requires deep technical knowledge of GPU/CPU architectures, PCIe, UALink, InfiniBand, and Ethernet networking. You will also design network topologies and storage solutions using technologies like Lustre and Ceph. To succeed, you must understand AI/ML framework characteristics and rack designs while solving complex problems related to the performance trade-offs inherent in large-scale cluster infrastructure for high-performance computing environments.

What you'll do

  • Design scalable AI and HPC clusters including compute, storage, and networking components.
  • Select CPUs, GPUs, accelerators, interconnects, and memory configurations to optimize cluster performance.
  • Design network topologies tailored to the specific performance needs of AI and HPC workloads.
  • Develop and optimize storage solutions such as Lustre or Ceph for high-performance computing.
  • Ensure all system designs comply with specific customer requirements and technical specifications.
  • Design power delivery solutions for racks and data center infrastructure.
  • Document technical specifications and provide clear communication across hardware, software, and operations teams.

What we're looking for

  • Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field.
  • Experience in HPC, AI systems/clusters, or data center engineering.
  • Knowledge of GPU/CPU architectures, PCIe, UALink, InfiniBand, and Ethernet networking.
  • Expertise in designing scalable AI/HPC clusters including compute, storage, and networking components.
  • Familiarity with AI/ML frameworks and workload characteristics.
  • Experience designing power delivery solutions for racks and data centers.
  • Knowledge of rack and cluster design.
  • Strong problem-solving, communication, and documentation skills.

More like this

Similar roles

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote

AI Systems Engineer, HPC

Amd

San Jose, CA 24 days ago $173,600$260,400
GPU HPC Kubernetes Python SLURM RoCEv2 KVM Ubuntu Shell Ansible Saltstack Terraform Prometheus Grafana Distributed ML LLMs 400G Networking

HPC Systems Engineer, AI Workloads

Amd

San Jose, CA 2 days ago $173,600$260,400
GPU HPC Kubernetes Python SLURM RoCEv2 KVM Ubuntu Shell Ansible Saltstack Terraform Prometheus Grafana Distributed ML LLMs 400G Networking

HPC Infrastructure & Cluster Engineer

General Dynamics

Springfield, VA 8 days ago $119,850$162,150
Linux Run:AI OpenShift Kubernetes InfiniBand SLURM Python Bash SAN HPC Bare-metal Parallel File Systems Cluster Administration Infrastructure Optimization Network Management Storage Area Network Container Orchestration
5+ yrs exp

AI / HPC Data Center Lab Engineer

Amd

Austin, TX 90 days ago $112,640$168,960
Python C/C++ Bash Linux Windows CI/CD Jenkins Ansible Perl Tcl LabView I2C SPI PCIe Gen 5 Redfish IPMI Oscilloscopes multi-meters Power Electronics Firmware BIOS

HPC Performance Engineer

Nvidia

Remote (OR) +3 53 days ago $152,000$241,500
HPC CUDA C++ C Fortran OpenMP MPI OpenACC Compiler Optimization Assembly Linear Algebra Numerical Methods Multi-GPU Systems
5+ yrs exp Remote