HPC Systems Engineer, AI Workloads

Amd

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
San Jose, CA
Salary
$173,600–$260,400 / yr
Posted
2 days ago
Freshness
Confirmed live yesterday
Closes
Sep 8, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $191k
This role $217k
$133k most similar roles pay here $274k

This role pays more than 70% of similar roles. Most pay $151,356–$231,362 — the shaded band above. At the midpoint, this role pays about $217k versus about $191k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · HPC Systems Engineer, AI Workloads

The HPC Systems Engineer - AI Workloads joins the IT compute platforms engineering team to design, develop, and administer high-performance computing infrastructure, GPU clusters, and AI workload schedulers. This role involves building scalable and performant HPC, AI, and data services while managing deployments, resource allocation, monitoring, and security for distributed ML services, LLMs, and AI inferencing. The engineer will automate system provisioning, manage cluster management end to end, and collaborate with cross-functional teams to meet infrastructure requirements. Key technical skills include Python, Shell, Ubuntu, and GPU drivers, alongside experience with SLURM, Kubernetes, RoCEv2, KVM, and 400G networking. The role also requires proficiency in automation and monitoring tools such as Ansible, Saltstack, Terraform, Prometheus, and Grafana to optimize performance for large-scale distributed computing within the AI and HPC domain.

What you'll do

  • Develop, implement, and maintain GPU-based clusters to ensure optimal performance.
  • Administer ML/AI platforms including distributed services, LLMs, and inference systems.
  • Automate end-to-end system provisioning and cluster management processes.
  • Monitor and evaluate AI system performance against industry best practices and company standards.
  • Manage resource allocation and security for high-performance computing infrastructure.
  • Utilize AI/ML to improve internal tools and service delivery processes.
  • Optimize GPU-based services, tools, and software within the HPC environment.

What we're looking for

  • Experience in designing, developing, and administering HPC infrastructure, GPU clusters, and AI workload schedulers.
  • Ability to manage ML/AI platforms including distributed services, LLMs, and inference monitoring.
  • Proficiency in RoCEv2, K8s, KVM, Ubuntu, Python, Shell, GPU drivers, and 400G networking (preferred).
  • Experience with automation and monitoring tools such as Ansible, Saltstack, Terraform, Prometheus, and Grafana (preferred).
  • Experience with SLURM and Kubernetes management (preferred).
  • Experience in developing Python-based AI apps, UI, and web services with HPC backends (preferred).
  • Strong organizational, problem-solving, troubleshooting, and communication skills.
  • Bachelor's or master's degree in computer science or computer engineering (preferred).

More like this

Similar roles

AI Systems Engineer, HPC

Amd

San Jose, CA 24 days ago $173,600$260,400
GPU HPC Kubernetes Python SLURM RoCEv2 KVM Ubuntu Shell Ansible Saltstack Terraform Prometheus Grafana Distributed ML LLMs 400G Networking

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote

AI/HPC Cluster Design Engineer

Amd

Austin, TX 53 days ago $133,200$199,800
HPC AI Systems GPU CPU InfiniBand Ethernet Lustre Ceph PCIe UALink Cluster Design Data Center Engineering Power Delivery System Architecture

System Software Engineer, HPC Performance

Nvidia

Champaign, IL +2 17 days ago $152,000$241,500
C C++ Python HPC Machine Learning Deep Learning Artificial Intelligence x86 ARM Linux Windows macOS Profiling Tools Cloud Computing
5+ yrs exp

AI Systems Performance Engineer

Broadcom

San Jose, CA 144 days ago $141,300$226,000
Ethernet MLPerf NCCL Python C++ Linux PyTorch RDMA RoCEv2 Docker Kubernetes CI/CD Performance Benchmarking Distributed Systems
10+ yrs exp

HPC Systems Engineer, Modeling & Simulation

Anduril Industries

Costa Mesa, CA 92 days ago $132,000$198,000
HPC Linux Unix Python Bash Slurm MPI OpenMP CMake NFS NAS TCP/IP ParaView VisIt Cubit CTH ALE3D Sierra Data Workflows
5+ yrs exp