AI Systems Engineer, HPC

Amd

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
San Jose, CA
Salary
$173,600–$260,400 / yr
Posted
24 days ago
Freshness
Confirmed live yesterday
Closes
Aug 17, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $190k
This role $217k
$132k most similar roles pay here $274k

This role pays more than 70% of similar roles. Most pay $150,874–$229,325 — the shaded band above. At the midpoint, this role pays about $217k versus about $190k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · AI Systems Engineer, HPC

The AI Systems Engineer joins the IT compute platforms engineering team to design, develop, and administer High-Performance Computing infrastructure, GPU clusters, and AI workload schedulers. The role involves building scalable HPC, AI, and data services while managing deployments, resource allocation, monitoring, and security for distributed ML services, LLMs, and AI inferencing. Key responsibilities include automating system provisioning, collaborating with cross-functional teams to meet infrastructure requirements, and using AI/ML to improve internal tools. The candidate will work with technologies including Python, Shell, SLURM, Kubernetes, KVM, Ubuntu, RoCEv2, and 400G networking. Additionally, the role requires proficiency in GPU drivers, Ansible, Saltstack, Terraform, Prometheus, and Grafana. This position focuses on solving technical challenges related to large-scale distributed computing for AI and HPC workloads within a high-performance environment.

What you'll do

  • Develop, implement, and maintain GPU-based clusters to ensure optimal performance.
  • Administer ML/AI platforms including distributed services, LLMs, and inference systems.
  • Automate end-to-end system provisioning and cluster management processes.
  • Monitor and evaluate the performance of AI systems against industry best practices.
  • Manage resource allocation, monitoring, and security for high-performance computing infrastructure.
  • Use AI/ML to improve internal tools and processes for service delivery.
  • Optimize GPU-based services and software using advanced networking and scheduling technologies.

What we're looking for

  • A Bachelor's or Master's degree in computer science or computer engineering is preferred.
  • Experience in HPC infrastructure engineering specifically for the AI/HPC domain.
  • Experience managing GPU clusters and optimizing GPU-based services, tools, and software.
  • Proficiency in SLURM and Kubernetes management.
  • Proficiency in RoCEv2, KVM, Ubuntu, Python, Shell, GPU drivers, and 400G networking cluster interconnects.
  • Experience with automation and monitoring tools including Ansible, Saltstack, Terraform, Prometheus, and Grafana.
  • Experience developing Python-based AI applications and user interfaces.
  • Demonstrated experience with AI workload schedulers and allocation optimization.

More like this

Similar roles

HPC Systems Engineer, AI Workloads

Amd

San Jose, CA 2 days ago $173,600$260,400
GPU HPC Kubernetes Python SLURM RoCEv2 KVM Ubuntu Shell Ansible Saltstack Terraform Prometheus Grafana Distributed ML LLMs 400G Networking

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote

AI/HPC Cluster Design Engineer

Amd

Austin, TX 53 days ago $133,200$199,800
HPC AI Systems GPU CPU InfiniBand Ethernet Lustre Ceph PCIe UALink Cluster Design Data Center Engineering Power Delivery System Architecture

System Software Engineer, HPC Performance

Nvidia

Champaign, IL +2 17 days ago $152,000$241,500
C C++ Python HPC Machine Learning Deep Learning Artificial Intelligence x86 ARM Linux Windows macOS Profiling Tools Cloud Computing
5+ yrs exp