PhD AI Training Systems and Performance Engineer Intern/Co-Op

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Jose, CASanta Clara, CA
Salary
$91,520–$137,280 / yr
Employment
Intern
Posted
6 days ago
Freshness
Confirmed live yesterday
Closes
Oct 5, 2027

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $166k
This role $114k
$76k $241k
below market most similar roles pay here above market

This role pays less than 71% of similar roles. Most pay $114,400–$218,500 — the blue band above. At the midpoint, this role pays about $114k versus about $166k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 510 open roles on FindRole.

Listed pay typically runs $172,000–$258,000 across 510 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · PhD AI Training Systems and Performance Engineer Intern/Co-Op

The 2027 PhD AI Training Systems and Performance Engineer Intern/Co-Op joins a team of software engineers, architects, and AI specialists to accelerate the adoption and optimization of cutting-edge AI training workloads on AMD Instinct GPUs. The intern will bring up new training workloads, analyze performance bottlenecks, and develop innovative tooling to improve scalability, efficiency, and developer productivity. Key responsibilities include optimizing large-scale training and fine-tuning workloads, profiling bottlenecks across GPUs, CPUs, and networking, and implementing optimization strategies for distributed multi-GPU environments. Required skills include Python and C++ programming, experience with PyTorch, JAX, TensorFlow, vLLM, or SGLang, and knowledge of distributed training, GPU performance optimization, and transformer-based architectures. The role focuses on improving training throughput, GPU utilization, and memory efficiency for foundation models and large language models.

What you'll do

  • Develop and optimize large-scale AI training and fine-tuning workloads on AMD GPU platforms.
  • Bring up newly released foundation models and training frameworks on AMD hardware.
  • Profile and analyze AI workloads to identify bottlenecks across GPUs, CPUs, memory, and networking.
  • Develop tools and workflows to automate training setup and debugging using LLM-powered agents.
  • Implement optimization strategies to improve training throughput, GPU utilization, and memory efficiency.
  • Benchmark and tune AI frameworks, libraries, and SDKs using advanced profiling tools.
  • Transform successful experiments into reusable workflows and software improvements for out-of-the-box performance.

What we're looking for

  • Currently pursuing a PhD in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, Machine Learning, or a related technical discipline.
  • Strong programming experience in Python and/or C++.
  • Hands-on experience implementing, training, and debugging deep learning models using frameworks such as PyTorch, JAX, TensorFlow, vLLM, or SGLang.
  • Experience with distributed training systems, GPU performance optimization, HPC, or large language model training and fine-tuning.
  • Understanding of transformer-based architectures, mixture-of-experts models, and modern LLM training techniques.
  • Experience profiling workloads using tools such as PyTorch Profiler, ROCm Profiler, VTune, Nsight, or similar tools.
  • Familiarity with distributed training technologies and communication libraries such as MPI, NCCL/RCCL, OpenMP, or related frameworks.
  • Experience identifying and resolving compute, memory, data-loading, or communication bottlenecks in large-scale AI workloads (preferred).

More like this

Similar roles

PhD AI Systems & GPU Performance Engineering Intern/Co-op

Amd

San Jose, CA +1 30 days ago $91,520–$137,280
Python C/C++ PyTorch JAX TensorFlow Triton ROCm HIP CUDA vLLM Linux Quantization Mixed Precision Operator Fusion Distributed Training rocProfiler Omniperf Nsight GPU Architecture Performance Engineering

PhD HPC & Sovereign AI Center of Excellence Intern/Co-op

Amd

Austin, TX 20 days ago $91,520–$137,280
High-Performance Computing Artificial Intelligence Python C/C++ Fortran PyTorch TensorFlow JAX ROCm CUDA MPI OpenMP SYCL GPU Programming Distributed Systems Parallel Computing Numerical Methods Performance Analysis Jira

PhD ML System Engineering Intern/Co-op

Amd

San Jose, CA +1 30 days ago $91,520–$137,280
LLM VLM Python C++ MLIR IRON AIE GEMM Transformer Graph Lowering Operator Fusion Tiling Memory Planning Profiling AI Compilers

PhD Optical & Photonics Engineering Intern/Co-Op

Amd

San Jose, CA +1 30 days ago $91,520–$137,280
Silicon Photonics Photonic Integrated Circuits (PICs) Optical Communications SerDes Signal Integrity Analog and Mixed-Signal Circuit Design Python MATLAB Optical Transceivers Fiber-optic Networking Oscilloscopes Network Analyzers Optical Spectrum Analyzers BER Testers Data Analysis Engineering Simulation Semiconductor Technologies