AI Research Scientist, Infrastructure Engineer, Reinforcement Learning

Amd

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$204,000–$306,000 / yr
Posted
71 days ago
Freshness
Confirmed live 2 days ago
Closes
Jul 2, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $220k
This role $255k
$159k most similar roles pay here $322k

This role pays more than 84% of similar roles. Most pay $186,200–$254,750 — the shaded band above. At the midpoint, this role pays about $255k versus about $220k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · AI Research Scientist, Infrastructure Engineer, Reinforcement Learning

As an AI Research Scientist - Infrastructure Engineer, Reinforcement Learning, you will own reinforcement learning infrastructure at scale across large GPU fleets. You will make research scientists productive by building reliable systems for distributed policy and value training, rollout generation, logging, checkpointing, and researcher-facing APIs. Your daily work involves designing and implementing distributed training stacks using data or pipeline parallelism integrated with schedulers and storage. You will build high-throughput rollout workers, trajectory stores, and reward computation pipelines while instrumenting jobs to handle debugging, autoscaling, and preemption-safe checkpointing. The role requires expertise in PyTorch or JAX, NCCL/MPI-style distributed training, GPU cluster orchestration, and C++/Python performance tuning. You will solve technical challenges related to RL training infrastructure, LLM post-training pipelines, and large-scale experiment management to improve throughput, fault tolerance, and observability for research programs.

What you'll do

  • Design and implement distributed RL training stacks using data/pipeline parallelism integrated with AMD schedulers.
  • Build high-throughput rollout workers, trajectory stores, and reward computation pipelines with versioning.
  • Instrument jobs to detect and debug issues like NaNs, stragglers, and Out-of-Memory errors.
  • Implement autoscaling and preemption-safe checkpointing for large GPU fleets.
  • Develop researcher-facing APIs and experiment templates to improve scientist productivity.
  • Manage infrastructure reliability through on-call rotations, runbooks, and postmortems for training incidents.
  • Optimize performance through C++/Python tuning and I/O optimization for containerized workloads.

What we're looking for

  • A Bachelor's degree in Computer Science is required.
  • A Master's or PhD in Computer Science is preferred.
  • Experience with PyTorch, JAX, and NCCL/MPI-style distributed training is required.
  • Expertise in GPU cluster orchestration and large-scale experiment management is required.
  • Proficiency in C++ and Python performance tuning and I/O optimization is required.
  • Experience building high-throughput rollout workers and reward computation pipelines is preferred.
  • Experience with LLM post-training pipelines or RL training infrastructure is preferred.
  • Demonstrated track record in machine learning platforms with deep systems expertise is preferred.

More like this

Similar roles

ML Systems Research Engineer, RL / Inference / Agent Systems

Amd

Santa Clara, CA 44 days ago $204,000$306,000
Python PyTorch JAX TensorFlow Reinforcement Learning RLHF LLM Agents Kubernetes Ray Slurm CUDA ROCm HIP Distributed Systems Data Pipelines Model Serving Compiler Optimization Profiling Hardware Engineering
Hybrid

AI Infrastructure Engineer

Fortinet

New York, NY 14 days ago $215,000$350,000
Linux GPU Docker Kubernetes Python Bash CI/CD KVM FortiGate FortiManager FortiAnalyzer Monitoring Logging Alerting Networking Virtualization Performance Testing Benchmarking Automation

Reinforcement Learning AI Engineer

Booz Allen Hamilton

Huntsville, AL +3 50 days ago $99,000$225,000
Reinforcement Learning Multi-Agent Reinforcement Learning Python PyTorch TensorFlow JAX C++ Rust Gym PettingZoo CUDA Kubernetes Containerization Distributed Training Data Science Simulation Environments