AI Research Scientist, Reinforcement Learning (LLM) and Post-Training

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$204,000–$306,000 / yr
Posted
9 days ago
Freshness
Confirmed live yesterday
Closes
Sep 1, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $227k
This role $255k
$170k most similar roles pay here $321k

This role pays more than 76% of similar roles. Most pay $198,656–$255,000 — the shaded band above. At the midpoint, this role pays about $255k versus about $227k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · AI Research Scientist, Reinforcement Learning (LLM) and Post-Training

As an AI Research Scientist, Reinforcement Learning (LLM) and Post-Training, you will join the team to advance post-training and interactive learning for large generative models. You will focus on engineering and hardware-adjacent tasks such as code, optimization, tool use, and long-horizon decision making. Your daily responsibilities include inventing and analyzing RL algorithms like policy optimization, reward modeling, and preference-based methods while conducting rigorous empirical studies to improve task success. You will design reward models and training recipes for sparse or noisy labels while collaborating with infrastructure engineers to scale training. The role requires expertise in RL theory, RLHF/RLAIF, and policy optimization for language or code agents. You will utilize GPUs and distributed jobs to develop methods for large-scale systems, potentially involving compilers, kernels, or EDA-style workflows within the domain of advanced machine learning research.

What you'll do

  • Research and develop RL methods for post-training LLMs and code models on structured engineering tasks.
  • Design reward models, training curricula, and optimization recipes for sparse or noisy data.
  • Identify and mitigate failure modes such as reward hacking, degenerate policies, and training instability.
  • Partner with infrastructure engineers to scale training and define interfaces for rollout generation and logging.
  • Publish original research at top-tier machine learning venues like NeurIPS, ICML, and ICLR.
  • Provide technical leadership on the internal reinforcement learning roadmap.

What we're looking for

  • PhD in Computer Science, Machine Learning, or a related field is strongly preferred.
  • Proven track record of publishing research at top venues like NeurIPS, ICML, or ICLR.
  • Expertise in reinforcement learning theory and practical production-scale training.
  • Experience with LLM post-training, RLHF/RLAIF, or policy optimization for language or code agents.
  • Hands-on experience training RL or preference-optimized models at scale using GPUs and distributed jobs.
  • Ability to design reward models, curricula, and training recipes for sparse or noisy labels.
  • Experience identifying and mitigating failure modes such as reward hacking and policy instability.
  • Familiarity with compilers, kernels, EDA-style workflows, or large-scale codebases is a plus.

More like this

Similar roles

ML Systems Research Engineer, RL / Inference / Agent Systems

Amd

Santa Clara, CA 45 days ago $204,000$306,000
Python PyTorch JAX TensorFlow Reinforcement Learning RLHF LLM Agents Kubernetes Ray Slurm CUDA ROCm HIP Distributed Systems Data Pipelines Model Serving Compiler Optimization Profiling Hardware Engineering
Hybrid

Reinforcement Learning AI Engineer

Booz Allen Hamilton

Huntsville, AL +3 50 days ago $99,000$225,000
Reinforcement Learning Multi-Agent Reinforcement Learning Python PyTorch TensorFlow JAX C++ Rust Gym PettingZoo CUDA Kubernetes Containerization Distributed Training Data Science Simulation Environments

Director, AI Research - Recursive Self-Improvement

Amd

Santa Clara, CA 38 days ago $246,400$369,600
Reinforcement Learning AI Machine Learning Kernel Optimization Compiler Optimization Hardware Design Code Generation Design-Space Exploration Research Prototyping Training Pipelines Reward Modeling silicon, architecture
Hybrid