AI Research Scientist, Recursive Self Improvement, AI Safety and Reinforcement Learning

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$204,000–$306,000 / yr
Posted
71 days ago
Freshness
Confirmed live yesterday
Closes
Jul 2, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $222k
This role $255k
$165k most similar roles pay here $321k

This role pays more than 75% of similar roles. Most pay $187,390–$256,237 — the shaded band above. At the midpoint, this role pays about $255k versus about $222k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · AI Research Scientist, Recursive Self Improvement, AI Safety and Reinforcement Learning

As an AI Research Scientist, Recursive Self Improvement, AI Safety and Reinforcement Learning, you will join the team to research recursive self-improvement in a bounded, engineering-first context. You will investigate systems where models, data generators, or toolchains improve their own training signals, curricula, or verification under explicit governance and human oversight. Your daily work involves researching self-improving training loops like model-generated supervision and iterative distillation while designing measurement and containment for RSI pipelines to prevent bias or reward hacking. You will develop evaluations for capability drift and Goodhart effects, partner with RL scientists on policy optimization, and define red-team protocols for monitoring. The role requires expertise in machine learning, AI safety, reinforcement learning, and software skills to build experimental harnesses. This work addresses the technical challenges of ensuring auditable, safe training for hardware and generative-AI programs.

What you'll do

  • Research self-improving training loops including model-generated supervision, iterative distillation, and automated curriculum updates.
  • Develop systems-grounded evaluations to detect capability drift, Goodhart effects, and distributional shifts in closed-loop training.
  • Collaborate with RL scientists to align recursive self-improvement objectives with policy optimization and preference learning.
  • Design red-team protocols and monitoring systems for recursive self-improvement pilots.
  • Establish rollback criteria and safety guardrails before experiments interact with shared infrastructure.
  • Produce technical reports and publications that align internal research with responsible deployment standards.
  • Build controlled experimental harnesses and reproducible environments to test training pipelines.

What we're looking for

  • A PhD in Computer Science, Machine Learning, or a related field is strongly preferred.
  • Strong background in machine learning, AI safety, or reinforcement learning.
  • Experience with iterative training, self-training, or open-ended learning.
  • Experience with empirical safety evaluation, scalable oversight, or stress-testing of generative model training pipelines.
  • Strong software skills for building controlled experimental harnesses and reproducible RSI microcosms.

More like this

Similar roles

Director, AI Research - Recursive Self-Improvement

Amd

Santa Clara, CA 37 days ago $246,400$369,600
Reinforcement Learning AI Machine Learning Kernel Optimization Compiler Optimization Hardware Design Code Generation Design-Space Exploration Research Prototyping Training Pipelines Reward Modeling silicon, architecture
Hybrid

AI Research Scientist, Hardware AI Systems

Amd

Santa Clara, CA 71 days ago $204,000$306,000
Machine Learning Reinforcement Learning Large Language Models Transfer Learning Meta-learning RTL Verification Physical Design Multi-task Learning Reward Shaping Representation Learning Hardware Engineering Silicon Signoff
Hybrid

Staff Gen AI Research Scientist

Anduril Industries

Costa Mesa, CA +2 23 days ago $220,000$292,000
Python PyTorch JAX LLMs Generative AI Computer Vision NLP Robotics SLAM Quantization Distillation Pruning Axolotl Hugging Face DeepSpeed Megatron-LM LangChain LlamaIndex RLHF DPO
2+ yrs exp

Director, AI Research

Amd

Santa Clara, CA 37 days ago $246,400$369,600
AI Machine Learning Reinforcement Learning Model Training AI-for-systems hardware–software stack Compiler Inference Research Prototyping Training Pipelines Performance Optimization
Hybrid