ML Systems Research Engineer, RL / Inference / Agent Systems

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$204,000–$306,000 / yr
Posted
44 days ago
Freshness
Confirmed live yesterday
Closes
Jul 28, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $226k
This role $255k
$159k most similar roles pay here $322k

This role pays more than 83% of similar roles. Most pay $196,562–$254,750 — the shaded band above. At the midpoint, this role pays about $255k versus about $226k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · ML Systems Research Engineer, RL / Inference / Agent Systems

As an ML Systems Research Engineer, RL / Inference / Agent Systems, you will join the team to build reinforcement learning, inference, and evaluation infrastructure for AI-for-engineering systems. You will develop scalable systems that enable agents and models to improve engineering workflows by managing long-latency rewards, job orchestration, sampling, caching, and reproducible evaluation. Your daily work involves building tool-use pipelines for LLM agents interacting with compilers, profilers, and simulators while designing reward shaping and proxy graders. The role requires proficiency in Python and frameworks like PyTorch, JAX, or TensorFlow. You will navigate complex technical challenges involving distributed systems, Kubernetes, Ray, Slurm, and GPU technologies such as ROCm/HIP and CUDA. This position focuses on the specific domain of automating hardware engineering tasks, including verification, simulation, and compiler optimization to make research practical for production teams.

What you'll do

  • Build RL and inference systems for agentic engineering workflows including job orchestration, sampling, scoring, and caching.
  • Develop infrastructure to handle long-horizon and high-latency reward tasks where validation takes significant time.
  • Design staged rewards, proxy graders, and uncertainty-aware evaluation methods for complex engineering domains.
  • Create scalable inference and tool-use pipelines for LLM agents interacting with compilers, profilers, and simulators.
  • Support optimization workflows through candidate generation, benchmark execution, and reward modeling.
  • Standardize datasets, evaluation definitions, leaderboards, and failure taxonomies to facilitate future training.
  • Analyze experimental results to provide actionable guidance for model, agent, tool, and reward improvements.

What we're looking for

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or a related field.
  • Master's degree preferred; PhD is considered a plus.
  • Strong programming skills in Python and experience with ML frameworks like PyTorch, JAX, or TensorFlow.
  • Experience building ML systems, RL infrastructure, inference services, agent frameworks, or distributed experimentation systems.
  • Experience with reinforcement learning, RLHF, GRPO, reward modeling, or post-training systems.
  • Experience with LLM agents, tool-use systems, code generation, or compiler optimization.
  • Experience with distributed systems, job orchestration, Kubernetes, Ray, Slurm, or large-scale experiment management.
  • Familiarity with GPU systems, ROCm/HIP, CUDA, profiling, and model serving.

More like this

Similar roles

Machine Learning Engineer, Agentic AI

Apple Inc

Sunnyvale, CA 125 days ago $150,400$277,600
LLMs Agentic Systems Machine Learning Computer Vision Data Pipelines System Design Monitoring Observability

Senior AI/ML Platform Engineer

Amd

Santa Clara, CA 51 days ago $204,000$306,000
Python C++ Go Rust Kubernetes Ray Slurm vLLM SGLang Triton PyTorch JAX ROCm HIP CUDA MLflow Weights & Biases CI/CD Distributed Systems GPU Infrastructure
Hybrid

AI Engineer, Recursive Self-Improvement for Compute

Amd

Santa Clara, CA 45 days ago $204,000$306,000
Python C++ CUDA HIP ROCm Triton PyTorch JAX TensorFlow Reinforcement Learning GPU Kernels Compiler Optimization Performance Engineering Profiling Distributed Training Agentic Workflows Hardware-aware Optimization
Hybrid

Senior Forward Deployed AI Engineer

Amd

Santa Clara, CA 51 days ago $204,000$306,000
Python C++ Rust C TypeScript CUDA HIP LLM Reinforcement Learning RLHF PPO DPO GRPO PyTorch Hugging Face JAX TensorFlow Ray vLLM Distributed Systems
Hybrid

Software Engineer, GPU AI ML

Amd

Santa Clara, CA 77 days ago $204,000$306,000
C++ HIP CUDA ROCm PyTorch TensorFlow JAX GPU Architecture Kernel Optimization Distributed Systems LLMs SFT RLHF GRPO Quantization Verilog SystemVerilog RTL Design
Hybrid