Senior Software Engineer, RL Post-Training Frameworks

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CA
Salary
$184,000–$287,500 / yr
Posted
21 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $196k
This role $236k
$129k most similar roles pay here $305k

This role pays more than 85% of similar roles. Most pay $170,051–$221,250 — the shaded band above. At the midpoint, this role pays about $236k versus about $196k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, RL Post-Training Frameworks

As a Senior Software Engineer, RL Post-Training Frameworks, you will join the RL Frameworks engineering team to develop open-source tools and infrastructure for reinforcement learning post-training. You will architect and build scalable infrastructure for training-inference-rollout loops across GPUs, CPUs, and LPUs while contributing to frameworks like VeRL, Miles, and TorchTitan. Your daily work involves optimizing distributed runtimes such as Ray and Monarch, ensuring fault tolerance, and managing elastic scaling for long-running jobs. You will leverage Python, C/C++, and PyTorch internals to solve complex challenges in RLHF, PPO, GRPO, and DPO. The role addresses the technical hurdles of coordinating actor, critic, and reward models across heterogeneous hardware to enable models to reason through hard problems and act as autonomous agents within complex environments involving tool-use and code execution.

What does a Software Engineer earn in California?

Median $214000 from 775 postings across 63 companies.

See salary data

What you'll do

  • Architect and build RL post-training infrastructure that scales from single GPUs to thousands of nodes.
  • Optimize training-inference-rollout loops across heterogeneous hardware including GPUs, CPUs, and LPUs.
  • Contribute to and improve the performance of open-source RL frameworks like VeRL, Miles, and TorchTitan.
  • Implement fault tolerance, elastic scaling, and fast restarts for long-running distributed training jobs.
  • Develop systems for CPU-driven rollout workloads such as tool-use, code execution, and agentic environments.
  • Advocate for researcher needs to internal networking, math library, and compiler teams to prioritize RL capabilities.
  • Integrate high-performance inference engines into RL training loops for accelerated rollouts.

What we're looking for

  • MS or PhD in Computer Science, Computer Engineering, or a related field (or equivalent experience).
  • 5+ years of professional experience in distributed systems, high-performance computing, deep learning infrastructure, or ML systems engineering.
  • Strong proficiency in Python and C/C++.
  • Experience building or contributing to large-scale distributed systems or runtime frameworks at a frontier AI lab, hyperscaler, or major technology company.
  • Depth in RL for LLM post-training (RLHF, PPO, GRPO, DPO) and its associated distributed execution challenges.
  • Expertise in PyTorch internals, including distributed training primitives like FSDP, tensor parallelism, and pipeline parallelism.
  • Knowledge of Kubernetes runtime internals, including container lifecycle, pod scheduling, and GPU allocation.
  • Proficiency in end-to-end distributed systems design, including service boundaries, data flows, consistency models, and failure recovery.

More like this

Similar roles

Senior Research Engineer, Autonomous Vehicles

Nvidia

Santa Clara, CA 37 days ago $184,000$287,500
PyTorch JAX TensorFlow Python C++ CUDA Kubernetes SLURM Deep Learning Reinforcement Learning LLMs MLOps HPC Distributed Training Sim-to-Real Natural Language Processing Graphics
10+ yrs exp

Senior Software Engineer

Genworth Financial

Richmond, VA +1 61 days ago $96,700$145,000
Python Flask Azure PostgreSQL GitLab CI/CD DevSecOps JavaScript HTML CSS React Angular Vue Azure App Service Azure Synapse Pipelines SQL
5+ yrs exp Hybrid

Senior Software Engineer

Datadog

96 days ago $192,000$240,000
Go Python C Redis Cassandra Kafka Data Pipelines AI Backend Systems Feature Flags Statistics
6+ yrs exp Hybrid

Senior Software Engineer

Morningstar Inc

Chicago, IL 67 days ago
Python FastAPI Flask PostgreSQL SQL Server Weaviate Pinecone AWS S3 Aurora RDS Lambda Redis Celery Git CI/CD Pandas Plotly Matplotlib QlikView GraphQL LangGraph AutoGen OpenAI Anthropic MistralAI Transformers
5+ yrs exp Hybrid

Senior Software Engineer

Microsoft

Redmond, WA 58 days ago $119,800$234,700
C++ Win32 COM Windows Services WinRT RPC C C# ETW IPC RDP PCoIP ICA Kernel Driver Development Threading System APIs Telemetry Agile
4+ yrs exp