Senior Manager, Software Engineering - RL Post-Training Frameworks

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Salary
$272,000–$431,250 / yr
Posted
17 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $199k
This role $352k
$94k most similar roles pay here $467k

This role pays more than 99% of similar roles. Most pay $173,537–$223,750 — the shaded band above. At the midpoint, this role pays about $352k versus about $199k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Manager, Software Engineering - RL Post-Training Frameworks

As a Senior Manager, Software Engineering - RL Post-Training Frameworks, you will join the RL Frameworks engineering team to lead the strategy and development of open-source tools and infrastructure for reinforcement learning. You will manage distributed teams to build, extend, and harden frameworks such as VeRL, Miles, Slime, SkyRL, and TorchTitan while integrating with systems like Megatron-Core, Ray, Monarch, NIXL, SGLang, and Kubernetes. Your role involves defining investment strategies, evaluating architecture performance across training and inference, and establishing benchmarking criteria for distributed runtimes. You will oversee the technical roadmap to ensure RL workloads scale reliably across GPUs and networking layers. Key skills include expertise in distributed systems, high-performance computing, and open-source collaboration. You will also be responsible for recruiting, coaching engineers, and translating complex partner signals into actionable engineering priorities within the reinforcement learning ecosystem.

What does a Software Engineering Manager earn?

Median $223750 from 106 postings across 38 companies.

See salary data

What you'll do

  • Define and own the strategic roadmap for NVIDIA's RL post-training frameworks and investment priorities.
  • Evaluate architecture and performance claims across training, inference, rollout, and orchestration systems.
  • Translate technical requirements from customers and partners into concrete engineering execution plans.
  • Recruit, develop, and manage a high-performing team of managers and senior individual contributors.
  • Establish operational models and decision rights for distributed teams across different geographic regions.
  • Drive integration efforts with open-source frameworks and distributed runtimes to improve RL framework quality.
  • Coach engineers to contribute effectively to open-source ecosystems while maintaining NVIDIA's technical standards.
  • Establish benchmarking, reproducibility criteria, and success metrics for large-scale RL workloads on NVIDIA platforms.

What we're looking for

  • MS or PhD in Computer Science, Computer Engineering, or a related field (or equivalent experience).
  • 10+ years of software engineering experience in distributed systems, AI frameworks, ML infrastructure, high-performance computing, or systems software.
  • 4+ years of experience as an engineering manager for software teams.
  • Strong technical background in distributed AI systems including training, inference, orchestration, and end-to-end performance.
  • Experience defining domain-level technical strategy, making investment decisions, and creating multi-team execution plans.
  • Ability to drive work across organizational boundaries and influence without direct authority.
  • Experience hiring and leading engineering teams and developing technical leaders or new managers.
  • Experience collaborating with open-source communities, research teams, external partners, or customer-facing teams.
  • Hands-on experience with RL post-training frameworks or algorithms (preferred).
  • Background with runtime and orchestration systems such as Ray, Monarch, Kubernetes, Slurm, or comparable systems (preferred).
  • Experience scaling workloads across thousands of GPUs or heterogeneous systems (preferred).
  • Familiarity with NVIDIA platform components such as CUDA, NCCL, cuDNN, TensorRT-LLM, Transformer Engine, Nsight, NeMo, or Megatron-Core (preferred).

More like this

Similar roles

Senior Software Engineer, RL Post-Training Frameworks

Nvidia

Remote (Santa Clara, CA) 21 days ago $184,000$287,500
Reinforcement Learning Python C/C++ PyTorch Kubernetes Ray vLLM SGLang TensorRT-LLM DeepSpeed Megatron-LM NCCL InfiniBand FSDP Distributed Systems High-Performance Computing
5+ yrs exp Remote

Senior Software Engineer

The Walt Disney Company

Remote 16 days ago $148,700$199,400
JavaScript HLS DASH Web Development PlayReady Widevine AVC HEVC AAC EAC3 cross platform development Android Gaming Consoles
5+ yrs exp Remote

Senior Software Engineer

Microsoft

Redmond, WA 113 days ago $119,800$234,700
Python C# Java JavaScript C++ C Azure LLMs GitHub Copilot GenAI CI/CD Infrastructure as Code Distributed Systems Telemetry A/B Testing Feature Flagging Security-as-Code WCAG
4+ yrs exp

Senior Software Engineer

Coinbase

Remote 90 days ago
React TypeScript Go API Design Distributed Systems AI/ML Observability Debugging Incident Response
5+ yrs exp Remote