Senior System Software Engineer, Agentic Inference

Nvidia

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$224,000–$356,500 / yr
Posted
46 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $216k
This role $290k
$163k most similar roles pay here $377k

This role pays more than 93% of similar roles. Most pay $185,600–$246,150 — the shaded band above. At the midpoint, this role pays about $290k versus about $216k for comparable roles.

Based on 239 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior System Software Engineer, Agentic Inference

Senior System Software Engineer, Agentic Inference - Dynamo joins the GPU-accelerated deep learning software team to develop open source software for serving inference of trained AI models on GPUs. The role focuses on building disaggregated serving for Dynamo-supported engines like vLLM, SGLang, and TRT-LLM while expanding capabilities for agentic inference workloads including long-horizon reasoning, tool calling, and stateful multi-turn execution. Key responsibilities include innovating in inference-state management using NIXL to optimize KV-cache reuse across memory hierarchies, developing distributed inference frontends, and balancing load across resources to improve throughput and latency. The position requires expert skills in Rust and Python programming, performance analysis, and test design. Candidates must understand modern LLM API semantics, including structured outputs and context management, while addressing technical challenges like bursty tool-call traffic and multi-turn state reuse for self-hosted LLMs.

What does a System Software Engineer earn in California?

Median $235750 from 37 postings across 4 companies.

See salary data

What you'll do

  • Develop open source software to serve inference of trained AI models running on GPUs.
  • Contribute to disaggregated serving for Dynamo-supported engines like vLLM, SGLang, and TRT-LLM.
  • Expand capabilities to support agentic inference workloads including long-horizon reasoning and tool calling.
  • Innovate in inference-state management using KV-cache and prefix-cache reuse across heterogeneous memory hierarchies.
  • Build and evolve Dynamo’s distributed inference frontend for new models and upstream API compatibility.
  • Load-balance asynchronous requests across available resources to optimize throughput under latency constraints.
  • Integrate the latest open source technologies into the generative AI inference platform.

What we're looking for

  • Master's degree, PhD, or equivalent experience in Computer Science, Computer Engineering, or a related field.
  • 10+ years of experience in Computer Science, Computer Engineering, or a related field.
  • Proficiency in Rust and Python programming for software design, debugging, performance analysis, and test design.
  • Understanding of modern LLM API semantics including structured outputs, tool calling, reasoning controls, and context management.
  • Experience contributing to open-source AI inference frameworks such as vLLM, TensorRT-LLM, or SGLang.
  • Experience optimizing GPU memory, KV/prefix caches, or high-performance networking for long-context workloads.
  • Understanding of LLM-specific inference challenges for agentic workloads and multi-turn state reuse.
  • Experience integrating self-hosted LLM serving stacks with agent harnesses like OpenCode, Codex, Claude Code, or Pi.

More like this

Similar roles

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 136 days ago $184,000$287,500
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (CA) +4 15 days ago $152,000$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization Profiling Agile
5+ yrs exp Remote

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp

Software Engineering Intern

Nvidia

Santa Clara, CA 37 days ago
Rust Python Kubernetes vLLM SGLang TensorRT-LLM llama.cpp mistral.rs RESTful APIs Machine Learning NLP GitHub Algorithms Data Structures

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid