Senior System Software Engineer, Dynamo-Triton Inference Server

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CASeattle, WA
Salary
$152,000–$241,500 / yr
Posted
71 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $196k
This role $197k
$141k most similar roles pay here $252k

This role pays less than 54% of similar roles. Most pay $155,675–$235,750 — the shaded band above. At the midpoint, this role pays about $197k versus about $196k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior System Software Engineer, Dynamo-Triton Inference Server

As a Senior System Software Engineer on the GPU-accelerated deep learning software team, you will develop world-class inference serving software for the Dynamo-Triton Inference Server. You will contribute to feature development and drive the convergence of Triton and NVIDIA Dynamo stacks to create a unified platform supporting both Large Language Model and non-LLM workloads. Your daily responsibilities include building robust software for production server or cloud environments, optimizing prediction throughput and latency, and adopting next-generation inference technologies while participating in the open source community. The role requires expert skills in Rust and C++, familiarity with Python, and experience in high-scale distributed systems and machine learning systems. You will utilize tools like TensorRT, PyTorch, ONNX, vLLM, or TRT-LLM to solve complex problems involving GPU memory management, cache management, and high-performance networking for AI model deployment.

What does a System Software Engineer earn in California?

Median $235750 from 37 postings across 4 companies.

See salary data

What you'll do

  • Develop high-performance GPU-accelerated AI inference serving software.
  • Drive the convergence of Triton Inference Server and NVIDIA Dynamo stacks into a unified platform.
  • Ensure feature parity for both Large Language Model (LLM) and non-LLM workloads.
  • Optimize and balance prediction throughput and latency for production environments.
  • Develop and adopt next-generation inference technologies.
  • Contribute to and participate in the open source deep learning software community.

What we're looking for

  • MS or PhD in Computer Science or a relevant field (or equivalent experience).
  • 5+ years of professional experience working on deep learning software.
  • Excellent Rust and C++ programming skills.
  • Familiarity with Python.
  • Strong software design skills including debugging, performance analysis, and test design.
  • Experience with high-scale distributed systems and ML systems.
  • Prior experience with AI frameworks such as TensorRT, PyTorch, ONNX, OpenVINO, vLLM, or TRT-LLM.
  • Knowledge of GPU memory management, cache management, or high-performance networking.

More like this

Similar roles

Senior System Software Engineer, Agentic Inference

Nvidia

Santa Clara, CA 46 days ago $224,000$356,500
Rust Python vLLM SGLang TensorRT-LLM GPU Deep Learning Generative AI Performance Analysis NIXL Open Source Software high-performance networking
10+ yrs exp Hybrid

Senior System Software Engineer

Nvidia

Remote (Santa Clara, CA) +3 66 days ago $184,000$287,500
C C++ System Architecture Embedded Systems Operating Systems Hardware Acceleration ISO 26262 ASPICE ISO 21434 Generative AI
8+ yrs exp Remote

Senior System Software Engineer, Neural Graphics SDKs

Nvidia

Santa Clara, CA +1 136 days ago $184,000$287,500
Python C++ CUDA Gaussian Splatting NeRFs Kubernetes CI/CD micro-services GLSL HLSL Metal Slang Source Control Real-time Graphics Computer Vision Neural Reconstruction
5+ yrs exp

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid