Senior Software Engineer, AI Inference Systems

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$184,000–$287,500 / yr
Posted
136 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $185k
This role $236k
$106k most similar roles pay here $307k

This role pays more than 89% of similar roles. Most pay $151,000–$219,562 — the shaded band above. At the midpoint, this role pays about $236k versus about $185k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, AI Inference Systems

As a Senior Software Engineer, AI Inference Systems, you will join a specialized team to build high-performance inference stacks and optimize GPU kernels for large-scale models. Your daily responsibilities include contributing features to vLLM, implementing speculative decoding, and developing high-level DSLs and compiler infrastructure to improve kernel developer productivity. You will also design benchmarking methodologies, manage containerized deployments across multi-cloud environments, and conduct original research in ML Systems. The role requires proficiency in Python, C/C++, and familiarity with Go or Rust. Key technical requirements include experience with CUDA, PyTorch, vLLM, SGLang, and tools like Nsight Systems. You will work on solving complex problems in accelerated computing by optimizing memory layouts, utilizing Triton or MLIR, and managing large-scale inference deployments across distributed GPU clusters to push the frontier of efficient AI performance.

What does a Software Engineer earn in California?

Median $214000 from 775 postings across 63 companies.

See salary data

What you'll do

  • Develop and contribute features to the vLLM framework to support new models and hardware.
  • Optimize inference frameworks using speculative decoding and various forms of parallelism.
  • Develop, optimize, and benchmark hand-tuned and compiler-generated GPU kernels.
  • Build and extend high-level DSLs and compiler infrastructure to improve kernel developer productivity.
  • Create and maintain benchmarking methodologies for the MLPerf Inference suite.
  • Architect and orchestrate containerized large-scale inference deployments across multi-node GPU clusters.
  • Conduct and publish original research to advance the state of ML Systems.

What we're looking for

  • Bachelor's degree in CS/CE/SE with 7+ years of experience, a Master's with 5+ years, or a PhD with relevant publications.
  • Strong programming skills in Python and C/C++.
  • Proficiency in Go or Rust is preferred.
  • Solid knowledge of algorithms, data structures, operating systems, computer architecture, parallel programming, and distributed systems.
  • Experience with GPU programming (CUDA), memory hierarchy, streams, NCCL, and profiling tools like Nsight Systems/Compute.
  • Familiarity with ML frameworks (PyTorch) and inference engines (vLLM, SGLang).
  • Experience with container orchestration (Docker, Kubernetes, Slurm) and Linux namespaces or cgroups.
  • Experience with ML compilers (Triton, MLIR/LLVM), GPU libraries (CUTLASS), and cloud infrastructure (AWS/GCP/Azure).

More like this

Similar roles

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (CA) +4 15 days ago $152,000$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization Profiling Agile
5+ yrs exp Remote

AI Inference Engineer

F5 Inc

San Jose, CA +1 15 days ago $176,600$265,000
vLLM TGI NVIDIA Triton Python C++ Rust TensorRT Llama.cpp Ollama Kubernetes Docker AWS GCP Azure CUDA TPUs MLOps
Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid