Senior Software Engineer, AI Inference Performance

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$184,000–$287,500 / yr
Posted
16 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $185k
This role $236k
$104k most similar roles pay here $307k

This role pays more than 85% of similar roles. Most pay $151,475–$217,725 — the shaded band above. At the midpoint, this role pays about $236k versus about $185k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, AI Inference Performance

Senior Software Engineer - AI Inference Performance will join the team to advance innovative LLM and VLM inference on GPU-accelerated systems. This role involves pushing workloads toward practical performance limits by analyzing inference processes, defining prefill and decode workloads, and optimizing metrics like latency, throughput, and KV-cache capacity. The engineer will build speed-of-light and roofline models, profile workloads using NVIDIA Nsight Systems and Nsight Compute, and eliminate bottlenecks in host code, CUDA kernels, and memory. Key responsibilities include tuning serving hyperparameters such as quantization and speculative decoding while developing performance-critical kernels for attention and matrix multiplication. Required skills include proficiency in Python, Rust, or C++, along with expertise in CUDA, CUTLASS, and Triton. The role addresses the technical challenge of optimizing large-scale distributed inference across complex hardware architectures and networking topologies.

What does a Software Engineer earn in California?

Median $214000 from 775 postings across 63 companies.

See salary data

What you'll do

  • Analyze end-to-end LLM and VLM inference processes to optimize latency, throughput, and processing efficiency.
  • Build speed-of-light and roofline models to quantify performance headroom and identify optimization hypotheses.
  • Profile workloads using tools like NVIDIA Nsight Systems and PyTorch Profiler to eliminate bottlenecks in host code and CUDA kernels.
  • Tune serving hyperparameters including batching, KV-cache management, quantization, and speculative decoding.
  • Develop and optimize performance-critical kernels for attention, matrix multiplication, and mixture-of-experts routing using CUDA or Triton.
  • Establish repeatable benchmarks and performance regression gates to ensure consistent production quality.
  • Contribute high-quality code improvements to open-source inference engines like TensorRT-LLM, vLLM, and SGLang.

What we're looking for

  • More than 6 years of experience in full-stack LLM/VLM inference performance involving models, serving, distributed runtimes, kernels, and hardware.
  • Strong programming skills in Python, Rust, and/or C++.
  • Hands-on experience with CUDA or another GPU programming environment.
  • Demonstrated expertise in speed-of-light analysis, roofline models, microbenchmarks, and tools including NVIDIA Nsight Systems and Nsight Compute.
  • Deep understanding of GPU architecture, including Tensor Cores, memory hierarchy, caches, occupancy, synchronization, and numerical formats.
  • Understanding of distributed systems and networking for accelerated computing, including collectives, topology, and scale-up/scale-out performance.
  • BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • Contributions to high-performance AI projects like TensorRT-LLM, vLLM, SGLang, PyTorch, CUDA, Triton, or NCCL (preferred).

More like this

Similar roles

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 136 days ago $184,000$287,500
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (CA) +4 15 days ago $152,000$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization Profiling Agile
5+ yrs exp Remote

Software Engineer, Inference (AI Data Engineering)

SpaceX

Palo Alto, CA 23 days ago $135,000$175,000
SGLang vLLM TensorRT-LLM Triton Rust C++ Python Go gRPC REST Docker Kubernetes PostgreSQL ClickHouse MongoDB CI/CD GPU Kernels Quantization Speculative Decoding
2+ yrs exp

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid