Senior Inference Engineer, GPU Kernel Optimization

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CAAustin, TXNew York, NYSeattle, WA
Salary
$184,000–$287,500 / yr
Posted
46 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $226k
This role $236k
$167k most similar roles pay here $300k

This role pays more than 62% of similar roles. Most pay $196,750–$254,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $226k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Inference Engineer, GPU Kernel Optimization

The Senior Inference Engineer, GPU Kernel Optimization joins the LLM Inference Performance Analysis and Optimization team to push inference operations toward their performance ceilings. This role involves developing silicon-measured kernel benchmarking infrastructure, model-level performance projection tooling, and agentic optimization systems that improve GPU kernels at the assembly layer. The engineer will perform microbenchmarking of competing implementations, analyze end-to-end model serving economics, and implement AI-driven analysis to diagnose performance gaps. Key technical requirements include proficiency in Python, C++, and CUDA, along with experience using tools like CUPTI, NSYS, and NCU. Candidates should possess expertise in Triton or CUTLASS, the ability to read PTX or SASS, and familiarity with frameworks such as TRT-LLM, SGLang, or vLLM. The work focuses on identifying bottlenecks across kernel execution, compiler decisions, and runtime scheduling for large language model inference.

What you'll do

  • Measure competing GPU kernel implementations at real-silicon fidelity across production configuration spaces.
  • Analyze end-to-end model performance to identify high-value optimization opportunities and produce serving policies.
  • Develop agentic AI systems to diagnose performance gaps and automate the exploration of kernel optimizations.
  • Perform deep-dive profiling using tools like CUPTI, NSYS, and NCU to attribute bottlenecks in execution and scheduling.
  • Optimize GPU kernels at the assembly level by analyzing PTX or SASS output.
  • Implement improvements across the LLM inference stack using frameworks like TRT-LLM, SGLang, or vLLM.
  • Collaborate with compiler, hardware, and framework teams to deliver upstream performance gains.

What we're looking for

  • Master's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 6+ years of relevant industry experience.
  • Experience building or directing agentic AI systems including code generation and automated optimization workflows.
  • Strong Python and C++ skills with proven software engineering fundamentals.
  • Hands-on GPU profiling experience using tools such as CUPTI, NSYS, and NCU.
  • Direct experience with LLM inference frameworks like TRT-LLM, SGLang, or vLLM.
  • Working knowledge of GPU kernel optimization including CUDA, CUTLASS, Triton, and the ability to read PTX or SASS.
  • Knowledge of compiler middle-end optimization or GPU code generation pipelines such as LLVM or MLIR.

More like this

Similar roles

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (CA) +4 15 days ago $152,000$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization Profiling Agile
5+ yrs exp Remote

Senior Deep Learning Frameworks CUDA Software Engineer

Nvidia

Remote (Santa Clara, CA) +1 11 days ago $184,000$287,500
CUDA PyTorch JAX C++ Python TRT-LLM vLLM SGLang TensorRT Triton NCCL MPI UCX XLA HPC Kernel Authoring NVIDIA Nsight Systems Compiler Technologies
8+ yrs exp Remote

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 136 days ago $184,000$287,500
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid