Senior Deep Learning Software Engineer, Inference
Nvidia
Quick summary
Market check
How this pay compares to similar roles
This role pays more than 62% of similar roles. Most pay $196,750–$254,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $226k for comparable roles.
Based on 240 similar postings.
Employer
Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing
Nvidia currently has 896 open roles on FindRole.
Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.
Most-posted roles
At a glance
The Senior Inference Engineer, GPU Kernel Optimization joins the LLM Inference Performance Analysis and Optimization team to push inference operations toward their performance ceilings. This role involves developing silicon-measured kernel benchmarking infrastructure, model-level performance projection tooling, and agentic optimization systems that improve GPU kernels at the assembly layer. The engineer will perform microbenchmarking of competing implementations, analyze end-to-end model serving economics, and implement AI-driven analysis to diagnose performance gaps. Key technical requirements include proficiency in Python, C++, and CUDA, along with experience using tools like CUPTI, NSYS, and NCU. Candidates should possess expertise in Triton or CUTLASS, the ability to read PTX or SASS, and familiarity with frameworks such as TRT-LLM, SGLang, or vLLM. The work focuses on identifying bottlenecks across kernel execution, compiler decisions, and runtime scheduling for large language model inference.
Skills
What you'll do
What we're looking for
More like this
Nvidia
Nvidia
Nvidia
Nvidia
Nvidia
Nvidia