Senior Software Engineer, AI Inference Performance
Nvidia
Quick summary
Market check
How this pay compares to similar roles
This role pays less than 71% of similar roles. Most pay $196,750–$254,750 — the shaded band above. At the midpoint, this role pays about $197k versus about $226k for comparable roles.
Based on 240 similar postings.
Employer
Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing
Nvidia currently has 896 open roles on FindRole.
Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.
Most-posted roles
At a glance
As a Senior Deep Learning Software Engineer, Inference, you will join the team responsible for developing and maintaining high-performance open-source frameworks for large-scale model serving. You will design, build, and optimize GPU-accelerated software to facilitate the deployment of groundbreaking language models. Your daily work involves performance optimization, analysis, and tuning of deep learning models across various domains including LLM, Multimodal, and Generative AI. You will contribute code to inference libraries such as vLLM, SGLang, and FlashInfer while implementing algorithms for public release. The role requires expert C/C++ programming skills and experience with Python. You will utilize tools like CUTLASS, OAI Triton, NCCL, and CUDA kernels to optimize model serving pipelines across datacenter GPUs and edge SoCs. Your work addresses the technical challenge of scaling performance for state-of-the-art models across diverse hardware architectures.
Skills
What you'll do
What we're looking for
More like this
Nvidia
Nvidia
Nvidia
Nvidia
Nvidia
Nvidia