Senior Deep Learning Software Engineer, Inference

Nvidia

Confirmed live today High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$184,000–$287,500 / yr
Employment
Full-time
Posted
17 days ago
Freshness
Confirmed live today

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $220k
This role $236k
$162k most similar roles pay here $301k

This role pays more than 65% of similar roles. Most pay $188,987–$251,875 — the shaded band above. At the midpoint, this role pays about $236k versus about $220k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 892 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 870 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Deep Learning Software Engineer, Inference

Senior Deep Learning Software Engineer, Inference will join the team to design, build, and optimize GPU-accelerated software powering sophisticated AI applications. The role focuses on developing high-performance deep learning frameworks, specifically SGLang and vLLM, to facilitate the deployment of large language models. You will perform performance optimization, analysis, and tuning for LLM, Multimodal, and Generative AI models across various NVIDIA accelerators from datacenter GPUs to edge SoCs. Key responsibilities include contributing code to inference libraries like FlashInfer and collaborating with cross-functional teams on innovative solutions. Required skills include expert C/C++ programming, software design, and Python. You will utilize tools such as CUTLASS, OAI Triton, NCCL, and CUDA kernels to optimize model serving pipelines. The work addresses the technical challenge of efficient large-scale model serving and high-performance inference for state-of-the-art generative models.

What you'll do

  • Design, build, and optimize GPU-accelerated software for high-performance deep learning inference.
  • Develop and maintain core inference libraries including SGLang, vLLM, and FlashInfer.
  • Implement the latest deep learning algorithms for public release in open-source frameworks.
  • Identify and drive performance improvements for LLM and Generative AI models across NVIDIA accelerators.
  • Optimize model serving pipelines using tools like CUTLASS, OAI Triton, NCCL, and CUDA kernels.
  • Scale the performance of deep learning models across various hardware architectures from datacenter GPUs to edge SoCs.
  • Perform performance modeling, profiling, debugging, and code optimization for production inference systems.

What we're looking for

  • Master's or PhD degree in Computer Engineering, Computer Science, EECS, or AI (or equivalent experience).
  • 6+ years of relevant software development experience.
  • Excellent C/C++ programming and software design skills.
  • Experience with Python is a plus.
  • Prior experience training, deploying, or optimizing the inference of DL models in production is a plus.
  • Background in performance modeling, profiling, debugging, code optimization, or CPU/GPU architectural knowledge is a plus.
  • Contributions to Deep Learning Software projects like PyTorch, vLLM, and SGLang are preferred.
  • Experience with Multi-GPU Communications (NCCL, NVSHMEM) and GPU programming (CUDA, OAI TRITON, or CUTLASS) is preferred.

More like this

Similar roles

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (CA) +4 31 days ago $152,000–$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization Profiling Agile
5+ yrs exp Remote

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 53 days ago $184,000–$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 61 days ago $224,000–$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Senior Deep Learning Frameworks CUDA Software Engineer

Nvidia

Remote (Santa Clara, CA) +1 27 days ago $184,000–$287,500
CUDA PyTorch JAX C++ Python TRT-LLM vLLM SGLang TensorRT Triton NCCL MPI UCX XLA HPC Kernel Authoring NVIDIA Nsight Systems Compiler Technologies
8+ yrs exp Remote

Senior Deep Learning Software Engineer

Nvidia

Santa Clara, CA +2 34 days ago $152,000–$241,500
MLIR C++ Python Compiler Optimization NVIDIA GPU LLM Inference Computational Graph Optimization Kernel Code Generation Performance Analysis AI Workloads Software Design Debugging Test Development
3+ yrs exp

Senior Software Engineer, PyTorch - Deep Learning

Nvidia

Santa Clara, CA +5 3 days ago $184,000–$287,500
PyTorch C++ Python CUDA Deep Learning Distributed Parallel Programming Systems Software Deep Learning Compilers Artificial Intelligence
3+ yrs exp Hybrid