Senior Deep Learning Software Engineer, Inference

Nvidia

Confirmed live 2 days ago High trust
Remote

Quick summary

Work type
Remote
Location
CATXNYWAMA
Salary
$152,000–$241,500 / yr
Posted
15 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $226k
This role $197k
$137k most similar roles pay here $293k

This role pays less than 71% of similar roles. Most pay $196,750–$254,750 — the shaded band above. At the midpoint, this role pays about $197k versus about $226k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Deep Learning Software Engineer, Inference

As a Senior Deep Learning Software Engineer, Inference, you will join the team responsible for developing and maintaining high-performance open-source frameworks for large-scale model serving. You will design, build, and optimize GPU-accelerated software to facilitate the deployment of groundbreaking language models. Your daily work involves performance optimization, analysis, and tuning of deep learning models across various domains including LLM, Multimodal, and Generative AI. You will contribute code to inference libraries such as vLLM, SGLang, and FlashInfer while implementing algorithms for public release. The role requires expert C/C++ programming skills and experience with Python. You will utilize tools like CUTLASS, OAI Triton, NCCL, and CUDA kernels to optimize model serving pipelines across datacenter GPUs and edge SoCs. Your work addresses the technical challenge of scaling performance for state-of-the-art models across diverse hardware architectures.

What you'll do

  • Design and optimize GPU-accelerated software for large-scale model serving and inference.
  • Perform performance optimization, analysis, and tuning of LLM, Multimodal, and Generative AI models.
  • Scale the performance of deep learning models across various NVIDIA accelerator architectures.
  • Contribute features and code to inference libraries such as vLLM, SGLang, and FlashInfer.
  • Implement and optimize model serving pipelines using CUTLASS, OAI Triton, NCCL, and CUDA kernels.
  • Integrate the latest algorithms from the deep learning community into public inference frameworks.
  • Identify and drive performance improvements for state-of-the-art models across datacenter GPUs and edge SoCs.

What we're looking for

  • Master's or PhD degree in Computer Engineering, Computer Science, EECS, or a related field.
  • At least 5 years of relevant software development experience.
  • Excellent C/C++ programming and software design skills.
  • Experience with GPU programming using CUDA, OAI Triton, or CUTLASS.
  • Experience with multi-GPU communications such as NCCL or NVSHMEM.
  • Proficiency in Python is preferred.
  • Experience deploying or optimizing the inference of deep learning models in production is a plus.
  • Background in performance modeling, profiling, debugging, and code optimization is a plus.

More like this

Similar roles

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 136 days ago $184,000$287,500
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Senior Deep Learning Software Engineer

Nvidia

Santa Clara, CA +2 18 days ago $152,000$241,500
MLIR C++ Python Compiler Optimization NVIDIA GPU LLM Inference Computational Graph Optimization Kernel Code Generation Performance Analysis AI Workloads Software Design Debugging Test Development
3+ yrs exp