Engineering Manager, Deep Learning Inference

Nvidia

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CADCTXNYWA
Salary
$184,000–$287,500 / yr
Posted
37 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $234k
This role $236k
$171k most similar roles pay here $310k

This role pays less than 53% of similar roles. Most pay $196,750–$270,500 — the shaded band above. At the midpoint, this role pays about $236k versus about $234k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Engineering Manager, Deep Learning Inference

Engineering Manager, Deep Learning Inference will lead a high-performing engineering team focused on deep learning inference and GPU-accelerated software. This role involves driving the strategy, roadmap, and execution of inference frameworks for Client AI while partnering with internal compiler, library, and research teams to deliver optimized pipelines across NVIDIA accelerators. The manager will oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications. Key technical requirements include proficiency in C/C++, Python, and GPU programming using CUDA, Triton, and CUTLASS. Candidates must also possess expertise in multi-GPU communications via NIXL, NCCL, and NVSHMEM. The role addresses the critical challenge of making AI deployment scalable and efficient by optimizing open-source frameworks like vLLM, SGLang, and FlashInfer to enable real-time inference across datacenter clusters and edge devices.

What does a Engineering Manager earn in California?

Median $290250 from 64 postings across 19 companies.

See salary data

What you'll do

  • Lead and mentor a high-performing engineering team focused on deep learning inference and GPU-accelerated software.
  • Drive the strategy, roadmap, and execution of NVIDIA’s inference frameworks for Client AI.
  • Partner with internal compiler, library, and research teams to deliver optimized end-to-end inference pipelines.
  • Oversee performance tuning, profiling, and optimization of large-scale models for LLM and generative AI applications.
  • Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications.
  • Represent the team in roadmap and planning discussions to align with NVIDIA’s broader AI and software strategies.

What we're looking for

  • MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field.
  • Software development experience.
  • 3+ years of experience in technical leadership or engineering management.
  • Strong background in C/C++ software design and development.
  • Proficiency in Python is a plus.
  • Hands-on experience with GPU programming using CUDA, Triton, and CUTLASS.
  • Proven record of deploying or optimizing deep learning models in production environments.
  • Experience leading teams using Agile or collaborative software development practices.

More like this

Similar roles

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (CA) +4 15 days ago $152,000$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization Profiling Agile
5+ yrs exp Remote

Manager, Deep Learning - Autonomous Vehicles and Robotics

Nvidia

Santa Clara, CA 142 days ago $224,000$356,500
Deep Learning TensorRT Transformer Vision-Language Models State Space Models MLIR TVM XLA Triton LangChain LangGraph NVIDIA DRIVE Jetson Inference Optimization Neural Network Compilation ISO 26262
8+ yrs exp

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 136 days ago $184,000$287,500
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp