Engineering Manager, Deep Learning Inference

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$224,000–$356,500 / yr
Posted
45 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $233k
This role $290k
$164k most similar roles pay here $377k

This role pays more than 84% of similar roles. Most pay $196,750–$269,103 — the shaded band above. At the midpoint, this role pays about $290k versus about $233k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Engineering Manager, Deep Learning Inference

Engineering Manager, Deep Learning Inference will lead a high-performing engineering team focused on developing and optimizing open-source frameworks for deep learning inference. The role involves guiding the strategy, roadmap, and execution of software powering large language models and multimodal generative AI systems. You will oversee performance tuning, profiling, and optimization of large-scale models while partnering with internal compiler, library, and research teams to deliver end-to-end optimized pipelines. Key technical requirements include proficiency in C/C++, Python, and GPU programming using CUDA, Triton, and CUTLASS. The position also requires expertise in multi-GPU communications via NIXL, NCCL, and NVSHMEM. This role addresses the critical challenge of making AI deployment scalable and efficient across various hardware environments, from datacenter clusters to edge devices, by optimizing inference for advanced models on NVIDIA GPUs.

What does a Engineering Manager earn in California?

Median $290250 from 64 postings across 19 companies.

See salary data

What you'll do

  • Lead and mentor a high-performing engineering team focused on deep learning inference and GPU-accelerated software.
  • Define the strategy, roadmap, and execution for NVIDIA's open-source inference frameworks.
  • Partner with internal compiler, library, and research teams to deliver optimized end-to-end inference pipelines.
  • Oversee performance tuning, profiling, and optimization of large-scale models for LLM and generative AI applications.
  • Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications.
  • Represent the team in roadmap discussions to ensure alignment with NVIDIA’s broader AI and software strategies.

What we're looking for

  • MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field.
  • Software development experience.
  • 3+ years of experience in technical leadership or engineering management.
  • Strong background in C/C++ software design and development.
  • Proficiency in Python is a plus.
  • Hands-on experience with GPU programming using CUDA, Triton, and CUTLASS.
  • Proven record of deploying or optimizing deep learning models in production environments.
  • Experience leading teams using Agile or collaborative software development practices.

More like this

Similar roles

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (CA) +4 15 days ago $152,000$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization Profiling Agile
5+ yrs exp Remote

Manager, Deep Learning - Autonomous Vehicles and Robotics

Nvidia

Santa Clara, CA 142 days ago $224,000$356,500
Deep Learning TensorRT Transformer Vision-Language Models State Space Models MLIR TVM XLA Triton LangChain LangGraph NVIDIA DRIVE Jetson Inference Optimization Neural Network Compilation ISO 26262
8+ yrs exp

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 136 days ago $184,000$287,500
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp