Engineering Manager, LLM Inference & Deployment at Scale

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$224,000–$356,500 / yr
Posted
14 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $239k
This role $290k
$166k most similar roles pay here $377k

This role pays more than 85% of similar roles. Most pay $208,600–$270,000 — the shaded band above. At the midpoint, this role pays about $290k versus about $239k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Engineering Manager, LLM Inference & Deployment at Scale

As an Engineering Manager, LLM Inference & Deployment at Scale, you will lead and mentor a high-performing team responsible for building and operating an AI inference platform. You will manage the transition of Large Language Models and Vision-Language Models from research prototypes into optimized, reliable, and scalable production deployments. Your daily work involves driving technical strategy for low-latency, high-throughput, and cost-efficient inference while establishing engineering standards for benchmarking and profiling. You will utilize expertise in quantization, speculative decoding, continuous batching, prefix caching, and KV-cache optimization to improve performance across distributed systems and GPU hardware. The role requires proficiency with tools like TensorRT, TensorRT-LLM, vLLM, and SGLang. You will solve complex problems at the intersection of model optimization, inference systems, and production infrastructure to enable seamless deployment of rapidly evolving generative AI models.

What does a Engineering Manager earn in California?

Median $290250 from 64 postings across 19 companies.

See salary data

What you'll do

  • Lead and mentor a high-performing engineering team building and operating an AI inference platform at scale.
  • Drive the deployment and optimization of LLMs and VLMs for low-latency, high-throughput, and cost-efficient inference.
  • Analyze and optimize end-to-end deep learning workloads across model stacks, distributed systems, and GPU hardware.
  • Transition new generative AI architectures from research prototypes into reliable, production-ready deployments.
  • Define the technical strategy and roadmap for model deployment, scalability, reliability, and performance.
  • Establish engineering standards for benchmarking, profiling, and continuous performance optimization of inference services.
  • Manage the integration of advanced optimization techniques like quantization, speculative decoding, and KV-cache management.

What we're looking for

  • Master’s or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 8+ years of overall experience.
  • 3 years of management or leadership experience.
  • Hands-on experience with LLMs and/or VLMs and understanding of modern deep learning architectures.
  • Expertise in inference optimization techniques like quantization, speculative decoding, continuous batching, prefix caching, and KV-cache optimization.
  • Experience with disaggregated inference, distributed inference, multi-node deployments, and GPU cluster orchestration.
  • Strong technical leadership, communication, and people-management skills.
  • Experience with TensorRT, TensorRT-LLM, vLLM, SGLang, or comparable inference/serving frameworks.

More like this

Similar roles

Engineering Manager, LLM Performance

Nvidia

Santa Clara, CA 37 days ago $224,000$356,500
LLM VLM TensorRT-LLM vLLM SGLang CUDA C++ Python GPU Architecture Software Design Dynamo
7+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Senior Staff LLM Serving Engineer, Cloud AI Engineering

Qualcomm

San Diego, CA +1 155 days ago $158,400$237,600
LLM PyTorch Python Triton-Inference Server vLLM SGLang CUDA Triton torch.compile torchDynamo Distributed Systems Kernel Design Deep Learning KV-Cache Management Model Optimization
4+ yrs exp

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp