Engineering Manager, LLM Performance

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$224,000–$356,500 / yr
Posted
37 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $223k
This role $290k
$156k most similar roles pay here $378k

This role pays more than 87% of similar roles. Most pay $177,650–$268,234 — the shaded band above. At the midpoint, this role pays about $290k versus about $223k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Engineering Manager, LLM Performance

Engineering Manager, LLM Performance will lead and grow a team focused on accelerating the performance of large language model inference across various frameworks including TensorRT LLM, vLLM, SGLang, and Dynamo for datacenter products. The role involves driving the design, implementation, and optimization of features critical to inference performance on current and upcoming GPU architectures while collaborating with researchers and architects to deliver production-grade software. Key responsibilities include project planning, milestone delivery, and cross-functional coordination to improve foundation model performance and provide an intuitive developer experience for deployment. Candidates should possess expertise in C++, Python, and software design, alongside a deep understanding of CUDA programming, GPU architecture, and system-level performance tuning. The role addresses the technical challenge of optimizing inference for large language models and vision language models within the AI ecosystem.

What does a Engineering Manager earn in California?

Median $290250 from 64 postings across 19 companies.

See salary data

What you'll do

  • Lead and grow a team responsible for optimizing LLM inference across frameworks like TensorRT-LLM, vLLM, and SGLang.
  • Drive the design, implementation, and optimization of features critical to high-performance LLM inference.
  • Improve inference performance on current and upcoming NVIDIA datacenter architectures and GPUs.
  • Optimize the performance of key foundation models for large language models and vision language models.
  • Partner with benchmark teams to tune performance for specific workloads.
  • Integrate cutting-edge technologies to provide an intuitive developer experience for LLM deployment.
  • Manage software development execution, including project planning, milestone delivery, and cross-functional coordination.

What we're looking for

  • MS, PhD, or equivalent experience in Computer Science, Computer Engineering, AI, or a related technical field.
  • Overall software engineering experience.
  • 3+ years of technical leadership experience.
  • Proven ability to lead and scale high-performing engineering teams across distributed and cross-functional groups.
  • Strong background in C++ or Python with expertise in software design and production-quality libraries.
  • Demonstrated expertise in large language models (LLM), vision language models (VLM), or inference in general.
  • Deep understanding of GPU architecture, CUDA programming, and system-level performance tuning.
  • Experience with frameworks such as TensorRT-LLM, vLLM, or SGLang.

More like this

Similar roles

Engineering Manager, LLM Inference & Deployment at Scale

Nvidia

Santa Clara, CA 14 days ago $224,000$356,500
LLMs VLMs TensorRT TensorRT-LLM vLLM SGLang Quantization Speculative Decoding Continuous Batching Prefix Caching KV-cache Optimization Distributed Computing GPU Cluster Orchestration Model Serving Inference Optimization Deep Learning
8+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp

Manager, Software Architecture

Nvidia

Santa Clara, CA 52 days ago $184,000$287,500
C C++ Rust Python Distributed Systems Networking Computer Architecture vLLM SGLang TensorRT-LLM Agile Hardware-Software Co-design Transformer Architectures
8+ yrs exp