Engineering Manager, LLM Inference & Deployment at Scale
Nvidia
Quick summary
Market check
How this pay compares to similar roles
This role pays more than 87% of similar roles. Most pay $177,650–$268,234 — the shaded band above. At the midpoint, this role pays about $290k versus about $223k for comparable roles.
Based on 240 similar postings.
Employer
Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing
Nvidia currently has 896 open roles on FindRole.
Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.
Most-posted roles
At a glance
Engineering Manager, LLM Performance will lead and grow a team focused on accelerating the performance of large language model inference across various frameworks including TensorRT LLM, vLLM, SGLang, and Dynamo for datacenter products. The role involves driving the design, implementation, and optimization of features critical to inference performance on current and upcoming GPU architectures while collaborating with researchers and architects to deliver production-grade software. Key responsibilities include project planning, milestone delivery, and cross-functional coordination to improve foundation model performance and provide an intuitive developer experience for deployment. Candidates should possess expertise in C++, Python, and software design, alongside a deep understanding of CUDA programming, GPU architecture, and system-level performance tuning. The role addresses the technical challenge of optimizing inference for large language models and vision language models within the AI ecosystem.
What does a Engineering Manager earn in California?
Median $290250 from 64 postings across 19 companies.
What you'll do
What we're looking for
More like this
Nvidia
Nvidia
Nvidia
Nvidia
Nvidia
Nvidia