Senior Product Manager, AI Inference Performance

Nvidia

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$208,000–$327,750 / yr
Posted
30 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $202k
This role $268k
$147k most similar roles pay here $347k

This role pays more than 94% of similar roles. Most pay $177,250–$227,750 — the shaded band above. At the midpoint, this role pays about $268k versus about $202k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Product Manager, AI Inference Performance

Senior Product Manager – AI Inference Performance joins the Product Management team to own the inference performance roadmap and drive the company’s Deep Learning and Generative AI strategy. This role involves managing the entire inference stack, including optimization techniques, frameworks, and benchmarking tools to improve latency, efficiency, and cost per token for various deployment scales. The candidate will define strategies for agentic applications, manage framework integrations like TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo, and oversee performance metrics such as TTFT and throughput. Key technical requirements include expertise in KV caching, quantization, speculative decoding, and disaggregated serving. The role requires translating low-level capabilities into business value while managing release readiness and quality bars. This position focuses on the technical challenge of optimizing AI models and applications running on NVIDIA hardware across diverse customer environments.

What does a Product Manager earn in California?

Median $223300 from 163 postings across 44 companies.

See salary data

What you'll do

  • Own the inference performance roadmap across model representation, memory management, and request scheduling.
  • Translate complex optimization techniques into scalable products for a wide range of customer sizes.
  • Define performance strategies for agentic applications including cross-turn cache reuse and request prioritization.
  • Develop integration strategies for optimizations across frameworks like TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo.
  • Establish benchmark methodologies and metrics to ensure credible performance claims regarding latency and cost.
  • Manage the daily product lifecycle including release readiness, quality bars, and regression tracking.
  • Translate low-level technical capabilities into business value for both engineers and executive stakeholders.

What we're looking for

  • 12+ years of experience in product management at a technology company, as a founder, engineering lead, or technical product owner.
  • Depth in AI inference optimization including KV caching, quantization, speculative decoding, and disaggregated serving.
  • Familiarity with inference and orchestration frameworks such as TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo.
  • Proven ability to work independently to define strategy and drive ambiguous problems to shipped outcomes.
  • Experience in operational product management including release management, quality bars, regression tracking, and customer issue resolution.
  • Ability to translate low-level technical capabilities into business value for both engineers and executives.
  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related field of study (or equivalent experience).
  • Engineering experience with LLM inference performance, open-source contributions, or production experience at scale (preferred).

More like this

Similar roles

Senior Product Manager, AI Platform Inference

Nvidia

Santa Clara, CA 38 days ago $168,000$258,750
vLLM SGLang FlashInfer TensorRT-LLM Triton Dynamo TorchAO NVIDIA GPUs Generative AI Machine Learning GPU Architecture Performance Profiling Software Development Open Source GitHub
6+ yrs exp

Technical Marketing Engineer, AI Platform Software

Nvidia

Santa Clara, CA 44 days ago $136,000$212,750
PyTorch TensorRT-LLM Megatron JAX vLLM SGLang Python C/C++ cuDNN NCCL Deep Learning Machine Learning Multi-GPU Quantization Benchmarking Software Development
4+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 38 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Senior Product Manager AI Inference Software

Amd

San Francisco, CA +1 72 days ago $205,680$308,520
ROCm vLLM SGLang AMD ATOM llm-d GPU Computing Distributed Systems Cloud Infrastructure High-Performance Computing Memory Management open-source Inference Engines Technical Writing
Hybrid