Engineering Manager, Inference Benchmarking

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CAAustin, TX
Salary
$224,000–$356,500 / yr
Posted
106 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $227k
This role $290k
$163k most similar roles pay here $377k

This role pays more than 83% of similar roles. Most pay $184,156–$270,000 — the shaded band above. At the midpoint, this role pays about $290k versus about $227k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Engineering Manager, Inference Benchmarking

As the Engineering Manager, Inference Benchmarking — AI Perf, you will serve as a Technical Lead Manager within the Dynamo organization to advance the AIPerf open-source benchmarking platform. You will lead an engineering team in building core infrastructure for load generation, ZMQ-based microservices, and GPU telemetry using DCGM, PyNVML, Prometheus metrics, and Kubernetes-native deployments. Your role involves ensuring the accuracy of benchmark results for LLM, multimodal, diffusion, and computer vision inference while managing upstream integrations with vLLM, TRT-LLM, and SGLang. You will mentor engineers in a high-velocity open-source environment to solve problems regarding production inference decisions like cost optimization and latency reduction. The role requires expertise in systems engineering, distributed systems, and deep knowledge of LLM inference mechanics including KV caching, speculative decoding, and measurement reproducibility across datacenter, local, and edge use cases.

What does a Engineering Manager earn in California?

Median $290250 from 64 postings across 19 companies.

See salary data

What you'll do

  • Drive the technical roadmap for AIPerf's core infrastructure including load generation and GPU telemetry.
  • Ensure the accuracy and statistical soundness of benchmark results used by industry engineering groups.
  • Advise on upstream engine integrations with vLLM, TRT-LLM, and SGLang to maintain platform relevance.
  • Hire, mentor, and grow a team of senior engineers in a high-velocity open-source environment.
  • Manage the development of benchmarking tools for LLM, multimodal, diffusion, and computer vision inference.
  • Oversee Kubernetes-native infrastructure deployments including operators and GPU observability tooling.
  • Lead technical strategy to establish AIPerf as the primary standard for datacenter, local, and edge use cases.

What we're looking for

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • Software engineering experience building performance-critical infrastructure, ML tooling, or distributed systems.
  • 3+ years of engineering leadership experience as a tech lead, TLM, or engineering manager.
  • Deep understanding of LLM inference mechanics including TTFT, ITL, KV caching, and speculative decoding.
  • Ability to ensure measurement correctness and reproducibility for benchmark results.
  • Proven track record of collaborating across multi-functional groups in high-velocity, high-visibility environments.
  • Experience with vLLM, TRT-LLM, or SGLang internals and contributions to their upstream projects.
  • Experience building Kubernetes-native infrastructure including operators, Helm charts, and GPU observability tooling.

More like this

Similar roles

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Senior AI Inference Platform Engineer

Apple Inc

Seattle, WA 31 days ago $175,000$308,500
Python Go C++ Triton TensorRT-LLM vLLM Kubernetes Prometheus Grafana Splunk CI/CD Data Pipelines Performance Benchmarking Distributed Systems Capacity Planning
7+ yrs exp

Senior AI Inference Platform Engineer

Apple Inc

Seattle, WA 30 days ago $175,000$308,500
Python Go C++ Triton TensorRT-LLM vLLM Kubernetes Prometheus Grafana Splunk CI/CD Data Pipelines Performance Benchmarking Distributed Systems Capacity Planning
7+ yrs exp

Engineering Manager

MongoDB

Gurugram, India 36 days ago
Python TypeScript LangChain LlamaIndex n8n LangGraph Mastra MongoDB CI/CD Distributed Systems Open Source Vector Search Query Optimization Automated Testing
8+ yrs exp Hybrid