Distinguished Engineer, Scaled Out Inferencing

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$320,000–$488,750 / yr
Posted
31 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $221k
This role $404k
$127k most similar roles pay here $527k

This role pays more than 99% of similar roles. Most pay $169,625–$272,500 — the shaded band above. At the midpoint, this role pays about $404k versus about $221k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Distinguished Engineer, Scaled Out Inferencing

As a Distinguished Engineer, Scaled Out Inferencing, you will join the team to lead the global strategy for scaled-out AI inferencing. You will architect high-throughput, low-latency distributed pipelines and model serving strategies designed for massive scale and production reliability across enterprise and cloud environments. Your daily work involves defining technical roadmaps for full-lifecycle management, including automated deployment, versioning, and intelligent scaling. You will perform hardware-software co-optimization, tuning performance at the kernel and driver levels while managing GPU resources. The role requires expertise in CUDA, TensorRT-LLM, vLLM, SGLang, Linux, Kubernetes, and Ray. You will solve complex problems regarding distributed systems for AI workloads by ensuring high availability and durability. By collaborating on open-source projects and infrastructure partnerships, you will ensure advanced models run with peak efficiency on accelerated computing hardware.

What you'll do

  • Architect high-throughput, low-latency distributed pipelines for massive-scale AI inference workloads.
  • Define and drive the technical roadmap for full-lifecycle model management including deployment, versioning, and automated scaling.
  • Perform hardware-software co-optimization by tuning performance at the kernel and driver levels to optimize GPU resource management.
  • Guide and influence open-source projects like TensorRT-LLM, vLLM, and SGLang to improve inference on NVIDIA hardware.
  • Establish systems and orchestration layers to ensure advanced AI models run with peak efficiency on accelerated computing hardware.
  • Design and manage the full software lifecycle from initial architecture and development to deployment and operations.
  • Collaborate with customers and infrastructure providers to establish industry standards for performance and availability.

What we're looking for

  • Experience in technical roles with a focus on AI infrastructure and large-scale inference orchestration.
  • 7 to 10+ years of leadership experience.
  • BS/MS or higher degree in systems engineering, software engineering, or related fields.
  • Proficiency in GPU architecture, hardware acceleration, and low-level performance tuning including CUDA and kernels.
  • Proven track record building secure, highly available, and durable production distributed systems.
  • Ability to synthesize cross-functional needs into architecture while guiding execution across diverse teams.
  • Experience designing and operating scaled-out systems in enterprise and cloud environments (preferred).
  • Familiarity with open source ecosystems such as Dynamo, TensorRT-LLM, vLLM, SGLang, Ray, Linux, and Kubernetes (preferred).

More like this

Similar roles

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Engineering Manager, LLM Performance

Nvidia

Santa Clara, CA 37 days ago $224,000$356,500
LLM VLM TensorRT-LLM vLLM SGLang CUDA C++ Python GPU Architecture Software Design Dynamo
7+ yrs exp Hybrid

Distinguished Engineer

Capital One Financial

McLean, VA +2 121 days ago $269,100$307,200
AWS Microsoft Azure Google Cloud Python Java Go JavaScript TypeScript Swift Machine Learning Cloud Computing Platform Engineering Observability Scalability
7+ yrs exp

Distinguished Engineer

Elevance Health

Atlanta, GA +2 52 days ago
AI LLM Agentic Systems Distributed Systems Python Java Microservices Event-Driven Architecture APIs Cloud Platforms CI/CD DevOps SDLC
10+ yrs exp Hybrid

Distinguished Engineer

Elevance Health

Chicago, IL 23 days ago $254,584$381,876
LLMs Agentic AI RAG Python Java Microservices Event-Driven Architecture Distributed Systems Vector Databases Apache Iceberg CI/CD DevSecOps FHIR HL7 Knowledge Graphs Machine Learning
10+ yrs exp Hybrid

Distinguished Engineer

Capital One Financial

McLean, VA +2 17 days ago $269,100$307,200
AWS Microsoft Azure Google Cloud Java Python Go JavaScript TypeScript Swift micro-services MFE Architecture AI-Augmented Development Intent-Driven Development Platform Engineering Solution Architecture Cloud Computing
7+ yrs exp