Senior DL Performance Efficiency Architect

Nvidia

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$184,000–$287,500 / yr
Posted
8 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $204k
This role $236k
$149k most similar roles pay here $302k

This role pays more than 76% of similar roles. Most pay $171,487–$235,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $204k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior DL Performance Efficiency Architect

As a Senior DL Performance Efficiency Architect on the DL Architecture team, you will drive a unified strategy to improve large language model efficiency from research through deployment. You will lead cross-layer efforts involving model architecture, training, and inference systems while analyzing how workloads map to GPUs, memory systems, interconnects, and distributed infrastructure. Your role involves identifying opportunities for model-system-hardware co-design and establishing a measurement-driven roadmap. You will collaborate with researchers, systems engineers, compiler developers, and hardware architects to influence future roadmaps. Key technical competencies include roofline modeling, workload characterization, benchmarking, and hardware-aware optimization. You will address challenges in computational efficiency regarding memory, power, and cost by leveraging expertise in low-precision computation, quantization, sparsity, Mixture-of-Experts, long-context inference, and speculative decoding to deliver scalable improvements for complex large language model workloads.

What you'll do

  • Lead cross-layer efforts to improve LLM efficiency across model architecture, training systems, and inference.
  • Analyze how LLM workloads map to GPUs, memory systems, interconnects, and distributed infrastructure.
  • Identify and execute opportunities for model-system-hardware co-design.
  • Establish a measurement-driven roadmap for LLM efficiency from initial investigation to production deployment.
  • Influence future roadmaps for models, software, and hardware by partnering with cross-functional engineering teams.
  • Perform performance analysis, roofline modeling, and workload characterization to identify fundamental bottlenecks.
  • Implement optimizations for low-precision computation, quantization, sparsity, and speculative decoding.

What we're looking for

  • MS or PhD degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • 5+ years of relevant experience in AI systems, model architecture, computer architecture, high-performance computing, or performance optimization.
  • Strong understanding of LLM architectures and the trade-offs between model quality, cost, memory footprint, latency, throughput, and power.
  • Strong background in performance analysis, roofline modeling, workload characterization, benchmarking, and hardware-aware optimization.
  • Proven ability to provide technical leadership and drive complex optimization projects from concept to production.
  • Experience with low-precision computation, quantization, sparsity, Mixture-of-Experts, long-context inference, and speculative decoding.
  • Experience co-designing model architectures with training, inference, compiler, or hardware constraints.
  • Experience influencing accelerator, system, or datacenter architecture based on future AI workload requirements.

More like this

Similar roles

Senior AI Performance and Efficiency Engineer

Nvidia

Remote (Santa Clara, CA) +2 176 days ago $152,000$241,500
Python Go Bash CUDA NCCL PyTorch TensorFlow NSight Systems NSight Compute AWS GCP Azure InfiniBand RDMA Lustre GPFS MLPerf Distributed Training Parallel Computing Machine Learning
5+ yrs exp Remote

Senior Deep Learning Performance Architect

Nvidia

CA, Canada 31 days ago $152,000$241,500
Deep Learning AI Inference C C++ Python CUDA HPC MPI OpenMP Computer Architecture Performance Modeling Profiling Compiler ASIC Machine Learning
5+ yrs exp Hybrid

Engineering Manager, LLM Inference & Deployment at Scale

Nvidia

Santa Clara, CA 14 days ago $224,000$356,500
LLMs VLMs TensorRT TensorRT-LLM vLLM SGLang Quantization Speculative Decoding Continuous Batching Prefix Caching KV-cache Optimization Distributed Computing GPU Cluster Orchestration Model Serving Inference Optimization Deep Learning
8+ yrs exp Hybrid

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp

Senior CPU Performance Architect

Nvidia

Santa Clara, CA +2 156 days ago $224,000$356,500
CPU Microarchitecture System Architecture GPU PyTorch Benchmarking Performance Analysis HPC AI Deep Learning I/O
10+ yrs exp