Senior Performance Architect, Nemotron

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CAHillsboro, ORRedmond, WA
Salary
$152,000–$241,500 / yr
Posted
115 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $207k
This role $197k
$139k most similar roles pay here $272k

This role pays less than 56% of similar roles. Most pay $177,250–$236,187 — the shaded band above. At the midpoint, this role pays about $197k versus about $207k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Performance Architect, Nemotron

Senior Performance Architect, Nemotron joins the team to shape the next generation of Nemotron models through performance modeling, analysis, and forward projections. The role involves developing high-fidelity analytical models to prototype emerging algorithmic techniques and hardware optimizations for the Nemotron family. You will model end-to-end performance impacts of generative AI workflows like Speculative Decoding, Agentic Pipelines, Inference-time compute scaling, and RL to determine datacenter needs while prioritizing features to guide software and hardware roadmaps. The position requires expertise in computer architecture, roofline modeling, queuing theory, and statistical performance analysis. You will utilize Python, C++, PyTorch, TRT-LLM, VLLM, SGLang, and CUDA for simulator design and data analysis. This role addresses the technical challenge of model-system-hardware co-design to ensure optimal trade-offs between accuracy, throughput, and interactivity while ensuring intelligence scales efficiently in production environments.

What you'll do

  • Develop high-fidelity analytical performance models to prototype emerging algorithmic techniques and hardware optimizations.
  • Model the end-to-end performance impact of GenAI workflows like Speculative Decoding and Agentic Pipelines.
  • Predict how architectural choices translate into real-world deployment efficiency for Nemotron models.
  • Ensure model designs achieve Pareto-optimal trade-offs between accuracy, throughput, and interactivity on target platforms.
  • Prioritize features to guide future software and hardware roadmaps based on detailed performance modeling.
  • Identify resource bottlenecks by designing experiments and visualizing large performance datasets.
  • Translate complex technical analyses into clear recommendations for cross-functional hardware and software teams.

What we're looking for

  • A minimum of a Master's degree in Computer Science, Electrical Engineering, or a related field is required.
  • Candidates must have 3+ years of experience in system evaluation of AI/ML workloads or performance analysis and modeling for AI.
  • Strong background in computer architecture, roofline modeling, queuing theory, and statistical performance analysis techniques is required.
  • Solid understanding of ML fundamentals, model parallelism, and inference serving techniques is required.
  • Proficiency in Python is required, with C++ skills being optional for simulator design and data analysis.
  • Experience with deep learning frameworks such as PyTorch, TRT-LLM, VLLM, or SGLang is required.
  • Experience with GPU computing (CUDA) is preferred to stand out.
  • Ability to distill complex analyses into clear recommendations for both technical and non-technical collaborators is required.

More like this

Similar roles

Senior Accelerated Computing Architect

Nvidia

Santa Clara, CA 127 days ago $184,000$287,500
CUDA OpenCL C C++ Python MPI NVSHMEM OpenSHMEM IPC Benchmarking Profiling Linear Algebra Numerical Methods High Performance Computing Machine Learning AI
6+ yrs exp

Senior Accelerated Computing Architect

Nvidia

Santa Clara, CA 150 days ago $184,000$287,500
CUDA OpenCL C C++ Python MPI NVSHMEM OpenSHMEM IPC Linear Algebra Numerical Methods Software-Hardware Co-design Data Structures Algorithms
6+ yrs exp

Senior CPU Performance Architect

Nvidia

Santa Clara, CA +2 156 days ago $224,000$356,500
CPU Microarchitecture System Architecture GPU PyTorch Benchmarking Performance Analysis HPC AI Deep Learning I/O
10+ yrs exp

Senior AI Training Performance Architect

Nvidia

Santa Clara, CA 49 days ago $184,000$287,500
CUDA C++ Python GPU Architecture Deep Learning Neural Networks Performance Modeling MLPerf Training Computer Architecture Optimization
5+ yrs exp

Senior Performance Architect

Microsoft

Hillsboro, OR 7 days ago $119,800$234,700
C++ Python System on Chip (SOC) Intellectual Property (IP) RTL Computer Architecture Performance Modeling AI Memory Subsystems Interconnect Quality of Service (QoS) Microarchitecture
5+ yrs exp