Senior AI Training Performance Architect

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$184,000–$287,500 / yr
Posted
49 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $204k
This role $236k
$158k most similar roles pay here $301k

This role pays more than 75% of similar roles. Most pay $172,300–$235,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $204k for comparable roles.

Based on 238 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior AI Training Performance Architect

As a Senior AI Training Performance Architect, you will join the team to analyze, profile, and optimize AI training workloads on state-of-the-art hardware and software platforms. You will be responsible for identifying performance bottlenecks on GPUs, prioritizing solutions across key training workloads, and implementing production-quality software across multiple layers of the deep learning platform stack, from drivers to frameworks. Your daily work includes building and supporting submissions for MLPerf Training benchmarks, implementing training workloads in proprietary processor and system simulators for architecture studies, and developing tools to automate workload analysis and optimization workflows. To succeed, you must possess a strong background in deep learning, neural networks, and computer architecture while demonstrating proficiency in C++, Python, and CUDA. This role focuses on the technical challenge of maximizing performance across the hardware and software stack for large-scale compute systems.

What you'll do

  • Analyze and profile AI training workloads on state-of-the-art hardware and software platforms.
  • Identify and resolve performance bottlenecks across key AI training workloads on GPUs.
  • Implement production-quality software across the deep learning platform stack from drivers to frameworks.
  • Build and support NVIDIA submissions for MLPerf Training benchmarks.
  • Implement DL training workloads in proprietary processor and system simulators for architecture studies.
  • Develop tools to automate workload analysis, optimization, and other critical workflows.

What we're looking for

  • PhD in CS, EE, or CSEE with 5+ years of relevant experience.
  • MS degree with 8+ years of relevant experience.
  • Strong background in deep learning and neural networks, specifically in training.
  • Solid understanding of computer architecture and GPU architecture fundamentals.
  • Proven experience in analyzing and tuning application performance.
  • Proven experience in processor and system-level performance modeling.
  • Proficiency in programming with C++, Python, and CUDA.

More like this

Similar roles

Senior Deep Learning Systems Architect

Nvidia

Santa Clara, CA 58 days ago $224,000$356,500
Deep Learning Neural Networks PyTorch TensorFlow JAX C++ Python CUDA OpenCL OpenACC MPI OpenMP HPC Computer Architecture Performance Analysis Numerical Analysis GPU Computing
10+ yrs exp

Senior Deep Learning Performance Architect

Nvidia

CA, Canada 31 days ago $152,000$241,500
Deep Learning AI Inference C C++ Python CUDA HPC MPI OpenMP Computer Architecture Performance Modeling Profiling Compiler ASIC Machine Learning
5+ yrs exp Hybrid

Senior Accelerated Computing Architect

Nvidia

Santa Clara, CA 150 days ago $184,000$287,500
CUDA OpenCL C C++ Python MPI NVSHMEM OpenSHMEM IPC Linear Algebra Numerical Methods Software-Hardware Co-design Data Structures Algorithms
6+ yrs exp

Senior Accelerated Computing Architect

Nvidia

Santa Clara, CA 127 days ago $184,000$287,500
CUDA OpenCL C C++ Python MPI NVSHMEM OpenSHMEM IPC Benchmarking Profiling Linear Algebra Numerical Methods High Performance Computing Machine Learning AI
6+ yrs exp