Senior High-Performance LLM Training Engineer

Nvidia

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$184,000–$287,500 / yr
Posted
156 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $214k
This role $236k
$160k most similar roles pay here $301k

This role pays more than 59% of similar roles. Most pay $174,081–$254,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $214k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior High-Performance LLM Training Engineer

As a Senior High-Performance LLM Training Engineer, you will join the team focused on optimizing high-performance LLM software stacks for training on thousands of GPUs. Your daily responsibilities include analyzing, profiling, and optimizing AI training workloads across various neural networks while implementing production-quality software across multiple layers of the deep learning platform stack, from drivers to frameworks. You will build and support submissions to the MLPerf Training benchmark suite, develop tools to automate workload analysis, and implement key training workloads in proprietary processor and system simulators for architecture studies. The role requires expertise in C++, Python, and CUDA, alongside a strong background in computer architecture, GPU fundamentals, and deep learning. This position addresses the technical challenge of improving efficiency for large language model training on advanced hardware and software platforms.

What you'll do

  • Analyze and profile AI training workloads on innovative hardware and software platforms.
  • Optimize LLM training performance across PyTorch and JAX frameworks for large-scale GPU clusters.
  • Develop production-quality software across the deep learning stack from drivers to high-level frameworks.
  • Manage and support NVIDIA submissions to the MLPerf Training benchmark suite.
  • Implement key training workloads in proprietary processor and system simulators for architecture studies.
  • Build automated tools for workload analysis, optimization, and other critical development workflows.
  • Identify and solve performance bottlenecks across state-of-the-art neural networks.

What we're looking for

  • A PhD in Computer Science, Electrical Engineering, or Computer Engineering and 5+ years of experience is required.
  • An MS degree (or equivalent experience) and 8+ years of meaningful work experience are required.
  • Candidates must have a strong background in deep learning and neural networks, specifically regarding training.
  • A deep background in computer architecture and familiarity with GPU architecture fundamentals is required.
  • Proven experience analyzing and tuning application performance and processor/system-level performance modeling is required.
  • Proficiency in C++, Python, and CUDA programming languages is required.
  • Experience optimizing LLM training workloads in frameworks like PyTorch and JAX is required.

More like this

Similar roles

Principal High-Performance LLM Training Engineer

Nvidia

Santa Clara, CA 136 days ago $272,000$431,250
PyTorch JAX NeMo CUDA LLM Distributed Training GPU Architecture High-Performance Computing Profiling Benchmarking Networking Memory Systems Mixed Precision Training
10+ yrs exp

Senior AI Training Performance Architect

Nvidia

Santa Clara, CA 49 days ago $184,000$287,500
CUDA C++ Python GPU Architecture Deep Learning Neural Networks Performance Modeling MLPerf Training Computer Architecture Optimization
5+ yrs exp

Senior Performance Engineer, Deep Learning

Nvidia

Santa Clara, CA 28 days ago $152,000$241,500
C++ Python PyTorch JAX CUDA OpenAI Triton cuBLAS cuDNN cuSOLVER Parallel Systems Code Optimization LLM Transformer Engine MLPerf Computer Architecture Operating Systems Multi-node Systems

Senior Deep Learning Systems Architect

Nvidia

Santa Clara, CA 58 days ago $224,000$356,500
Deep Learning Neural Networks PyTorch TensorFlow JAX C++ Python CUDA OpenCL OpenACC MPI OpenMP HPC Computer Architecture Performance Analysis Numerical Analysis GPU Computing
10+ yrs exp

Senior Lead AI Engineer

Capital One Financial

New York, NY +4 16 days ago $229,900$262,400
LLM Python PyTorch Huggingface AWS VectorDBs Nemo Guardrails Go Scala Java C++ C# Machine Learning Reinforcement Learning Similarity Search Model Evaluation
6+ yrs exp