Principal High-Performance LLM Training Engineer

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$272,000–$431,250 / yr
Posted
136 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $213k
This role $352k
$134k most similar roles pay here $463k

This role pays more than 98% of similar roles. Most pay $180,500–$246,150 — the shaded band above. At the midpoint, this role pays about $352k versus about $213k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Principal High-Performance LLM Training Engineer

As a Principal High-Performance LLM Training Engineer, you will join the team to drive performance for large-scale AI training and post-training workloads across the full hardware and software stack. You will perform end-to-end analysis and optimization of frontier-scale LLM pre-training and post-training workloads, identifying bottlenecks in compute, memory, communication, scheduling, and kernel efficiency. Your daily work involves developing production-quality software, tools, models, and benchmarking infrastructure to improve training performance and developer velocity. You will utilize PyTorch, JAX, NeMo, and NeMo RL while leveraging expertise in CUDA libraries, distributed training techniques like tensor and pipeline parallelism, and GPU architecture. This role solves the technical challenge of achieving speed-of-light performance for transformer-based models by translating workload insights into concrete hardware and software recommendations to shape future infrastructure across the AI ecosystem.

What you'll do

  • Analyze and optimize large-scale LLM pre-training and post-training workloads across the full hardware and software stack.
  • Identify and remove performance bottlenecks in compute, memory, communication, scheduling, and kernel efficiency.
  • Develop production-quality software, tools, models, and benchmarking infrastructure to improve training performance and developer velocity.
  • Build and refine performance models and simulation methodologies to guide future GPU, networking, and system architecture decisions.
  • Serve as a technical authority to provide hardware and software recommendations based on workload insights.
  • Mentor engineers and establish best practices for large-scale AI performance analysis and optimization.
  • Influence multi-functional decisions across the organization to improve performance across the broader AI ecosystem.

What we're looking for

  • A Master's degree or PhD in Computer Science, Electrical Engineering, Computer Engineering, or a related field is required.
  • Candidates must have 12+ years of relevant work or research experience.
  • Demonstrated principal-level impact in areas such as large-scale AI training systems, GPU performance optimization, or distributed systems.
  • Deep hands-on experience optimizing large-scale deep learning workloads, including transformer-based models and LLM pre-training.
  • Strong understanding of GPU and AI accelerator architecture from individual units to datacenter-scale systems.
  • Experience with distributed training techniques like data, tensor, pipeline, and expert parallelism.
  • Proven track record using profiling, tracing, benchmarking, and performance modeling tools to diagnose complex bottlenecks.
  • Excellent communication and technical leadership skills to influence multi-functional decisions across the organization.

More like this

Similar roles

Fellow GPU Performance Optimization Engineer

Amd

San Jose, CA 167 days ago $268,000$402,000
GPU Distributed Training PyTorch JAX TensorFlow CUDA HIP ROCm Python C++ RDMA NCCL RCCL Megatron-LM Torchtitan MaxText Kernel Optimization Compiler Stack Performance Profiling
Hybrid

Senior Performance Engineer

Nvidia

Remote (Santa Clara, CA) +2 45 days ago $224,000$356,500
C++ Python CUDA PyTorch JAX XLA GPU Computing Distributed Systems Performance Engineering Benchmarking Profiling Observability High-Performance Computing Data Analysis Automation Workflows
Remote

Principal Developer, AI Networking

Nvidia

Remote (Santa Clara, CA) +3 91 days ago $272,000$431,250
NCCL RDMA CUDA PyTorch TensorFlow C++ Python Bash RoCE MPI SHARP Distributed Systems LLM Training High-Performance Networking Profiling Benchmarking
10+ yrs exp Remote

Senior Principal Software Engineer, LLM Engineering

JPMorgan Chase

Palo Alto, CA 18 days ago
LLMs GNNs MLOps LLMOps Python PyTorch TensorFlow Hugging Face AWS SageMaker Bedrock EKS Triton Inference Server vLLM Kubernetes Docker Terraform CI/CD infrastructure‑as‑code Java
10+ yrs exp

Engineering Manager, LLM Performance

Nvidia

Santa Clara, CA 37 days ago $224,000$356,500
LLM VLM TensorRT-LLM vLLM SGLang CUDA C++ Python GPU Architecture Software Design Dynamo
7+ yrs exp Hybrid