Performance Engineer, Deep Learning

Nvidia

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$124,000–$195,500 / yr
Employment
Full-time
Posted
4 days ago
Freshness
Confirmed live today

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $217k
This role $160k
$107k $280k
below market most similar roles pay here above market

This role pays less than 92% of similar roles. Most pay $187,850–$246,150 — the blue band above. At the midpoint, this role pays about $160k versus about $217k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1448 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 1088 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Performance Engineer, Deep Learning

Performance Engineer - Deep Learning joins the Deep Learning models performance engineering team to build and optimize libraries and tools for designing and deploying efficient AI applications. You will build and support the Transformer Engine open-source library to accelerate Large Language Model training while collaborating on systems research regarding low precision training and parallelism methods. Responsibilities include implementing, benchmarking, and optimizing new models to scale on GPUs, contributing to MLPerf benchmarks, and influencing future hardware designs. The role requires proficiency in C++ and Python, with a strong background in parallel systems programming, computer architecture, and code optimization. Candidates should possess experience with PyTorch, JAX, and GPU kernels using CUDA, OpenAI Triton, CuTeDSL, or Pallas to solve complex performance bottlenecks within mainstream open-source Deep Learning frameworks.

What you'll do

  • Build and support the Transformer Engine library for accelerating Large Language Model training.
  • Implement and optimize new Deep Learning models to scale efficiently on NVIDIA GPUs and systems.
  • Collaborate on systems research to improve performance through low precision training and parallelism methods.
  • Build and contribute to NVIDIA submissions for community benchmarks like MLPerf.
  • Influence the design of new hardware generations and core platform software components.
  • Support enterprise customers and partners by delivering NVIDIA’s latest hardware and software innovations.
  • Contribute optimizations directly into mainstream open-source frameworks like PyTorch and JAX.

What we're looking for

  • BS or equivalent experience in Computer Science, Electrical Engineering, or a related field.
  • 2+ years of experience in C++ and Python programming.
  • Strong background, experience, or coursework in parallel systems programming, preferably on GPUs (preferred).
  • Knowledge of Computer Architecture, Code Optimization, and/or Operating Systems.
  • Proven experience in developing large software projects.
  • Excellent verbal and written communication skills.
  • Experience in PyTorch, JAX, or any other DL framework (preferred).
  • Experience with performance analysis, profiling, and code optimization techniques, especially with multi-GPU or multi-node systems (preferred).

More like this

Similar roles

Senior Performance Engineer, Deep Learning

Nvidia

Santa Clara, CA 56 days ago $152,000–$241,500
C++ Python PyTorch JAX CUDA OpenAI Triton cuBLAS cuDNN cuSOLVER Parallel Systems Deep Learning LLMs Transformer Engine MLPerf Code Optimization Computer Architecture Operating Systems

Deep Learning Compiler Engineer

Nvidia

Remote (Santa Clara, CA) +5 47 days ago $152,000–$241,500
CUDA C++ MLIR LLVM TensorIR XLA TVM OpenCL GPU Architecture Compiler Optimization Deep Learning Performance Analysis IR Design C Software Design
3+ yrs exp Remote

Senior Deep Learning Compiler Engineer

Nvidia

Remote (Santa Clara, CA) +2 162 days ago $152,000–$241,500
C++ Python CUDA OpenCL MLIR XLA TVM LLVM PyTorch GPU Architecture Deep Learning Compiler Optimization cross compilation Performance Analysis Kernel Generation
3+ yrs exp Remote

Senior Deep Learning Software Engineer, Inference

Nvidia

CA +5 43 days ago $152,000–$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization FlashInfer Agile
5+ yrs exp Hybrid

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (Santa Clara, CA) 29 days ago $184,000–$287,500
CUDA C++ Python vLLM SGLang PyTorch Triton CUTLASS NCCL NVSHMEM FlashInfer Deep Learning GPU Programming High-Performance Computing Model Serving Performance Optimization Agile
6+ yrs exp Remote

Senior Deep Learning Compiler Engineer, HW-SW Codesign

Nvidia

Remote (Santa Clara, CA) +2 24 days ago $152,000–$241,500
C++ CUDA XLA MLIR LLVM OpenAI Triton JAX PyTorch TensorFlow TVM OpenCL Distributed Programming High-Performance Computing Compiler Optimization Deep Learning Frameworks
4+ yrs exp Remote