Senior Performance Engineer, Deep Learning

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$152,000–$241,500 / yr
Posted
28 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $211k
This role $197k
$139k most similar roles pay here $276k

This role pays less than 63% of similar roles. Most pay $176,000–$246,150 — the shaded band above. At the midpoint, this role pays about $197k versus about $211k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Performance Engineer, Deep Learning

As a Senior Performance Engineer - Deep Learning, you will join the Deep Learning models performance engineering team to build and optimize libraries and tools that enable researchers to develop efficient AI applications. You will focus on building and supporting Transformer Engine for accelerating Large Language Model training while collaborating on systems research involving low precision and parallelism methods. Your daily work involves implementing, benchmarking, and optimizing new models like LLMs to scale on GPUs, contributing to community benchmarks like MLPerf, and influencing hardware design. The role requires proficiency in C++, Python, and parallel systems programming, with knowledge of Computer Architecture and Operating Systems. You will utilize frameworks such as PyTorch and JAX, along with libraries like cuBLAS, cuDNN, and cuSOLVER, while potentially utilizing CUDA or OpenAI Triton to develop GPU kernels for high-performance computing.

What you'll do

  • Build and support the Transformer Engine library to accelerate Large Language Model training.
  • Implement and optimize new Deep Learning models to scale efficiently on NVIDIA GPUs and systems.
  • Conduct systems research to improve performance through low precision training and parallelism methods.
  • Benchmark and optimize code for community benchmarks such as MLPerf.
  • Contribute to open-source projects including PyTorch and JAX frameworks.
  • Engage with the open-source community and enterprise customers to deliver hardware and software innovations.
  • Influence the design of new hardware generations and core platform software components.

What we're looking for

  • Bachelor's degree or equivalent experience in Computer Science, Electrical Engineering, or a related field.
  • 3+ years of experience in C++ and Python programming.
  • Strong background or experience in parallel systems programming, preferably on GPUs.
  • Knowledge of Computer Architecture, Code Optimization, and/or Operating Systems.
  • Proven experience in developing large software projects.
  • Excellent verbal and written communication skills.
  • Experience with PyTorch, JAX, or other Deep Learning frameworks.
  • Experience with performance analysis, profiling, and code optimization for multi-GPU or multi-node systems.

More like this

Similar roles

Senior Deep Learning Compiler Engineer

Nvidia

Remote (Santa Clara, CA) +1 134 days ago $152,000$241,500
C++ Python CUDA OpenCL MLIR XLA TVM LLVM PyTorch GPU Architecture Compiler Optimization Deep Learning Kernel Generation cross compilation Performance Analysis
3+ yrs exp Remote

Senior Deep Learning Compiler Engineer, XLA

Nvidia

Remote 28 days ago $152,000$241,500
C++ CUDA XLA MLIR LLVM OpenAI Triton JAX PyTorch TensorFlow TVM High-Performance Computing Distributed Programming Compiler Optimization Deep Learning
4+ yrs exp Remote

Senior Deep Learning Frameworks CUDA Software Engineer

Nvidia

Remote (Santa Clara, CA) +1 11 days ago $184,000$287,500
CUDA PyTorch JAX C++ Python TRT-LLM vLLM SGLang TensorRT Triton NCCL MPI UCX XLA HPC Kernel Authoring NVIDIA Nsight Systems Compiler Technologies
8+ yrs exp Remote