Senior DL Algorithms Engineer, Inference Performance

Nvidia

Confirmed live 2 days ago High trust
Remote

Quick summary

Work type
Remote
Location
Canada
Salary
$152,000–$241,500 / yr
Posted
129 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $226k
This role $197k
$136k most similar roles pay here $304k

This role pays less than 72% of similar roles. Most pay $196,750–$254,750 — the shaded band above. At the midpoint, this role pays about $197k versus about $226k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior DL Algorithms Engineer, Inference Performance

As a Senior DL Algorithms Engineer - Inference Performance, you will join the team to optimize LLM and Omni model performance across the entire hardware and software stack. Your daily responsibilities include enabling state-of-the-art open models on accelerated inference stacks, contributing features and production code to frameworks like TRT-LLM, vLLM, SGLang, and FlashInfer. You will profile bottlenecks, perform competitive benchmarking, and co-design next-generation AI services. The role requires expertise in deep learning, neural networks, and GPU architecture. You must be proficient in PyTorch or similar high-performance computing frameworks while utilizing skills in performance profiling and analysis. Candidates should possess a strong understanding of computer architecture and algorithms, with preferred experience in CUDA or OpenCL programming to solve complex inference performance challenges within the domain of large language model and diffusion architectures.

What you'll do

  • Optimize state-of-the-art open models on NVIDIA’s accelerated inference software stack.
  • Develop new features and fix bugs in open-source frameworks like TRT-LLM, vLLM, and SGLang.
  • Profile and analyze performance bottlenecks across the entire inference stack to improve speed.
  • Benchmark state-of-the-art offerings and perform competitive analysis for NVIDIA’s software and hardware.
  • Co-design next-generation AI models and services with internal partner teams.
  • Implement optimizations across the full hardware/software stack from GPU architecture to deep learning frameworks.

What we're looking for

  • PhD in Computer Science, Electrical Engineering, CSEE, or equivalent experience.
  • At least 3 years of relevant professional experience.
  • Strong background in deep learning and neural networks, specifically regarding inference.
  • Experience with performance profiling, analysis, and optimization for GPU-based applications.
  • Proficiency in PyTorch or other frameworks for AI and HPC-heavy application development.
  • Deep understanding of computer architecture and fundamentals of GPU architecture.
  • Proven experience with processor and system-level performance optimization.
  • Strong fundamentals in algorithms and deep understanding of modern LLM/Diffusion architectures.

More like this

Similar roles

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (CA) +4 15 days ago $152,000$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization Profiling Agile
5+ yrs exp Remote

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp

Senior Performance Engineer, Deep Learning

Nvidia

Santa Clara, CA 28 days ago $152,000$241,500
C++ Python PyTorch JAX CUDA OpenAI Triton cuBLAS cuDNN cuSOLVER Parallel Systems Code Optimization LLM Transformer Engine MLPerf Computer Architecture Operating Systems Multi-node Systems

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 136 days ago $184,000$287,500
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid

Senior Deep Learning Compiler Engineer, XLA

Nvidia

Remote 28 days ago $152,000$241,500
C++ CUDA XLA MLIR LLVM OpenAI Triton JAX PyTorch TensorFlow TVM High-Performance Computing Distributed Programming Compiler Optimization Deep Learning
4+ yrs exp Remote