Senior Deep Learning Software Infrastructure Engineer

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Canada
Salary
$224,000–$356,500 / yr
Posted
4 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $201k
This role $290k
$136k most similar roles pay here $380k

This role pays more than 97% of similar roles. Most pay $159,437–$241,750 — the shaded band above. At the midpoint, this role pays about $290k versus about $201k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Deep Learning Software Infrastructure Engineer

As a Senior Deep Learning Software Infrastructure Engineer, you will join the Autonomous Vehicles project to build and scale training libraries and infrastructure for end-to-end autonomous driving models. You will be responsible for crafting, scaling, and hardening deep learning infrastructure frameworks for multi-thousand GPU clusters while improving efficiency across data loaders, distributed training, scheduling, and performance monitoring. Your daily work involves building robust pipelines to handle massive video datasets and owning core components like orchestration libraries and fault-resilient systems. The role requires expertise in PyTorch, DDP/FSDP, NCCL, tensor/pipeline parallelism, and Python for production-grade tools. You will also leverage knowledge of datacenter networking including RoCE and IB, parallel filesystems like Lustre, and schedulers such as Slurm or Kubernetes to solve complex problems related to large-scale training and high availability in a distributed environment.

What you'll do

  • Build and scale deep learning infrastructure libraries for training on multi-thousand GPU clusters.
  • Optimize the training stack including data loaders, distributed training, scheduling, and performance monitoring.
  • Develop robust pipelines to handle massive video datasets and enable rapid experimentation.
  • Own core components such as orchestration libraries, distributed training frameworks, and fault-resilient systems.
  • Improve training availability by minimizing stalls and enhancing infrastructure efficiency.
  • Ensure infrastructure scales with growing GPU capacity while maintaining developer stability and efficiency.

What we're looking for

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, or a related field, or equivalent experience.
  • 12+ years of professional experience building and scaling high-performance distributed systems in ML, HPC, or large-scale data infrastructure.
  • Extensive knowledge of deep learning frameworks (PyTorch preferred) and large-scale training techniques like DDP, FSDP, NCCL, and parallelism.
  • Strong systems background including datacenter networking (RoCE, IB), parallel filesystems (Lustre), storage systems, and schedulers (Slurm, Kubernetes).
  • Proficiency in Python for writing production-grade libraries, orchestration layers, and automation tools.
  • Experience scaling large GPU training clusters with more than 1,000 GPUs.
  • Expertise in fault resilience, high availability, elastic training, and large-scale observability.
  • Demonstrated leadership skills as a hands-on technical authority to establish guidelines for ML systems engineering.

More like this

Similar roles

Senior Deep Learning Frameworks CUDA Software Engineer

Nvidia

Remote (Santa Clara, CA) +1 11 days ago $184,000$287,500
CUDA PyTorch JAX C++ Python TRT-LLM vLLM SGLang TensorRT Triton NCCL MPI UCX XLA HPC Kernel Authoring NVIDIA Nsight Systems Compiler Technologies
8+ yrs exp Remote

Senior Research Engineer, Autonomous Vehicles

Nvidia

Santa Clara, CA 37 days ago $184,000$287,500
PyTorch JAX TensorFlow Python C++ CUDA Kubernetes SLURM Deep Learning Reinforcement Learning LLMs MLOps HPC Distributed Training Sim-to-Real Natural Language Processing Graphics
10+ yrs exp

Senior Software Engineer, RL Post-Training Frameworks

Nvidia

Remote (Santa Clara, CA) 21 days ago $184,000$287,500
Reinforcement Learning Python C/C++ PyTorch Kubernetes Ray vLLM SGLang TensorRT-LLM DeepSpeed Megatron-LM NCCL InfiniBand FSDP Distributed Systems High-Performance Computing
5+ yrs exp Remote

Principal Developer, AI Networking

Nvidia

Remote (Santa Clara, CA) +3 91 days ago $272,000$431,250
NCCL RDMA CUDA PyTorch TensorFlow C++ Python Bash RoCE MPI SHARP Distributed Systems LLM Training High-Performance Networking Profiling Benchmarking
10+ yrs exp Remote

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (CA) +4 15 days ago $152,000$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization Profiling Agile
5+ yrs exp Remote