Senior Deep Learning Frameworks CUDA Software Engineer

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CAAustin, TX
Salary
$184,000–$287,500 / yr
Posted
11 days ago
Freshness
Confirmed live yesterday
Closes
Nov 1, 2026

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $212k
This role $236k
$159k most similar roles pay here $301k

This role pays more than 73% of similar roles. Most pay $181,593–$241,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $212k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Deep Learning Frameworks CUDA Software Engineer

As a Senior Deep Learning Frameworks CUDA Software Engineer, you will join the team responsible for core CUDA features and runtimes to scale Deep Learning and HPC applications. You will integrate new CUDA features and runtime abstractions into AI stacks like PyTorch, TRT-LLM, vLLM, SGLang, and JAX. Your daily responsibilities include performing deep analysis of AI workloads, designing fault-tolerant solutions for large-scale multi-GPU systems, and improving the AI compiler-runtime interface to achieve speed-of-light performance. You will develop exploratory tools to accelerate new paradigms while ensuring code transitions smoothly into production or open-source releases. The role requires expertise in C++, Python, CUDA, and distributed machine learning techniques like pipeline and tensor parallelism. You will solve complex problems regarding multi-GPU demands ranging from massive training scales to microsecond latency inference across various AI models.

What you'll do

  • Integrate new CUDA features and runtime abstractions into major AI frameworks like PyTorch, JAX, and vLLM.
  • Perform deep analysis of AI workloads to identify opportunities for innovation in lower layers of the software stack.
  • Drive improvements in the AI compiler-runtime interface to build high-performance multi-GPU and multi-node solutions.
  • Design fault-tolerant and elastic systems for large-scale or dynamic AI training and inference workloads.
  • Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms.
  • Write maintainable code to transition prototypes into open-source releases, internal tools, or commercial products.
  • Influence the roadmap of core CUDA features to facilitate the development of next-generation DL frameworks.

What we're looking for

  • BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • 8+ years of relevant industry experience or equivalent academic experience after completing a degree.
  • Experience developing with Deep Learning frameworks such as PyTorch and JAX.
  • Experience with inference engines including TRT-LLM, vLLM, and SGLang.
  • Proficiency in rapid prototyping and development using Python, C++, CUDA, or related DSLs.
  • Solid grasp of AI models, parallelisms, and compiler technologies like torch.compile.
  • Experience conducting performance benchmarking on AI clusters and using profiler tools like NVIDIA Nsight Systems.
  • Strong understanding of computer system architecture, hardware-software interactions, and operating systems principles.

More like this

Similar roles

Senior Software Engineer, CUDA Deep Learning Systems

Nvidia

Remote (Santa Clara, CA) +1 11 days ago $184,000$287,500
CUDA C++ Python Deep Learning PyTorch JAX TensorRT vLLM Triton XLA NCCL MPI UCX Distributed Computing Kernel Optimization Model Optimization FP8 INT8
8+ yrs exp Remote

Software Engineer, CUDA Deep Learning Systems

Nvidia

Remote (Santa Clara, CA) +1 11 days ago $124,000$195,500
CUDA C++ Python Deep Learning PyTorch JAX TensorRT vLLM Triton XLA NCCL MPI UCX Distributed Computing Kernel Optimization Model Optimization FP8 INT8
2+ yrs exp Remote

Senior Deep Learning Compiler Engineer

Nvidia

Remote (Santa Clara, CA) +1 134 days ago $152,000$241,500
C++ Python CUDA OpenCL MLIR XLA TVM LLVM PyTorch GPU Architecture Compiler Optimization Deep Learning Kernel Generation cross compilation Performance Analysis
3+ yrs exp Remote

Senior Deep Learning Compiler Engineer, XLA

Nvidia

Remote 28 days ago $152,000$241,500
C++ CUDA XLA MLIR LLVM OpenAI Triton JAX PyTorch TensorFlow TVM High-Performance Computing Distributed Programming Compiler Optimization Deep Learning
4+ yrs exp Remote

Senior Deep Learning Software Engineer

Nvidia

Santa Clara, CA +2 18 days ago $152,000$241,500
MLIR C++ Python Compiler Optimization NVIDIA GPU LLM Inference Computational Graph Optimization Kernel Code Generation Performance Analysis AI Workloads Software Design Debugging Test Development
3+ yrs exp