Senior Deep Learning Framework Communications Engineer

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CAWestford, MAAustin, TXDurham, NC
Salary
$184,000–$287,500 / yr
Posted
7 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $221k
This role $236k
$162k most similar roles pay here $301k

This role pays more than 63% of similar roles. Most pay $187,250–$254,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $221k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Deep Learning Framework Communications Engineer

As a Senior Deep Learning Framework Communications Engineer, you will join the team responsible for communication libraries like NCCL and NVSHMEM to integrate advanced communication features into AI stacks including PyTorch, TRT-LLM, vLLM, SGLang, and JAX. Your daily responsibilities involve performing deep analysis of AI workloads to identify multi-GPU requirements, improving compilers to hide communications through automatic fusion, designing fault-tolerant solutions for large-scale workloads, and authoring custom communication or fused compute-communication kernels. You will utilize Python, C++, CUDA, and DSLs like Triton or cuTe while leveraging tools such as NVIDIA Nsight Systems. The role addresses the critical challenge of optimizing communication performance between GPUs to support diverse demands ranging from massive-scale training on 100K GPUs to low-latency inference, ultimately improving the ease of use for the broader AI community.

What you'll do

  • Integrate new communication library features into AI frameworks from proof of concept to production.
  • Analyze AI workloads and frameworks to identify multi-GPU communication requirements and optimization opportunities.
  • Improve AI compilers to hide communications or perform automatic fusion.
  • Conduct in-depth performance characterization of AI workloads on multi-GPU clusters.
  • Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads.
  • Author custom communication or fused compute-communication kernels to optimize performance on NV platforms.
  • Influence the roadmap of core communication libraries like NCCL and NVSHMEM.

What we're looking for

  • Bachelor's, Master's, or PhD in Computer Science or a related field (or equivalent experience).
  • Software engineering and HPC/AI experience.
  • Experience developing or integrating Deep Learning frameworks like PyTorch, JAX, TRT-LLM, vLLM, and SGLang.
  • Proficiency in rapid prototyping with Python, C++, CUDA, or related DSLs such as Triton and cuTe.
  • Solid grasp of AI models, parallelisms, and compiler technologies like torch.compile.
  • Experience conducting performance benchmarking on AI clusters using tools like PyTorch profiler or NVIDIA Nsight Systems.
  • Understanding of HPC/AI communication concepts including 1-sided/2-sided communication, elasticity, resiliency, and topology discovery.
  • Familiarity with parallel programming on communication runtimes such as NCCL, NVSHMEM, or MPI.

More like this

Similar roles

Senior Deep Learning Frameworks CUDA Software Engineer

Nvidia

Remote (Santa Clara, CA) +1 11 days ago $184,000$287,500
CUDA PyTorch JAX C++ Python TRT-LLM vLLM SGLang TensorRT Triton NCCL MPI UCX XLA HPC Kernel Authoring NVIDIA Nsight Systems Compiler Technologies
8+ yrs exp Remote

Senior Deep Learning Communication Architect

Nvidia

Santa Clara, CA +1 114 days ago $184,000$287,500
Deep Learning PyTorch TensorRT-LLM vLLM SGLang C++ Python CUDA OpenCL InfiniBand NVLink NCCL MPI UCX UCC NVSHMEM RoCE Triton FSDP
6+ yrs exp

Senior Performance Engineer, Deep Learning

Nvidia

Santa Clara, CA 28 days ago $152,000$241,500
C++ Python PyTorch JAX CUDA OpenAI Triton cuBLAS cuDNN cuSOLVER Parallel Systems Code Optimization LLM Transformer Engine MLPerf Computer Architecture Operating Systems Multi-node Systems

Senior Deep Learning Software Engineer

Nvidia

Santa Clara, CA +2 18 days ago $152,000$241,500
MLIR C++ Python Compiler Optimization NVIDIA GPU LLM Inference Computational Graph Optimization Kernel Code Generation Performance Analysis AI Workloads Software Design Debugging Test Development
3+ yrs exp