Senior Deep Learning Communication Architect

Nvidia

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Santa Clara, CAAustin, TX
Salary
$184,000–$287,500 / yr
Posted
114 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $226k
This role $236k
$172k most similar roles pay here $300k

This role pays more than 61% of similar roles. Most pay $196,750–$254,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $226k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Deep Learning Communication Architect

As a Senior Deep Learning Communication Architect within the software architecture group, you will focus on scaling deep learning models and training or inference frameworks across systems containing hundreds of thousands of nodes. You will be responsible for identifying data transfer bottlenecks, designing efficient communication protocols tailored for deep learning workloads, and collaborating with hardware teams to integrate high-speed interconnects like NVLink and InfiniBand. Your daily work involves developing proofs-of-concept and performing quantitative modeling to validate new communication strategies. The role requires expertise in C++, Python, CUDA, and OpenCL, along with experience in PyTorch, TensorRT-LLM, vLLM, or SGLang. You will address the technical challenges of optimizing large language model training and inference performance using parallelism techniques like Data, Pipeline, Tensor, Expert Parallelism, and FSDP across complex distributed networks including RoCE.

What you'll do

  • Identify and eliminate bottlenecks in data transfer and synchronization during distributed deep learning training and inference.
  • Develop and implement communication algorithms and protocols tailored for deep learning workloads to minimize overhead and latency.
  • Co-craft systems with hardware and software teams that utilize high-speed interconnects like NVLink, InfiniBand, and SPC-X.
  • Integrate communication libraries such as MPI, NCCL, UCX, UCC, and NVSHMEM into system architectures.
  • Research and evaluate new communication technologies to enhance the performance and scalability of deep learning systems.
  • Build proofs-of-concept and conduct experiments to validate and deploy new communication strategies.
  • Perform quantitative modeling to optimize the scaling of large language models on massive distributed systems.

What we're looking for

  • Bachelor's, Master's, or Ph.D. in Computer Science, Electrical Engineering, CSEE, or a related field.
  • 6+ years of experience building, scaling, and parallelizing DNN frameworks for training and inference workloads.
  • Experience optimizing LLM training and inference performance on cutting-edge hardware.
  • Deep understanding of parallelism techniques including Data, Pipeline, Tensor, Expert Parallelism, and FSDP.
  • Knowledge of emerging serving architectures like Disaggregated Serving and inference servers such as Dynamo and Triton.
  • Proficiency in developing code for DNN frameworks like PyTorch, TensorRT-LLM, vLLM, or SGLang.
  • Strong programming skills in C++ and Python.
  • Familiarity with GPU computing (CUDA, OpenCL) and high-speed networks (InfiniBand, RoCE).

More like this

Similar roles

Principal Deep Learning Communication Architect

Nvidia

Remote (Santa Clara, CA) +1 150 days ago $272,000$431,250
NCCL UCX UCC NVSHMEM CUDA InfiniBand RDMA RoCE TensorRT-LLM vLLM SGLang Megatron-Core DeepSpeed JAX XLA PyTorch Distributed Ray HPC High-Performance Computing
10+ yrs exp Remote

Principal Architect, AI Networking

Nvidia

Remote (Santa Clara, CA) +1 141 days ago $272,000$431,250
RDMA NVLink GPUDirect InfiniBand RoCE NCCL UCX MPI NVSHMEM CUDA C C++ Rust Python vLLM SGLang TensorRT-LLM Distributed Training Computer Architecture
10+ yrs exp Remote