Principal Deep Learning Communication Architect

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CAAustin, TX
Salary
$272,000–$431,250 / yr
Posted
150 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $233k
This role $352k
$149k most similar roles pay here $462k

This role pays more than 96% of similar roles. Most pay $194,565–$271,475 — the shaded band above. At the midpoint, this role pays about $352k versus about $233k for comparable roles.

Based on 239 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Principal Deep Learning Communication Architect

As a Principal Deep Learning Communication Architect, you will join the team to define the long-term technical roadmap for communication libraries across next-generation platforms. You will lead the development of next-generation communication primitives and collective algorithms while ensuring seamless model scaling across clusters comprising hundreds of thousands of nodes. Your daily work involves collaborating with silicon architects and application developers to co-design specialized primitives for trillion-parameter models and Agentic AI. You will utilize technologies including NCCL, NVSHMEM, UCX, UCC, and CUDA programming models, while optimizing for heterogeneous interconnects like NVLink, Spectrum-X, and Quantum-X. The role requires expertise in 3D parallelism, RDMA, RoCE, and InfiniBand verbs to solve complex high-performance computing challenges. You will also develop high-fidelity analytical models to predict system behavior under emerging workloads within the distributed deep learning domain.

What you'll do

  • Define the long-term technical roadmap for communication libraries across next-generation platforms.
  • Lead the development of next-generation communication primitives and collective algorithms for heterogeneous interconnects.
  • Partner with application developers to architect and implement specialized communication primitives for large-scale AI models.
  • Influence hardware specifications for next-generation networking through collaboration with silicon architects and software engineers.
  • Develop high-fidelity analytical models and simulators to predict system behavior under emerging workloads.
  • Ensure the seamless scaling of models to clusters comprising hundreds of thousands of nodes.
  • Optimize AI and HPC libraries like NCCL, NVSHMEM, and UCX for trillion-parameter models.

What we're looking for

  • Ph.D. or M.S. in Computer Science, Electrical Engineering, or a related field (or equivalent experience).
  • 12+ years of industry experience in high-performance computing (HPC) or distributed deep learning.
  • Deep understanding of 3D parallelism and advanced strategies like Context Parallelism, Expert Parallelism, and ZeRO variants.
  • Technical proficiency with NCCL, UCX, UCC, NVSHMEM, or MPI.
  • Experience with RDMA, RoCE, and low-level InfiniBand verbs.
  • Advanced knowledge of high-throughput inference engines such as TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo.
  • Expert knowledge of the NVIDIA GPU memory hierarchy and CUDA programming models.

More like this

Similar roles

Senior Deep Learning Communication Architect

Nvidia

Santa Clara, CA +1 114 days ago $184,000$287,500
Deep Learning PyTorch TensorRT-LLM vLLM SGLang C++ Python CUDA OpenCL InfiniBand NVLink NCCL MPI UCX UCC NVSHMEM RoCE Triton FSDP
6+ yrs exp

Principal Architect, AI Networking

Nvidia

Remote (Santa Clara, CA) +1 141 days ago $272,000$431,250
RDMA NVLink GPUDirect InfiniBand RoCE NCCL UCX MPI NVSHMEM CUDA C C++ Rust Python vLLM SGLang TensorRT-LLM Distributed Training Computer Architecture
10+ yrs exp Remote