Principal Developer, AI Networking

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CATXCOWA
Salary
$272,000–$431,250 / yr
Posted
91 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $214k
This role $352k
$121k most similar roles pay here $465k

This role pays more than 97% of similar roles. Most pay $170,902–$256,500 — the shaded band above. At the midpoint, this role pays about $352k versus about $214k for comparable roles.

Based on 239 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Principal Developer, AI Networking

Principal Developer, AI Networking joins the AI Networking Codesign and Benchmarking R&D group to profile, analyze, and optimize AI workloads on large-scale GPU and CPU clusters for distributed Deep Learning LLM training and inference. The role focuses on collective communication and networking across hardware components like HCAs, Switches, CPUs, GPUs, and Systems, while engaging with software layers including machine learning frameworks and computing libraries. Key responsibilities include building performance analysis tools, defining test plans, and developing PyTorch trace-based profiling toolsets to identify bottlenecks in high-performance networking. The candidate will utilize Python, Bash, C++, CUDA, NCCL, RoCE, RDMA, MPI, and SHARP. This position addresses the technical challenges of distributed systems and large-scale LLM workloads by providing performance insights and co-designing network systems to meet specific performance targets for advanced AI infrastructure.

What you'll do

  • Profile, analyze, and optimize AI workloads on large-scale GPU and CPU clusters for distributed LLM training and inference.
  • Characterize deep learning models and workloads specifically aimed at high-performance networking and NVIDIA communication libraries.
  • Identify performance bottlenecks and areas for improvement within collective communications and network infrastructure.
  • Develop a PyTorch trace-based profiling, analysis, and replaying toolset to aid in benchmarking and debugging.
  • Define performance test plans and set expectations for new technologies to ensure they meet specific performance targets.
  • Build performance analysis tools and strategies to clarify system limitations and hardware/software interactions.
  • Analyze the performance of various components including HCAs, switches, CPUs, GPUs, and memory systems.

What we're looking for

  • Bachelor's degree in Computer Science, Software Engineering, or equivalent experience.
  • 15+ years of experience with high-performance networking including RDMA, MPI, NCCL, and SHARP.
  • Proficiency in programming languages including Python, Bash, and C++.
  • Experience with NVIDIA GPUs and the CUDA library.
  • Expertise in networking collective communication libraries such as NCCL and protocols like RoCE and RDMA.
  • Knowledge of deep learning frameworks such as TensorFlow or PyTorch.
  • Demonstrated ability in performance evaluation techniques and tools to identify bottlenecks.
  • Experience in container-based development environments.

More like this

Similar roles

Principal Architect, AI Networking

Nvidia

Remote (Santa Clara, CA) +1 141 days ago $272,000$431,250
RDMA NVLink GPUDirect InfiniBand RoCE NCCL UCX MPI NVSHMEM CUDA C C++ Rust Python vLLM SGLang TensorRT-LLM Distributed Training Computer Architecture
10+ yrs exp Remote

Principal Deep Learning Communication Architect

Nvidia

Remote (Santa Clara, CA) +1 150 days ago $272,000$431,250
NCCL UCX UCC NVSHMEM CUDA InfiniBand RDMA RoCE TensorRT-LLM vLLM SGLang Megatron-Core DeepSpeed JAX XLA PyTorch Distributed Ray HPC High-Performance Computing
10+ yrs exp Remote

Senior System Software Engineer, GPU Performance

Nvidia

Remote 6 days ago $152,000$241,500
HPC C++ Python CUDA NCCL UCX NVSHMEM MPI Infiniband Ethernet RDMA PyTorch TensorFlow Kubernetes Docker SLURM Ansible Performance Engineering Parallel Programming
Remote

Principal Software Engineer, AI Networking

Nvidia

Remote (Santa Clara, CA) 67 days ago $272,000$431,250
RoCE InfiniBand RDMA C C++ DOCA DPDK NCCL CUDA Networking Protocols Distributed Systems Firmware Performance Tuning QoS Telemetry Automation
10+ yrs exp Remote

Principal Network Developer

Oracle

Nashville, TN 66 days ago $102,300$209,500
Network Architecture Network Engineering Oracle Cloud Infrastructure (OCI) Scripting Automation Frameworks Configuration Management Telemetry SRE Network Hardware Firmware ASIC Data Center Networking Root Cause Analysis (RCA)
6+ yrs exp