Senior Manager, GPU Cloud Infrastructure

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$256,000–$414,000 / yr
Posted
157 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $223k
This role $335k
$147k most similar roles pay here $443k

This role pays more than 95% of similar roles. Most pay $195,150–$250,725 — the shaded band above. At the midpoint, this role pays about $335k versus about $223k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Manager, GPU Cloud Infrastructure

As a Senior Manager, GPU Cloud Infrastructure - GeForce NOW, you will lead and mentor a specialized team of network architects focused on high-performance GPU infrastructure. You will oversee the design of intra-cluster and inter-cluster connectivity while driving technical tuning to reduce latency and jitter for cloud gaming, AI/ML training, and real-time inference workloads. Your daily responsibilities include managing RoCE, Ethernet-based AI fabrics, and high-bandwidth data center interconnects, as well as implementing Infrastructure as Code via Ansible or Terraform. You will utilize tools like Prometheus and Grafana for observability while working with BGP, EVPN/VXLAN, and SR-IOV technologies. The role addresses the critical challenge of delivering ultra-low-latency, high-throughput networking across data centers to support demanding interactive entertainment and large-scale AI platforms through robust infrastructure automation and proactive fault tolerance strategies.

What you'll do

  • Build and mentor a specialized team of network architects focused on high-performance GPU infrastructure.
  • Oversee the design of intra-cluster and inter-cluster connectivity using RoCE, Ethernet-based AI fabrics, and high-bandwidth interconnects.
  • Perform technical tuning to reduce latency and jitter while implementing congestion control and packet-loss mitigation.
  • Define the networking roadmap for gaming, AI/ML training, and real-time inference at scale.
  • Engage with ISPs to optimize low-latency edge networks for seamless connections between data centers and end clients.
  • Implement Infrastructure as Code (IaC) and observability frameworks to automate provisioning and monitor cluster health.
  • Collaborate with hardware vendors and SRE groups to influence technology direction and vendor selection.
  • Establish fault tolerance protocols and lead incident response and root cause analysis for complex network issues.

What we're looking for

  • Experience in networking, cloud infrastructure, or distributed systems.
  • 5+ years of experience managing technical teams.
  • Mastery of data center networking including Clos/spine-leaf architectures and high-performance fabrics like RDMA, RoCE, or InfiniBand.
  • Hands-on experience with BGP, EVPN/VXLAN, and kernel-level development for routing and switching.
  • Proficiency in infrastructure automation using Ansible or Terraform and monitoring tools like Prometheus and Grafana.
  • Experience designing for large-scale configurations using SR-IOV, Xen virtualization, or Open Virtual Switch.
  • Bachelor’s or Master’s degree in Computer Science or a related engineering field (or equivalent experience).
  • Ability to ensure infrastructure meets internal policies and regulatory standards such as GDPR.

More like this

Similar roles

Senior Software Engineer, GeForce NOW Video Streaming

Nvidia

Santa Clara, CA 53 days ago $184,000$287,500
C C++ CUDA Vulkan DirectX OpenGL H.264 HEVC AV1 NvMedia Multi-threaded Programming Video Compression adaptive bit-rate Kernel-mode Drivers Telemetry Performance Monitoring DLSS RTX
6+ yrs exp

Senior Manager, Systems Software Engineering

Nvidia

Santa Clara, CA 53 days ago $272,000$431,250
Cloud Services SaaS Distributed Systems Software Engineering Agile Scalability Availability Maintainability Technical Leadership
10+ yrs exp

Senior System Software Engineer, Cloud

Nvidia

Santa Clara, CA 112 days ago $224,000$356,500
Kubernetes Golang C++ Python Java Docker Terraform Pulumi Helm Kustomize EKS GKE AKS gRPC Microservices Infrastructure as Code Observability
10+ yrs exp

Senior Systems Software Engineer

Nvidia

Santa Clara, CA 11 days ago $224,000$356,500
Distributed Systems Cloud Computing Java Golang Python Kubernetes Spring Boot Microservices gRPC Cassandra Redis Infrastructure as Code ECS OpenStack NoSQL SaaS PaaS Monitoring
10+ yrs exp