Senior Software Engineer, Distributed Systems Engineering

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CA
Salary
$152,000–$241,500 / yr
Employment
Full-time
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $186k
This role $197k
$122k $254k
below market most similar roles pay here above market

This role pays more than 63% of similar roles. Most pay $150,187–$222,000 — the blue band above. At the midpoint, this role pays about $197k versus about $186k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1472 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 1109 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, Distributed Systems Engineering

Senior Software Engineer, Distributed Systems Engineering - DGX Cloud joins the DGX Cloud team to scale AI infrastructure. This role involves developing production systems that enable large scalable GPU clusters for various AI workloads. Key responsibilities include building custom software for scheduling GPU resources on Kubernetes, implementing monitoring and health management capabilities, and harnessing data streams from GPU hardware diagnostics, cluster telemetry, and network telemetry. You will evaluate system failures and improve services using a defined incident management process. Required skills include software engineering experience with Kubernetes APIs, systems programming languages like Go or Python, and a solid understanding of data structures and algorithms. The work focuses on ensuring production AI clusters run reliably with maximum performance while managing large-scale distributed systems independent of cloud providers.

What does a Software Engineer earn in California?

Median $214500 from 1223 postings across 78 companies.

See salary data

What you'll do

  • Develop custom software for scheduling GPU resources on Kubernetes.
  • Implement monitoring and health management capabilities for GPU assets.
  • Harness multiple data streams including GPU hardware diagnostics and network telemetry.
  • Ensure production AI clusters run reliably and consistently with maximum performance.
  • Evaluate system failures and improve services based on incident management processes.
  • Develop and manage large-scale distributed systems independent of cloud providers.
  • Build and deploy infrastructure solutions for a broad range of AI-based applications.

What we're looking for

  • 5+ years of experience in a similar role and experience on large-scale production systems.
  • Direct experience in a software engineering role within a highly technical organization with demonstrable impact.
  • Software development experience with Kubernetes APIs and frameworks, including cluster operations and operator development.
  • Technical knowledge of a systems programming language such as Go or Python.
  • Solid understanding of data structures and algorithms.
  • BS in Computer Science, Engineering, Physics, Mathematics, or a comparable degree or equivalent experience.
  • Technical competency in managing and automating large-scale distributed systems independent of cloud providers (preferred).
  • Advanced hands-on experience and deep understanding of cluster management systems like Kubernetes, Slurm, or Bright Cluster Manager (preferred).

More like this

Similar roles

Senior Software Engineer, Distributed Systems Engineering

Nvidia

Remote (Santa Clara, CA) 9 days ago $184,000–$287,500
Kubernetes Go Python Distributed Systems Slurm Bright Cluster Manager Cluster Operations Operator Development Node Health Monitoring Data Structures Algorithms Systems Programming Network Telemetry Incident Management
8+ yrs exp Remote

Principal Software Engineer, Distributed Systems

Nvidia

Durham, NC 12 days ago
Kubernetes Go Python GPU Slurm Bright Cluster Manager Distributed Systems Data Structures Algorithms Systems Programming Cluster Management Monitoring Telemetry Incident Management
10+ yrs exp

Senior Software Engineer, DGX Cloud Production Engineering

Nvidia

Santa Clara, CA 18 days ago
Kubernetes Python Go GitOps Terraform ArgoCD GPU Infrastructure Linux Containers Cloud Infrastructure Infrastructure Automation Distributed Systems Observability BMaaS VMaaS Managed Kubernetes Multi-cloud APIs Incident Response
8+ yrs exp

Senior Software Engineer

Nvidia

Remote (Seattle, WA) +1 10 days ago
Kubernetes Go Linux Terraform Ansible CI/CD Gitlab Argo Flux Cloud Native Distributed Systems Observability Telemetry Container Orchestration Unix Data Structures and Algorithms
5+ yrs exp Remote

Senior Software Engineer, DGX Cloud Production Engineering

Nvidia

Remote (CA) +4 21 days ago
Kubernetes Python Go GitOps Terraform ArgoCD GPU Infrastructure Linux Containers Cloud Infrastructure Infrastructure Automation Distributed Systems Observability BMaaS VMaaS Managed Kubernetes Multi-cloud APIs Incident Response
8+ yrs exp Remote

Senior Software Engineer, Core Infrastructure Services

Nvidia

Remote (TX) +3 2 days ago $168,000–$270,250
Python Go Kubernetes Terraform Ansible FastAPI gRPC REST Temporal Redis Kafka NATS SQS Prometheus Grafana OpenTelemetry Linux BGP InfiniBand RDMA NetBox Nautobot DNS NTP RADIUS OAuth Firewalls iptables nftables
8+ yrs exp Remote