Senior Software Engineer, Distributed Systems Engineering

Nvidia

Confirmed live today High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CA
Salary
$184,000–$287,500 / yr
Employment
Full-time
Posted
2 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $187k
This role $236k
$113k most similar roles pay here $306k

This role pays more than 91% of similar roles. Most pay $152,406–$222,000 — the shaded band above. At the midpoint, this role pays about $236k versus about $187k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1391 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 1116 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, Distributed Systems Engineering

Senior Software Engineer, Distributed Systems Engineering - DGX Cloud joins the DGX Cloud team to develop production systems for large scalable GPU clusters supporting diverse AI workloads. The role involves building custom software for scheduling GPU resources on Kubernetes, implementing monitoring and health management capabilities, and processing data streams from hardware diagnostics and network telemetry. You will evaluate system failures using incident management processes to ensure reliability and performance of AI infrastructure. Required skills include extensive experience with Kubernetes APIs, frameworks, and large-scale production systems. Technical requirements include proficiency in systems programming languages like Go or Python, along with a solid understanding of data structures and algorithms. The work focuses on the technical challenge of managing distributed systems and optimizing GPU assets for high-performance computing environments within the AI infrastructure domain.

What does a Software Engineer earn in California?

Median $214000 from 925 postings across 71 companies.

See salary data

What you'll do

  • Develop custom software to schedule GPU resources on Kubernetes clusters.
  • Implement monitoring and health management capabilities for high-availability GPU assets.
  • Analyze multiple data streams including hardware diagnostics, cluster telemetry, and network data.
  • Ensure production AI clusters run reliably with maximum performance across various workloads.
  • Evaluate system failures and improve services through a defined incident management process.
  • Develop software using Kubernetes APIs and frameworks rather than just operating the cluster.
  • Automate and manage large-scale distributed systems independent of cloud providers.

What we're looking for

  • Experience in a software engineering role within a highly technical organization with demonstrable impact.
  • 8+ years of experience in a similar role and on large-scale production systems.
  • Software development experience with Kubernetes APIs and frameworks, not just operating a cluster.
  • Proficiency in a systems programming language such as Go or Python.
  • Solid understanding of data structures and algorithms.
  • A BS in Computer Science, Engineering, Physics, Mathematics, or a comparable degree or equivalent experience.
  • Technical competency in managing and automating large-scale distributed systems independent of cloud providers (preferred).
  • Advanced hands-on experience with cluster management systems like Kubernetes, Slurm, or Bright Cluster Manager (preferred).

More like this

Similar roles

Principal Software Engineer, Distributed Systems Engineer

Nvidia

Remote (Durham, NC) 5 days ago $272,000–$431,250
Kubernetes GPU Go Python Slurm Bright Cluster Manager Distributed Systems Cluster Management Monitoring Data Structures Algorithms Systems Programming Network Telemetry Incident Management
10+ yrs exp Remote

Principal Software Engineer, Distributed Systems Engineer

Nvidia

Remote (Durham, NC) 101 days ago
Kubernetes GPU Go Python Slurm Bright Cluster Manager Distributed Systems Cluster Management Monitoring Data Structures Algorithms Systems Programming Network Telemetry Incident Management
10+ yrs exp Remote

Senior Software Engineer, DGX Cloud Production Engineering

Nvidia

Remote (Santa Clara, CA) 11 days ago $184,000–$287,500
Kubernetes Python Go GitOps Terraform ArgoCD Linux Cloud Infrastructure GPU Infrastructure Distributed Systems Observability Infrastructure Automation BMaaS VMaaS
8+ yrs exp Remote

Principal Software Engineer, DGX Cloud Production Engineering

Nvidia

Remote (Santa Clara, CA) 139 days ago $272,000–$431,250
Kubernetes Go Python GitOps Linux GPU Clusters AI/ML Infrastructure Distributed Systems Infrastructure Automation APIs Observability SLOs Multi-cloud BMaaS VMaaS High-Performance Computing
10+ yrs exp Remote