Principal Software Engineer, Distributed Systems Engineer

Nvidia

Confirmed live today High trust
Remote

Quick summary

Work type
Remote
Location
Durham, NC
Salary
$272,000–$431,250 / yr
Employment
Full-time
Posted
3 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $208k
This role $352k
$114k most similar roles pay here $465k

This role pays more than 99% of similar roles. Most pay $174,600–$241,550 — the shaded band above. At the midpoint, this role pays about $352k versus about $208k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1150 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 912 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Principal Software Engineer, Distributed Systems Engineer

Principal Software Engineer, Distributed Systems Engineer - DGX Cloud joins the DGX Cloud team to develop production systems enabling large scalable GPU clusters for diverse AI workloads. This role involves building custom software for scheduling GPU resources on Kubernetes, implementing monitoring and health management capabilities, and processing data streams from hardware diagnostics and network telemetry. The engineer will ensure reliability, availability, and scalability of GPU assets while managing system failures through defined incident processes. Required skills include extensive experience with Kubernetes APIs, frameworks, and large-scale distributed systems. Candidates must possess proficiency in systems programming languages like Go or Python, along with a strong grasp of data structures and algorithms. The work focuses on the technical challenge of maintaining high-performance infrastructure for AI applications by managing complex cluster management systems such as Slurm or Bright Cluster Manager.

What you'll do

  • Develop custom software for scheduling GPU resources on Kubernetes clusters.
  • Implement monitoring and health management capabilities to ensure reliability of GPU assets.
  • Process multiple data streams including hardware diagnostics, cluster telemetry, and network data.
  • Maintain production AI clusters to ensure consistent performance and high availability.
  • Evaluate system failures and improve services through a defined incident management process.
  • Develop software using Kubernetes APIs and frameworks for large-scale distributed systems.
  • Automate the management of large-scale infrastructure independent of cloud providers.

What we're looking for

  • Experience in a software engineering role within a highly technical organization with demonstrable impact.
  • Software development experience with Kubernetes APIs and frameworks beyond basic cluster operations.
  • 15+ years of experience in similar roles and on large-scale production systems.
  • Proficiency in a systems programming language such as Go or Python.
  • Solid understanding of data structures and algorithms.
  • Bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or equivalent experience.
  • Technical competency in managing and automating large-scale distributed systems independent of cloud providers (preferred).
  • Advanced hands-on experience with cluster management systems like Kubernetes, Slurm, or Bright Cluster Manager (preferred).

More like this

Similar roles

Principal Software Engineer, Distributed Systems Engineer

Nvidia

Remote (Durham, NC) 99 days ago
Kubernetes GPU Go Python Slurm Bright Cluster Manager Distributed Systems Cluster Management Monitoring Data Structures Algorithms Systems Programming Network Telemetry Incident Management
10+ yrs exp Remote

Principal Software Engineer

Nvidia

Remote 56 days ago $272,000–$431,250
Kubernetes Go C Linux CI/CD GitLab Argo Flux Container Orchestration Distributed Systems Cloud Computing GPU DPU Confidential Computing
10+ yrs exp Remote