Principal Software Engineer, Distributed Systems Engineer

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Durham, NC
Salary
$272,000–$431,250 / yr
Posted
78 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $204k
This role $352k
$108k most similar roles pay here $466k

This role pays more than 99% of similar roles. Most pay $174,600–$232,850 — the shaded band above. At the midpoint, this role pays about $352k versus about $204k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Principal Software Engineer, Distributed Systems Engineer

Principal Software Engineer, Distributed Systems Engineer - DGX Cloud joins the DGX Cloud team to develop production systems for large scalable GPU clusters supporting diverse AI workloads. This role involves building custom software for scheduling GPU resources on Kubernetes, implementing monitoring and health management capabilities, and processing data streams from hardware diagnostics and network telemetry. The engineer will evaluate system failures and improve services using established incident management processes to ensure reliability and performance of AI infrastructure. Required skills include extensive experience with Kubernetes APIs, frameworks, and large-scale distributed systems. Candidates must possess proficiency in systems programming languages like Go or Python, along with a strong grasp of data structures and algorithms. The role focuses on the technical challenge of managing high-performance GPU assets and ensuring consistent operation across complex cluster environments for various AI applications.

What you'll do

  • Develop custom software to schedule GPU resources on Kubernetes clusters.
  • Implement monitoring and health management capabilities for high-availability GPU assets.
  • Analyze multiple data streams including hardware diagnostics, cluster telemetry, and network data.
  • Ensure production AI clusters run reliably with maximum performance across various workloads.
  • Evaluate system failures and improve services using a defined incident management process.
  • Develop software using Kubernetes APIs and frameworks to manage large-scale distributed systems.
  • Automate the management of large-scale infrastructure independent of specific cloud providers.

What we're looking for

  • Experience in a software engineering role within a highly technical organization with demonstrable impact.
  • 15+ years of experience in similar roles and on large-scale production systems.
  • Software development experience with Kubernetes APIs and frameworks beyond basic cluster operation.
  • Proficiency in systems programming languages such as Go or Python.
  • Solid understanding of data structures and algorithms.
  • Experience with common software engineering principles, tools, and techniques.
  • A BS in Computer Science, Engineering, Physics, Mathematics, or a comparable degree/equivalent experience.
  • Technical competency in managing and automating large-scale distributed systems independent of cloud providers.

More like this

Similar roles

Principal Software Engineer

Nvidia

Santa Clara, CA +1 135 days ago $272,000$431,250
Go Python Java Kubernetes Slurm Prometheus OpenTelemetry Grafana Docker AWS GCP Azure CUDA cuDNN Distributed Systems Infrastructure Automation Workflow Orchestration
10+ yrs exp

Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale

Nvidia

Remote (Santa Clara, CA) +1 74 days ago $152,000$241,500
Kubernetes Python Golang Distributed Systems Containers CI/CD GPU Operator NVIDIA Software Stack AWS Azure GCP OCI Performance Modeling Benchmarking Confidential Containers CNCF open‑source Networking Storage Systems Computer Architecture
5+ yrs exp Remote

Principal Software Engineer

Nvidia

Remote 35 days ago $272,000$431,250
Kubernetes Go C Linux CI/CD GitLab Argo Flux Container Orchestration Distributed Systems Cloud Computing GPU DPU Confidential Computing
10+ yrs exp Remote