Principal Software Engineer, DGX Cloud Production Engineering

Nvidia

Confirmed live yesterday Trusted
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$272,000–$431,250 / yr
Posted
116 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $203k
This role $352k
$108k most similar roles pay here $466k

This role pays more than 99% of similar roles. Most pay $174,600–$231,000 — the shaded band above. At the midpoint, this role pays about $352k versus about $203k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Principal Software Engineer, DGX Cloud Production Engineering

As a Principal Software Engineer, DGX Cloud Production Engineering, you will join the team responsible for scaling GPU infrastructure across internal, partner, and cloud environments. You will define technical strategies for cluster operations while building automation, GitOps, and Day 2 reliability systems to manage large-scale GPU clusters in both on-prem and partner settings. Your daily work involves designing systems for lifecycle management, validation, repair, upgrades, observability, and readiness while eliminating operational toil through software, APIs, and agent-assisted workflows. You will establish technical standards for production readiness, SLOs, and incident response while mentoring other engineers. The role requires expertise in Kubernetes, Linux, infrastructure automation, and programming in Go or Python. This position focuses on solving complex infrastructure problems by creating durable software and operating models specifically for high-performance GPU cluster management and multi-cloud fleet operations.

What does a Software Engineer earn in Remote?

Median $204500 from 415 postings across 57 companies.

See salary data

What you'll do

  • Define and execute the technical strategy for DGX Cloud cluster operations across partner and on-prem environments.
  • Build automation, GitOps workflows, and reliability systems to manage large-scale GPU clusters.
  • Design and implement systems for cluster lifecycle management, including validation, repair, upgrades, and observability.
  • Establish standardized patterns for Kubernetes-based GPU cluster operations.
  • Eliminate operational toil by developing software, APIs, and agent-assisted automated workflows.
  • Set technical standards for production readiness, SLOs, incident response, and operational acceptance.
  • Mentor engineers and influence cross-functional teams across platform, infrastructure, storage, networking, and security.

What we're looking for

  • 15+ years of experience building and operating large-scale distributed systems or cloud infrastructure.
  • Deep experience with Kubernetes, Linux, infrastructure automation, and production operations.
  • Strong programming experience in Go, Python, or similar languages.
  • Proven ability to lead complex cross-org technical initiatives.
  • Experience designing reliable systems with clear SLOs, observability, incident response, and automation.
  • BS/MS in Computer Science or equivalent experience.
  • Preferred experience with GPU clusters, AI/ML infrastructure, GitOps, or multi-cloud fleet operations.
  • Track record of turning operational pain into reusable software, APIs, and engineering standards.

More like this

Similar roles

Principal Software Engineer, Distributed Systems Engineer

Nvidia

Remote (Durham, NC) 78 days ago $272,000$431,250
Kubernetes GPU Go Python Slurm Bright Cluster Manager Distributed Systems Cluster Management Monitoring Data Structures Algorithms Systems Programming Network Telemetry Incident Management
10+ yrs exp Remote

Principal Software Engineer

Nvidia

Santa Clara, CA +1 135 days ago $272,000$431,250
Go Python Java Kubernetes Slurm Prometheus OpenTelemetry Grafana Docker AWS GCP Azure CUDA cuDNN Distributed Systems Infrastructure Automation Workflow Orchestration
10+ yrs exp

Principal Software Engineer

Nvidia

Remote 35 days ago $272,000$431,250
Kubernetes Go C Linux CI/CD GitLab Argo Flux Container Orchestration Distributed Systems Cloud Computing GPU DPU Confidential Computing
10+ yrs exp Remote

DGX Cloud Automation Engineer

Nvidia

Santa Clara, CA 33 days ago $184,000$287,500
Golang Kubernetes KubeVirt Docker AWS CI/CD Infrastructure as Code Bare Metal PXE Boot DHCP DNS Firecracker KVM OpenStack Nutanix AHV Redhat OpenShift PaaS IaaS
8+ yrs exp

Manager, Software Engineering

Nvidia

Remote (Santa Clara, CA) +1 9 days ago $224,000$356,500
Kubernetes Golang Java C C++ Rust Docker Containerd CRI-O Linux Kernel DevOps Identity and Access Management Cloud Computing Infrastructure Networking Storage
10+ yrs exp Remote