Principal Software Engineer, DGX Cloud Production Engineering

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$272,000–$431,250 / yr
Posted
31 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $203k
This role $352k
$108k most similar roles pay here $466k

This role pays more than 99% of similar roles. Most pay $174,600–$231,000 — the shaded band above. At the midpoint, this role pays about $352k versus about $203k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Principal Software Engineer, DGX Cloud Production Engineering

As a Principal Software Engineer, DGX Cloud Production Engineering, you will join the team building next-generation Kubernetes platforms for large-scale AI and computational environments. You will lead the architecture and development of core platform capabilities, including cluster management, control plane services, fleet lifecycle, and day-2 operations. Your daily work involves designing reliable distributed systems and APIs to automate provisioning, upgrading, and remediating clusters across cloud and on-premises environments. You will define technical requirements for declarative workflows while collaborating across teams to ensure consistency in runtime integration and infrastructure. The role requires deep expertise in Kubernetes internals, controllers, and operators, along with proficiency in Go, Python, Rust, or C++. You will solve complex problems involving networking, hardware, and scalability to support high-performance computing and accelerated computing platforms for large-scale AI deployments.

What does a Software Engineer earn in Remote?

Median $204500 from 415 postings across 57 companies.

See salary data

What you'll do

  • Lead the architecture and development of core Kubernetes platform capabilities including cluster management and control plane services.
  • Design and build highly reliable distributed systems and APIs for provisioning, managing, and upgrading clusters at scale.
  • Define technical requirements, validation criteria, and production-readiness practices for declarative workflows across the Kubernetes stack.
  • Develop automated solutions for fleet consistency, lifecycle orchestration, and runtime integration across cloud and on-premises environments.
  • Diagnose and resolve complex platform issues involving infrastructure, networking, hardware, and operations to support large-scale AI deployments.
  • Influence engineering standards, architectural decisions, and long-term strategy for the Kubernetes platform.
  • Mentor senior engineers and establish high standards for design quality and execution across the organization.

What we're looking for

  • BS or MS degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 15+ years of relevant software engineering experience building and operating large-scale production systems.
  • Deep expertise in Kubernetes internals, APIs, controllers, operators, and cluster lifecycle management.
  • Strong background in distributed systems design, reliability, scalability, and failure recovery.
  • Proven experience building platform software, infrastructure control planes, or foundations for managed services.
  • Strong programming skills in one or more systems or cloud-native languages like Go, Python, Rust, or C++.
  • Demonstrated ability to provide technical leadership across team boundaries and drive cross-functional initiatives.
  • Experience with fleet management, GitOps, or large-scale accelerated computing platforms (preferred).

More like this

Similar roles

Principal Software Engineer

Nvidia

Santa Clara, CA +1 135 days ago $272,000$431,250
Go Python Java Kubernetes Slurm Prometheus OpenTelemetry Grafana Docker AWS GCP Azure CUDA cuDNN Distributed Systems Infrastructure Automation Workflow Orchestration
10+ yrs exp

Principal Software Engineer

Nvidia

Remote 35 days ago $272,000$431,250
Kubernetes Go C Linux CI/CD GitLab Argo Flux Container Orchestration Distributed Systems Cloud Computing GPU DPU Confidential Computing
10+ yrs exp Remote

Principal Software Engineer, Distributed Systems Engineer

Nvidia

Remote (Durham, NC) 78 days ago $272,000$431,250
Kubernetes GPU Go Python Slurm Bright Cluster Manager Distributed Systems Cluster Management Monitoring Data Structures Algorithms Systems Programming Network Telemetry Incident Management
10+ yrs exp Remote

Manager, Software Engineering

Nvidia

Remote (Santa Clara, CA) +1 9 days ago $224,000$356,500
Kubernetes Golang Java C C++ Rust Docker Containerd CRI-O Linux Kernel DevOps Identity and Access Management Cloud Computing Infrastructure Networking Storage
10+ yrs exp Remote

Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale

Nvidia

Remote (Santa Clara, CA) +1 74 days ago $152,000$241,500
Kubernetes Python Golang Distributed Systems Containers CI/CD GPU Operator NVIDIA Software Stack AWS Azure GCP OCI Performance Modeling Benchmarking Confidential Containers CNCF open‑source Networking Storage Systems Computer Architecture
5+ yrs exp Remote