Senior Systems Software Engineer, Kubernetes Node Lifecycle

Nvidia

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Santa Clara, CASeattle, WA
Salary
$184,000–$287,500 / yr
Posted
93 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $187k
This role $236k
$122k most similar roles pay here $305k

This role pays more than 87% of similar roles. Most pay $144,350–$228,950 — the shaded band above. At the midpoint, this role pays about $236k versus about $187k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Systems Software Engineer, Kubernetes Node Lifecycle

Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud is a technical role within the DGX Cloud division focused on managing the node layer of the NVIDIA Kubernetes Engine. The engineer will build and refine Cluster API providers, develop bring-your-own-node workflows for diverse hardware integration, and manage OS image generation, packaging, and hardening pipelines to meet security compliance. Day-to-day responsibilities include managing nodepool lifecycles at scale, resolving production faults related to kubelet operations or driver packaging, and collaborating with upstream communities like the CNCF. The role requires expertise in Golang, Python, and tools such as containerd, cloud-init, and packer. This position addresses the technical challenge of maintaining cluster reliability for frontier AI workloads by ensuring consistent node provisioning and automated testing across various Kubernetes versions and hardware configurations.

What you'll do

  • Build and refine Cluster API (CAPI) providers to ensure scalable node provisioning across cloud environments.
  • Develop "bring-your-own-node" workflows to integrate diverse hardware into Kubernetes clusters while maintaining operational consistency.
  • Manage the end-to-end lifecycle of OS images, including generation, packaging, deployment, and updates for GPU workloads.
  • Create and maintain automated test suites to validate node images across various Kubernetes versions and hardware configurations.
  • Implement security hardening pipelines for node images using CIS benchmarks and automated CVE remediation.
  • Manage large-scale nodepool lifecycles, including provisioning, upgrades, and seamless node replacement in production clusters.
  • Troubleshoot and resolve complex node-layer faults involving driver packaging, kubelet operations, and hardware activation.
  • Contribute to upstream projects like Cluster API and Kubernetes to establish industry standards for node provisioning.

What we're looking for

  • 8 years of experience in systems software, cloud infrastructure, or Kubernetes node engineering.
  • Bachelor’s or Master’s degree in Engineering (Electrical, Computer, or Science) or equivalent experience.
  • Deep expertise in Cluster API (CAPI), including provider development and full machine lifecycle management.
  • Extensive experience with OS image build pipelines, packaging, and delivery systems for Kubernetes nodes.
  • Practical experience with bring-your-own-node models and large-scale nodepool lifecycle management.
  • Strong understanding of kubelet configuration, node bootstrap, and the Kubernetes node registration lifecycle.
  • Experience with node image security, including vulnerability scanning, patch automation, and compliance gating.
  • Proficiency in Golang and/or Python, plus experience with at least one major public cloud provider.

More like this

Similar roles

Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale

Nvidia

Remote (Santa Clara, CA) +1 74 days ago $152,000$241,500
Kubernetes Python Golang Distributed Systems Containers CI/CD GPU Operator NVIDIA Software Stack AWS Azure GCP OCI Performance Modeling Benchmarking Confidential Containers CNCF open‑source Networking Storage Systems Computer Architecture
5+ yrs exp Remote

Senior Systems Software Engineer, Containers and Kubernetes

Nvidia

Remote (Santa Clara, CA) +2 9 days ago $184,000$287,500
Go C Kubernetes Container Orchestration Linux Distributed Systems Cloud Computing K8s Operator Framework Container Device Interface (CDI) Dynamic Resource Allocation (DRA) CNCF Data Structures Algorithms Systems Programming
8+ yrs exp Remote

Principal Software Engineer

Nvidia

Remote 35 days ago $272,000$431,250
Kubernetes Go C Linux CI/CD GitLab Argo Flux Container Orchestration Distributed Systems Cloud Computing GPU DPU Confidential Computing
10+ yrs exp Remote