Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CASeattle, WA
Salary
$184,000–$287,500 / yr
Posted
78 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $190k
This role $236k
$135k most similar roles pay here $304k

This role pays more than 88% of similar roles. Most pay $151,000–$228,950 — the shaded band above. At the midpoint, this role pays about $236k versus about $190k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale

Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud joins the DGX Cloud team to solve complex technical problems regarding the scaling of AI infrastructure while minimizing total cost of ownership. This role involves leading end-to-end performance and scalability analysis across the Kubernetes-based accelerated runtime stack, including components like GPU Operator and Network Operator. The engineer will design upstream architectural changes for the Kubernetes control plane, improve container startup latency, and advance the performance of confidential containers. Key responsibilities include using simulation infrastructure to model AI-factory deployments and collaborating with communities to develop automated workload tests. Required skills include expertise in distributed systems, Kubernetes, and the CNCF ecosystem, along with proficiency in Golang or Python. The role focuses on optimizing large-scale, parallel, distributed accelerator systems for training and inference workloads across thousands of GPU nodes.

What you'll do

  • Lead end-to-end performance and scalability analysis across the Kubernetes-based accelerated runtime stack from orchestration down to the metal.
  • Design and contribute upstream architectural changes to the Kubernetes control plane to enable reliable operation at hyperscale cluster sizes.
  • Improve container startup and cold-start latency to ensure low-latency inference scaling across thousands of GPU nodes.
  • Assess and improve open-source projects to optimize Kubernetes as a platform for large-scale AI training and inference.
  • Advance the scalability and performance of confidential containers (CoCo) to meet strict efficiency requirements in production.
  • Use simulation infrastructure to model full AI-factory deployments and validate scalability across thousands of simulated GPUs.
  • Develop automated, at-scale workload tests and integrate continuous performance testing into modern CI/CD workflows.
  • Document findings and present technical results at internal meetings and industry events like KubeCon and GTC.

What we're looking for

  • Bachelor’s or Master’s degree in Engineering, Electrical Engineering, Computer Engineering, or Computer Science.
  • 8+ years of experience in computer architecture, networking, storage systems, and accelerator-based platforms.
  • Expertise in Kubernetes and familiarity with the broader CNCF ecosystem.
  • Deep experience with large-scale, parallel, distributed accelerator systems and performance optimization of AI workloads.
  • Experience with performance modeling and benchmarking for large-scale systems.
  • Proficiency in Golang and/or Python.
  • Strong familiarity with the NVIDIA software stack across training and inference.
  • Expertise with at least one major public cloud provider such as AWS, Azure, GCP, or OCI.

More like this

Similar roles

Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale

Nvidia

Remote (Santa Clara, CA) +1 74 days ago $152,000$241,500
Kubernetes Python Golang Distributed Systems Containers CI/CD GPU Operator NVIDIA Software Stack AWS Azure GCP OCI Performance Modeling Benchmarking Confidential Containers CNCF open‑source Networking Storage Systems Computer Architecture
5+ yrs exp Remote

Senior Systems Software Engineer, Kubernetes Node Lifecycle

Nvidia

Santa Clara, CA +1 93 days ago $184,000$287,500
Kubernetes Cluster API Golang Python CI/CD NVIDIA Kubernetes Engine OS Image Packaging Cloud Infrastructure Containerd Cloud-init Packer Kubelet GCP AWS Azure OCI Flatcar Bottlerocket SBOM CIS Benchmarks
8+ yrs exp

Senior Systems Software Engineer, Containers and Kubernetes

Nvidia

Remote (Santa Clara, CA) +2 9 days ago $184,000$287,500
Go C Kubernetes Container Orchestration Linux Distributed Systems Cloud Computing K8s Operator Framework Container Device Interface (CDI) Dynamic Resource Allocation (DRA) CNCF Data Structures Algorithms Systems Programming
8+ yrs exp Remote

Principal Software Engineer

Nvidia

Remote 35 days ago $272,000$431,250
Kubernetes Go C Linux CI/CD GitLab Argo Flux Container Orchestration Distributed Systems Cloud Computing GPU DPU Confidential Computing
10+ yrs exp Remote

Principal Software Engineer, Distributed Systems Engineer

Nvidia

Remote (Durham, NC) 78 days ago $272,000$431,250
Kubernetes GPU Go Python Slurm Bright Cluster Manager Distributed Systems Cluster Management Monitoring Data Structures Algorithms Systems Programming Network Telemetry Incident Management
10+ yrs exp Remote