Senior Software Engineer, Capacity Management

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$200,000–$322,000 / yr
Employment
Full-time
Posted
8 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $189k
This role $261k
$112k most similar roles pay here $345k

This role pays more than 95% of similar roles. Most pay $160,381–$216,937 — the shaded band above. At the midpoint, this role pays about $261k versus about $189k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 892 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 870 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, Capacity Management

Senior Software Engineer, Capacity Management - DGX Cloud joins the team to design and build distributed services and data pipelines for capacity planning, allocation, reservations, and utilization. This role involves developing a unified model of GPU capacity across various cloud providers, regions, and clusters while automating workflows currently dependent on manual coordination. The engineer will build APIs, tools, and integrations to enable capacity-aware decisions, improve forecasting, and establish monitoring and data-quality controls for capacity systems. Candidates should possess experience in Python, Go, or Java, along with expertise in distributed systems, backend services, and data-processing pipelines. The role requires proficiency with Kubernetes, cloud infrastructure, and large-scale resource-management systems to solve the challenge of managing constrained GPU infrastructure while maximizing utilization and ensuring reliable customer experiences within the DGX Cloud environment.

What does a Software Engineer earn in California?

Median $214000 from 777 postings across 68 companies.

See salary data

What you'll do

  • Design and build distributed services and data pipelines for capacity planning, allocation, and reservations.
  • Develop a unified model of GPU capacity across multiple cloud providers, regions, and products.
  • Automate manual capacity-management workflows into scalable software and automated decision-making systems.
  • Build APIs and tools to enable other internal systems to make capacity-aware decisions.
  • Improve forecasting and scenario planning by integrating demand signals with infrastructure supply data.
  • Establish monitoring, data-quality controls, and service-level indicators for all capacity systems.
  • Diagnose complex production issues to improve the reliability and performance of management services.
  • Lead technical design reviews and mentor other engineers on engineering standards.

What we're looking for

  • BS degree or equivalent experience in Computer Science, Computer Engineering, or a related technical field.
  • 12+ years of software engineering experience building production systems.
  • Strong programming experience in languages such as Python, Go, Java, or similar.
  • Experience designing distributed systems, backend services, APIs, and data-processing pipelines.
  • Experience with cloud infrastructure, Kubernetes, compute platforms, or large-scale resource-management systems.
  • Strong understanding of data modeling, system integration, observability, and production operations.
  • Ability to translate ambiguous business requirements into clear technical designs and communicate across cross-functional teams.
  • Experience with GPU infrastructure, AI/ML platforms, capacity planning, or optimization techniques (preferred).

More like this

Similar roles

Senior Cloud Software Engineer

Nvidia

Remote 14 days ago $152,000–$241,500
Kubernetes AWS GCP Azure Go Python Rust C++ Java Distributed Systems Cloud-Native Data Management Storage Systems Performance Engineering Observability
5+ yrs exp Remote

Principal Software Engineer

Nvidia

Remote 52 days ago $272,000–$431,250
Kubernetes Go C Linux CI/CD GitLab Argo Flux Container Orchestration Distributed Systems Cloud Computing GPU DPU Confidential Computing
10+ yrs exp Remote

Senior Performance Engineer

Nvidia

Remote (Santa Clara, CA) +2 62 days ago $224,000–$356,500
C++ Python CUDA PyTorch JAX XLA GPU Computing Distributed Systems Performance Engineering Benchmarking Profiling Observability High-Performance Computing Data Analysis Automation Workflows
Remote