Principal Software Engineer, At-Scale Reliability and Fleet Intelligence

Nvidia

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$272,000–$431,250 / yr
Posted
12 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $201k
This role $352k
$108k most similar roles pay here $466k

This role pays more than 99% of similar roles. Most pay $174,600–$228,187 — the shaded band above. At the midpoint, this role pays about $352k versus about $201k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Principal Software Engineer, At-Scale Reliability and Fleet Intelligence

Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements joins the CSP Engagements team as a technical focal point for fleet-scale reliability. This role involves collaborating with hyperscale customer engineering teams to ensure platforms achieve target MTBI in production environments. You will drive work streams regarding reliability software architecture, integrate fleet telemetry into improvement priorities, and validate lab improvements in real-world settings. Key responsibilities include performing statistical failure analysis using Pareto and Weibull methods, defining burn-in certification criteria, and developing predictive failure models from telemetry data. The role requires expertise in multi-NUMA rack-scale system software, firmware, and hardware failure modes across compute, interconnect, memory, power, and thermal domains. Required skills include experience with time-series databases, anomaly detection, health scoring, and the ability to translate complex statistical findings into actionable engineering priorities for hardware and software teams.

What does a Software Engineer earn in California?

Median $214000 from 775 postings across 63 companies.

See salary data

What you'll do

  • Drive reliability work streams with CSP engineering teams to align on MTBI measurement and failure classification.
  • Synthesize fleet telemetry data from multiple customers to identify and champion systemic hardware and software improvements.
  • Define consistent MTBI measurement methodologies that function across diverse customer monitoring environments and operational practices.
  • Perform statistical failure analysis using Pareto, survival, and Weibull methods to categorize system issues.
  • Design health monitoring architectures that integrate NVIDIA's telemetry and reporting with CSP automation workflows.
  • Establish burn-in reliability test environments and cluster certification criteria in collaboration with quality teams.
  • Develop predictive failure models based on fleet telemetry to proactively identify potential hardware failures.

What we're looking for

  • 15+ years of experience in systems software at datacenter scale or reliability engineering with a focus on at-scale challenges.
  • BS or MS in Computer Science, Electrical Engineering, Statistics, or a related field (or equivalent experience).
  • Deep expertise in multi-NUMA and rack-scale system software and firmware.
  • Proficiency in statistical failure analysis methods including MTBF/MTBI calculation, Pareto analysis, and root cause classification.
  • Experience with fleet-level telemetry and observability systems such as time-series databases and anomaly detection.
  • Understanding of hardware failure modes in large-scale GPU or accelerator deployments across compute, interconnect, memory, power, and thermal domains.
  • Experience defining or operating burn-in, stress testing, or certification frameworks for complex hardware systems.
  • Strong communication skills to present technical reliability findings to both technical audiences and executive leadership.

More like this

Similar roles

Principal Software Engineer, Rack-Scale System Software

Nvidia

Remote (Santa Clara, CA) +1 5 days ago $272,000$431,250
System Software Firmware rack-scale system S Fabric Management NVSwitch GPU Distributed Systems Telemetry APIs Cluster Management Orchestration Frameworks Error Handling Serviceability Power Management InfiniBand NVLink
10+ yrs exp Remote

Principal Software Engineer, CSP Engagements

Nvidia

Santa Clara, CA 6 days ago $272,000$431,250
System Software Architecture x86 ARM aarch64 Linux Kernel Device Drivers CUDA GPU Computing CXL Memory Fabric Performance Analysis High-Performance Computing Hardware/Software Interface
10+ yrs exp

Principal Software Engineer, GPU Firmware and GPU System Software

Nvidia

Remote 5 days ago $272,000$431,250
GPU Firmware GPU System Software VBIOS InfoROM NVLink Firmware Update Orchestration Multi-tenancy Isolation Secure Boot Attestation ECC Power Management Telemetry Xid Errors Driver Stack Fleet Management Rollback Strategy
10+ yrs exp Remote

Senior Site Reliability Engineer, Fleet Management

MongoDB

Remote (Austin, TX) +6 101 days ago $127,000$249,000
Kubernetes Go Python Terraform AWS GCP Azure Crossplane Helm Kustomize Gatekeeper Kyverno CRDs Operators Linux TCP/IP DNS TLS MongoDB
6+ yrs exp Remote