Senior System Architect, Infrastructure Reliability

Nvidia

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CAWestford, MAAustin, TXDurham, NCRedmond, WA
Salary
$184,000–$287,500 / yr
Posted
13 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $192k
This role $236k
$134k most similar roles pay here $304k

This role pays more than 86% of similar roles. Most pay $157,462–$226,350 — the shaded band above. At the midpoint, this role pays about $236k versus about $192k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior System Architect, Infrastructure Reliability

As a Senior System Architect, Infrastructure Reliability, you will join the team to address the challenge of failure attribution at scale for heterogeneous EDA systems. You will architect and build an automated framework that ingests telemetry from CPU and GPU clusters to identify root causes of job failures in real-time, distinguishing between hardware faults, infrastructure instability, and software defects. Your daily work involves building a high-fidelity flight recorder, implementing low-overhead tracing across Slurm or Kubernetes clusters, and developing machine learning models for automated root cause analysis. You will utilize C++, Python, and tools like DCGM and NVML to monitor system health. The role requires expertise in x86/ARM architectures, Linux kernel diagnostics, and memory fault handling to solve the problem of resource waste caused by failures across thousands of heterogeneous nodes.

What you'll do

  • Architect a scalable "flight recorder" to capture high-fidelity state across CPU, GPU, and Fabric during job failures.
  • Build automated diagnostics to correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions with system events.
  • Implement low-overhead tracing mechanisms for tracking job execution across multi-node Slurm or Kubernetes clusters.
  • Develop machine learning models and heuristics to automatically classify failures as hardware, software, or environment issues.
  • Define "signals of impending failure" to enable proactive job migration or checkpointing before a crash occurs.
  • Build high-performance daemons in C++ and Python to monitor system health without impacting workload performance.
  • Utilize NVIDIA DCGM and NVML to monitor device health and capture state-dumps for GPU infrastructure.

What we're looking for

  • BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience).
  • 6+ years of experience in systems programming.
  • Experience building automated Root Cause Analysis pipelines for HPC or cloud-scale environments.
  • Expert knowledge of x86/ARM node-level metrics including IPC, cache contention, NUMA imbalance, and hardware interrupts.
  • Proficiency in C++ and Python to build high-performance system monitoring daemons.
  • Familiarity with cluster resource managers such as Slurm, LSF, or Kubernetes.
  • Expert knowledge of the Linux kernel and its error-reporting interfaces like dmesg and journald.
  • Experience with NVIDIA DCGM, NVML, and checkpoint/restore technologies like CRIU.

More like this

Similar roles

Senior AI Solutions Architect, Semiconductors

Nvidia

Remote (Santa Clara, CA) 57 days ago $152,000$241,500
CUDA Python C/C++ PyTorch TensorFlow GPU Kubernetes HPC EDA cuLitho Distributed Systems GitHub PCIe FPGA Circuit Simulation Metrology
4+ yrs exp Remote

Senior Infrastructure Reliability Engineer

Anduril Industries

Costa Mesa, CA 8 days ago $166,000$220,000
Kubernetes Docker Terraform Python Go Bash AWS GCP Azure CI/CD Linux Ansible Puppet Chef Prometheus Grafana Datadog GitOps ArgoCD FluxCD VMware ESXi vSphere RKE2 Cilium RunAI Slurm Kubeflow Volcano

Senior Systems Reliability Engineer

The Walt Disney Company

Remote 59 days ago $141,900$190,300
infrastructure-as-code AWS Azure Terraform Ansible Python Go Ruby Swift Docker Kubernetes Jenkins GitLab CI/CD DataDog New Relic Grafana Active Directory LDAP Ping Identity VMWare KVM
5+ yrs exp Remote

Senior Director, Reliability Engineering

Nvidia

Santa Clara, CA 64 days ago $332,000$500,250
Reliability Engineering DfR DfX FMEA Physics of Failure Data Analytics Statistics Finite Element Analysis Simulation Risk Assessment AI Tools
10+ yrs exp

Senior Reliability Engineer

Anduril Industries

Atlanta, GA +1 79 days ago $143,000$191,000
FMEA Fault Tree Analysis Weibull Analysis MIL-STD-810 MIL-HDBK-217 MIL-HDBK-472 HALT HASS HITL SITL Root Cause Analysis CAPA Environmental Testing Mechanical Testing Data Acquisition Technical Writing Systems Engineering
8+ yrs exp

Senior Reliability Engineer

Anduril Industries

Costa Mesa, CA +1 79 days ago $166,000$220,000
FMEA Fault Tree Analysis Weibull Analysis MIL-STD-810 MIL-HDBK-217 MIL-HDBK-472 HALT HASS SITL HITL Root Cause Analysis CAPA Environmental Testing Mechanical Testing Data Acquisition Technical Writing
8+ yrs exp