Systems Quality and Reliability Lead

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$168,000–$264,500 / yr
Posted
67 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $181k
This role $216k
$126k most similar roles pay here $279k

This role pays more than 79% of similar roles. Most pay $147,200–$214,000 — the shaded band above. At the midpoint, this role pays about $216k versus about $181k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Systems Quality and Reliability Lead

The Systems Quality and Reliability Lead - LPU joins the LPU team to own, build, and manage RMA and FA debug and root-cause analysis for existing and new AI/ML products. This role involves conducting tests, performing root-cause analysis of field RMAs, identifying quality trends, and driving mitigation plans while overseeing hardware quality performance metrics like MTBF and Reliability Ratio. The position requires managing operational performance at contract manufacturers to ensure key performance indicators are met during the setup of new products into failure analysis operations. Candidates should possess experience with lab equipment such as oscilloscopes and logic analyzers, along with knowledge of techniques like FIB, SEM, TDR, VNA, and CSAM. Required skills include proficiency in Python, PERL, or C++ on UNIX/Linux systems and expertise in high-speed interfaces including SerDes, PCIe, and DDR.

What you'll do

  • Own and manage RMA and failure analysis (FA) root-cause investigations for new and existing AI/ML products.
  • Conduct technical testing and root-cause analysis to identify hardware and system failures.
  • Create formal FA result reports aligned with standard 8D processes or similar quality standards.
  • Analyze RMA, FA, and repair data to identify trends and issue quality alerts.
  • Develop and execute resolution, containment, and mitigation plans for identified quality issues.
  • Monitor hardware performance metrics including RMA rates, MTBF, and Reliability Ratios.
  • Manage the operational performance of contract manufacturers (CMs) regarding FA cycle times and fault isolation.
  • Oversee the integration and setup of new products into Failure Analysis operations.

What we're looking for

  • BS/MS in EE, Physics, or a related degree (or equivalent experience).
  • 8+ years of hands-on systems test and/or validation engineering experience.
  • Proven hands-on management and leadership experience.
  • Competence using lab equipment such as oscilloscopes, logic analyzers, and power analyzers.
  • Experience with reliability tests like HTOL and quality tests like Burn in.
  • Working knowledge of FA techniques and tools including FIB, SEM, TDR, VNA, and CSAM.
  • Strong knowledge of fault isolation techniques such as OBIRCH, DLS/LADA, LVP, and LVI.
  • Proficiency with high-speed interfaces (SerDes, PCIe, DDR) and programming languages like Python, PERL, or C++.

More like this

Similar roles

Systems Quality and Reliability Engineer

Nvidia

Santa Clara, CA 29 days ago $136,000$218,500
Failure Analysis root-cause analysis Python Perl C++ Linux Unix Oscilloscopes Logic Analyzers Power Analyzers FIB SEM TDR VNA CSAM OBIRCH DLS/LADA LVP LVI SerDes PCIe DDR HTOL Burn in
5+ yrs exp Hybrid

Failure Analysis & System Validation Engineer

Amd

Austin, TX 51 days ago $130,400$195,600
Failure Analysis Post Silicon Validation PCIe Memory Validation Python Linux Windows Firmware Power Management Root Cause Analysis NPI Oscilloscopes Logic Analyzers GPU Architecture Data Center Infrastructure AI Tools
Hybrid

Principal Systems Software Engineer

Nvidia

Santa Clara, CA 73 days ago $272,000$431,250
Rust Firmware Kernel Drivers RTOS RISC-V Linux gRPC MPI vLLM Kubernetes CI/CD Distributed Systems Hardware Bring-up Data Center Operations
10+ yrs exp

Senior System Level Test Engineer

Nvidia

Santa Clara, CA 67 days ago $196,000$310,500
SOC ASIC C# C/C++ Python Perl .NET Framework PCB Design Silicon Qualification HTOL Burn in Hardware Integration Network Topology Security Provisioning Fuse Programming
10+ yrs exp

Senior SoC Subsystem and I/O Architect

Nvidia

Remote (CA) +2 36 days ago $184,000$287,500
SoC Architecture PCIe CXL NVLink UCIe AXI CHI NoC SystemC C++ Python Memory Systems Virtualization IOMMU Firmware Boot Power Management RAS DFT
8+ yrs exp Remote

System Performance Lead

Amd

Austin, TX 128 days ago $200,000$300,000
System Architecture C++ Python Perl Ruby Performance Modeling Power Management Technical Debug Validation Strategy SoC Data Collection Modeling Frameworks