Systems Quality and Reliability Engineer

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$136,000–$218,500 / yr
Posted
29 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $169k
This role $177k
$126k most similar roles pay here $228k

This role pays more than 57% of similar roles. Most pay $140,117–$198,125 — the shaded band above. At the midpoint, this role pays about $177k versus about $169k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Systems Quality and Reliability Engineer

The Systems Quality and Reliability Engineer - LPU joins the LPU team to own, build, and manage RMA and FA debug and root-cause analysis for existing and new AI/ML products. This role involves conducting tests, performing root-cause analysis of field RMAs, and collaborating with systems, hardware, software, and operations engineers. The engineer will scale fault analysis capabilities, create 8D reports, analyze repair data to identify trends, and manage operational performance at contract manufacturers regarding cycle times and fault isolation rates. Required skills include proficiency in Python, PERL, or C++ on UNIX/Linux, and experience with high-speed interfaces like SerDes, PCIe, and DDR. Candidates must be competent with oscilloscopes, logic analyzers, and power analyzers while possessing knowledge of FIB, SEM, TDR, VNA, CSAM, OBIRCH, DLS/LADA, LVP, and LVI techniques for hardware quality performance.

What you'll do

  • Own and manage RMA and FA debug and root-cause analysis for new and existing AI/ML products.
  • Conduct and lead technical investigations to identify the root causes of field failures.
  • Create formal failure analysis reports aligned with standard 8D processes or similar methodologies.
  • Analyze RMA, FA, and repair data to identify trends and issue quality alerts.
  • Develop resolution, containment, and mitigation plans for identified quality issues.
  • Monitor hardware performance metrics including RMA rates, MTBF, and Reliability Ratios.
  • Manage the operational performance of contract manufacturers regarding cycle times and fault isolation rates.
  • Oversee the integration of new products into Failure Analysis operations.

What we're looking for

  • BS/MS in EE, Physics, or a related degree (or equivalent experience).
  • 5+ years of hands-on systems test and/or validation engineering experience.
  • Proven hands-on experience as a systems quality and reliability engineer.
  • Competence using lab equipment such as oscilloscopes, logic analyzers, and power analyzers.
  • Experience enabling reliability tests like HTOL and quality tests like Burn-in.
  • Working knowledge of FA techniques and tools including FIB, SEM, TDR, VNA, and CSAM.
  • Strong knowledge of fault isolation techniques such as OBIRCH, DLS/LADA, LVP, and LVI.
  • Proficiency with high-speed interfaces (SerDes, PCIe, DDR) and programming languages like Python, PERL, or C++.

More like this

Similar roles

Systems Quality and Reliability Lead

Nvidia

Santa Clara, CA 67 days ago $168,000$264,500
Failure Analysis root-cause analysis Python Perl C++ Linux Unix Oscilloscopes Logic Analyzers Power Analyzers FIB SEM TDR VNA CSAM OBIRCH DLS LADA LVP LVI SerDes PCIe DDR HTOL Burn in
8+ yrs exp

Principal Systems Software Engineer

Nvidia

Santa Clara, CA 73 days ago $272,000$431,250
Rust Firmware Kernel Drivers RTOS RISC-V Linux gRPC MPI vLLM Kubernetes CI/CD Distributed Systems Hardware Bring-up Data Center Operations
10+ yrs exp

Senior System Level Test Engineer

Nvidia

Santa Clara, CA 67 days ago $196,000$310,500
SOC ASIC C# C/C++ Python Perl .NET Framework PCB Design Silicon Qualification HTOL Burn in Hardware Integration Network Topology Security Provisioning Fuse Programming
10+ yrs exp

Senior Systems Reliability Engineer

The Walt Disney Company

Remote 59 days ago $141,900$190,300
infrastructure-as-code AWS Azure Terraform Ansible Python Go Ruby Swift Docker Kubernetes Jenkins GitLab CI/CD DataDog New Relic Grafana Active Directory LDAP Ping Identity VMWare KVM
5+ yrs exp Remote

Senior Systems Reliability Engineer

T-Mobile

Overland Park, KS +2 35 days ago $98,500$177,700
Python SRE CICD Azure DevOps Power Platform Microsoft Graph APIs Ansible Docker Kubernetes Splunk AppDynamics Shell Perl C# Java C Microservices Cloud Native Agile
5+ yrs exp