Staff Site Reliability Operations

Nvidia

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Hillsboro, OR
Salary
$144,000–$230,000 / yr
Employment
Full-time
Posted
10 days ago
Freshness
Confirmed live today

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $182k
This role $187k
$134k $240k
below market most similar roles pay here above market

This role pays less than 55% of similar roles. Most pay $145,375–$219,131 — the blue band above. At the midpoint, this role pays about $187k versus about $182k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1472 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 1109 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Staff Site Reliability Operations

The Staff Site Reliability Operations role serves as a senior technical individual contributor and technical lead for site reliability and support. You will own local technical service delivery, acting as the final escalation point for complex issues across Active Directory, Exchange, database platforms, and compute infrastructure. Daily responsibilities include managing incidents, ensuring SLA attainment, overseeing asset lifecycles, and leading site-level projects. You will utilize PowerShell, Python, or Bash to build automation for diagnostics and reporting while managing endpoints via Intune, Autopilot, and M365. The role requires deep expertise in hybrid Entra ID, virtualization, storage, and datacenter hardware to solve recurring problems. You will provide technical leadership to engineers, drive root cause analysis, and represent local requirements in global architecture forums to ensure infrastructure stability.

What you'll do

  • Own day-to-day site operations including incident management, SLA attainment, and service quality.
  • Serve as the Tier 3 escalation point for identity, messaging, compute, and endpoint infrastructure.
  • Lead root cause analysis and drive permanent fixes for recurring technical problems.
  • Manage endpoint compliance, vulnerability remediation, patch management, and audit readiness.
  • Provide technical leadership to site engineers by setting standards and mentoring on diagnostic rigor.
  • Build automation scripts in PowerShell, Python, or Bash for diagnostics and reporting.
  • Represent site priorities in regional and global IT architecture and standards forums.
  • Manage site assets, inventory, and procurement across the hardware lifecycle.

What we're looking for

  • 8+ years in enterprise support engineering, infrastructure, or end user services.
  • 5+ years in a senior, lead, or escalation-tier role in a multi-site environment.
  • Bachelor's degree in Computer Science, Information Systems, or related field, or equivalent experience.
  • Deep hands-on experience with Active Directory, hybrid Entra ID, Exchange hybrid, Windows/Linux servers, virtualization, storage, and datacenter hardware.
  • Proficiency in enterprise endpoint management (Intune, Autopilot, MECM/SCCM, Jamf), Windows 11, and the Microsoft 365 ecosystem.
  • Knowledge of database operations and networking fundamentals including DNS, DHCP, VLAN, wireless, firewall policy, and switch-level troubleshooting.
  • Scripting and automation skills in Python, PowerShell, or Bash, and experience with ServiceNow or similar ITSM tools.
  • Experience supporting engineering, lab, R&D, or manufacturing environments with specialized equipment (preferred).

More like this

Similar roles

Senior Staff Site Reliability Operations

Nvidia

Seattle, WA 30 days ago $184,000–$264,500
Active Directory Entra ID Exchange Hybrid Linux Virtualization Microsoft 365 Intune Autopilot MECM SCCM Jamf Python PowerShell Bash DNS DHCP VLAN ServiceNow ITSM Root Cause Analysis Storage Kerberos LDAP SMTP Firewall
10+ yrs exp

Senior Staff Site Reliability Operations Technical Lead

Nvidia

Durham, NC 30 days ago $184,000–$264,500
Active Directory Entra ID Exchange Hybrid Linux Virtualization Microsoft 365 Intune Autopilot MECM SCCM Jamf Python PowerShell Bash ServiceNow DNS DHCP VLAN Firewall Storage Root Cause Analysis ITSM
10+ yrs exp

Staff Site Reliability Engineer, AI Platform Runtime

Nvidia

Santa Clara, CA 33 days ago $168,000–$270,250
Kubernetes Python Go Terraform AWS Azure GCP infrastructure-as-code OpenTelemetry TypeScript JavaScript AWS CDK AWS CloudFormation CrossPlane Distributed Systems Observability Networking AI/ML High-Performance Computing
10+ yrs exp Hybrid

Staff Site Reliability Engineer

CME Group

Chicago, IL 72 days ago $132,100–$220,100
Python Go Kubernetes GCP Kafka Terraform ArgoCD Node.js GitOps Distributed Systems Gemini SLOs SLIs IAM GKE
10+ yrs exp Hybrid

Staff Site Reliability Engineer

Circle

Remote (San Francisco, CA) 67 days ago $195,000–$257,500
Kubernetes Terraform Pulumi Python Go CI/CD Infrastructure as Code Blockchain Distributed Systems SQL Helm Observability Chaos Engineering DNS Load Balancers Canary Releases MCP Servers
6+ yrs exp Remote

Staff Site Reliability Engineer

Anduril Industries

Costa Mesa, CA 17 days ago $191,000–$253,000
Site Reliability Engineering Kubernetes AWS GCP Azure Python Go Rust Prometheus Grafana OpenTelemetry Datadog Distributed Systems Canary Analysis Networking Storage
10+ yrs exp