Senior Staff Site Reliability Operations Technical Lead

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Durham, NC
Salary
$184,000–$264,500 / yr
Posted
4 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $183k
This role $224k
$132k most similar roles pay here $279k

This role pays more than 84% of similar roles. Most pay $151,862–$214,625 — the shaded band above. At the midpoint, this role pays about $224k versus about $183k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 929 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 915 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Staff Site Reliability Operations Technical Lead

Senior Staff Site Reliability Operations Technical Lead serves as a senior technical individual contributor and the primary point of last resort for site reliability and support. This role manages day-to-day operations including incident management, service quality, asset inventory, and vulnerability remediation while providing technical leadership to site support engineers. The position involves resolving complex issues across Active Directory, hybrid Entra ID, Exchange, database platforms, and compute infrastructure involving Windows, Linux, and macOS systems. Key responsibilities include building automation using PowerShell, Python, or Bash, performing root cause analysis, and managing endpoint compliance via Intune and Autopilot. The role addresses the technical challenges of maintaining stable enterprise infrastructure, ensuring security hardening, and coordinating with global teams to resolve recurring problems within a complex environment involving datacenter hardware, networking fundamentals like DNS and DHCP, and large-scale site expansions.

What you'll do

  • Serve as the Tier 3 escalation point for identity, messaging, compute, and endpoint management issues across the site and region.
  • Lead the site through major incidents by performing root cause analysis and driving permanent fixes for recurring problems.
  • Manage site assets and inventory throughout their lifecycle, including procurement, hardware refresh, and compliance audits.
  • Oversee endpoint security tasks such as vulnerability remediation, patch management, and hardening in partnership with InfoSec.
  • Provide technical leadership to site engineers by setting standards, reviewing work, and maintaining the knowledge base.
  • Develop automation scripts in PowerShell, Python, or Bash to improve diagnostics, reporting, and system health checks.
  • Represent local requirements and priorities in regional and global IT architecture and infrastructure forums.
  • Act as the primary technical liaison for executive leadership and internal departments during incidents and planned changes.

What we're looking for

  • 12+ years in enterprise support engineering, infrastructure, or end user services.
  • 5+ years in a senior, lead, or escalation-tier role in a multi-site environment.
  • Deep hands-on experience with Active Directory, hybrid Entra ID, Exchange hybrid, Windows/Linux servers, virtualization, and datacenter hardware.
  • Experience in enterprise endpoint management (Intune, Autopilot, MECM/SCCM, Jamf), M365 ecosystem, and vulnerability remediation.
  • Proficiency in networking fundamentals including DNS, DHCP, VLAN, wireless, firewall policy, and switch-level troubleshooting.
  • Ability to script and automate using Python, PowerShell, or Bash for diagnostics and reporting.
  • Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent experience.
  • Experience supporting engineering/lab environments, site buildouts, or executive support programs (preferred).

More like this

Similar roles

Senior Staff Site Reliability Operations

Nvidia

Seattle, WA 4 days ago $184,000$264,500
Active Directory Entra ID Exchange Microsoft 365 Intune Autopilot MECM SCCM Jamf PowerShell Python Bash Linux Windows macOS Virtualization DNS DHCP VLAN ServiceNow ITSM SLA Root Cause Analysis Patch Management
10+ yrs exp

Senior Site Reliability Engineer

Oracle

Nashville, TN 21 days ago $81,100$187,000
OCI AWS Azure GCP Terraform Chef Ansible Jenkins Docker CI/CD RESTful APIs Infrastructure-as-a-Service Log Analysis Agile Source Control Management
3+ yrs exp

Senior Staff Site Reliability Engineer

Cisco

Remote (Irvine, CA) +1 19 days ago $192,400$275,800
SRE Linux Administration AWS GCP Azure Python Go Distributed Systems Splunk SPL Indexer Clustering Search Head Clusters KVStore Monitoring Alerting Observability Root Cause Analysis
10+ yrs exp Remote

Senior Lead Site Reliability Engineer

JPMorgan Chase

Palo Alto, CA 63 days ago
Site Reliability Engineering Java Go Python Terraform Kubernetes Docker CI/CD GitOps Grafana Prometheus Dynatrace Datadog Splunk Kafka RabbitMQ SQS Neo4j Pinecone Weaviate Chroma LangChain LangGraph AutoGen CrewAI GitHub Copilot Fluentd Logstash Vector RESTful APIs RAG TensorFlow PyTorch scikit-learn Hadoop Spark Flink MongoDB Cassandra DynamoDB InfluxDB TimescaleDB AWS Azure GCP
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 49 days ago
Site Reliability Engineering Observability Monitoring Telemetry Service Level Objectives Alerting AI SDLC Automation
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 3 days ago
SRE DevOps Kubernetes AWS Terraform Python Bash Go CI/CD Spinnaker Harness EKS ECS AWS Lambda DynamoDB S3 Linux Distributed Systems Infrastructure as Code Observability
5+ yrs exp