Principal Site Reliability Engineer

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$248,000–$396,750 / yr
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $185k
This role $322k
$112k most similar roles pay here $427k

This role pays more than 99% of similar roles. Most pay $155,437–$213,939 — the shaded band above. At the midpoint, this role pays about $322k versus about $185k for comparable roles.

Based on 238 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Principal Site Reliability Engineer

As a Principal Site Reliability Engineer, you will join the team to shape the technical direction and roadmap for reliability across the AI Platform Runtime and related enterprise systems. You will architect highly available, secure, and scalable distributed platforms while leading the development of AI agents and intelligent automation to accelerate platform operations and incident response. Your daily work involves establishing platform-wide standards like service-level objectives and capacity models, identifying systemic risks, and advancing observability through OpenTelemetry. The role requires deep expertise in Linux, Kubernetes, networking, and public cloud platforms such as AWS, Azure, or GCP. You will utilize programming languages including Python, Go, TypeScript, JavaScript, or Java, alongside infrastructure-as-code tools like Terraform and Crossplane to solve complex problems within the domain of large-scale AI-powered services and high-performance computing environments.

What does a Site Reliability Engineer earn in California?

Median $214000 from 54 postings across 15 companies.

See salary data

What you'll do

  • Define and drive the long-term technical vision, architecture, and roadmap for reliability across the AI Platform Runtime.
  • Architect highly available, resilient, secure, and scalable distributed platforms to power next-generation AI products.
  • Design and develop AI agents and intelligent automation to accelerate platform operations, incident response, and troubleshooting.
  • Establish platform-wide reliability standards including service-level objectives, error budgets, capacity models, and resilience patterns.
  • Advance observability across complex distributed environments using OpenTelemetry, metrics, logs, traces, and automated anomaly detection.
  • Lead cross-functional programs to identify systemic risks and improve availability, performance, security, and developer productivity.
  • Provide technical leadership during critical incidents and translate post-mortem lessons into durable engineering improvements.
  • Develop reference architectures and automation frameworks that can be adopted across multiple engineering organizations.

What we're looking for

  • 15+ years of experience in Site Reliability Engineering, Platform Engineering, Distributed Systems, Cloud Architecture, or related roles.
  • BS or MS degree in Computer Science or a related technical field involving significant software development, or equivalent experience.
  • Demonstrated experience setting technical strategy and leading large-scale engineering initiatives across multiple teams or organizations.
  • Deep expertise in distributed systems architecture, networking, Linux, Kubernetes, and public cloud platforms like AWS, Azure, or GCP.
  • Strong proficiency in Python, Go, TypeScript, JavaScript, or Java for building production-grade automation and platform software.
  • Extensive experience with infrastructure-as-code and platform automation tools such as Terraform, Crossplane, AWS CDK, or CloudFormation.
  • Deep understanding of observability at scale including OpenTelemetry, metrics, logging, tracing, profiling, and analytics.
  • Expertise in reliability engineering practices like SLOs, error budgets, capacity planning, fault tolerance, and incident management.
  • Experience designing/operating AI/ML platforms, GPU infrastructure, or high-performance computing (preferred).
  • Hands-on experience building AI agents or intelligent automation for infrastructure operations (preferred).
  • Success establishing reliability architecture and engineering standards adopted across a large enterprise (preferred).
  • Experience applying machine learning to capacity management, anomaly detection, or autonomous remediation (preferred).
  • Recognized technical leadership through patents, publications, open-source contributions, or conference presentations (preferred).

More like this

Similar roles

Lead Principal Site Reliability Engineer

Oracle

Nashville, TN 57 days ago $96,300$264,100
Site Reliability Engineering Kubernetes Docker Terraform Ansible Chef Puppet Python Go Java JavaScript Bash Oracle Cloud Infrastructure Microsoft Azure Google Cloud Platform infrastructure-as-code Chaos Engineering
6+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 45 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Staff Site Reliability Engineer, AI Platform Runtime

Nvidia

Santa Clara, CA 3 days ago $168,000$270,250
Site Reliability Engineering Kubernetes Python Go TypeScript JavaScript Terraform AWS CDK CloudFormation CrossPlane OpenTelemetry infrastructure-as-code Observability Distributed Systems Networking AWS Azure GCP
10+ yrs exp Hybrid

Principal Site Reliability Engineer

Oracle

Nashville, TN 57 days ago $84,900$209,500
Site Reliability Engineering Oracle Cloud Infrastructure Linux Windows Server Python PowerShell Bash Ansible Chef infrastructure-as-code Networking DNS Firewalls Load Balancing Certificates incident-management Observability Capacity Planning
3+ yrs exp

Principal Site Reliability Engineer

Oracle

Reston, VA +1 45 days ago $84,900$209,500
Kubernetes Terraform Docker Python Bash Linux Unix Oracle Database RAC Chef Puppet DNS DHCP HTTP TCP/IP LLM VMware Cisco
6+ yrs exp