This role pays more than
99%
of similar roles. Most pay
$155,437–$213,939
— the shaded band above.
At the midpoint, this role pays about
$322k
versus about
$185k
for comparable roles.
Based on 238 similar postings.
Employer
About Nvidia
Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing
Nvidia currently has
896 open roles
on FindRole.
Listed pay typically runs
$184,000–$287,500
across 876 roles with salary data.
As a Principal Site Reliability Engineer, you will join the team to shape the technical direction and roadmap for reliability across the AI Platform Runtime and related enterprise systems. You will architect highly available, secure, and scalable distributed platforms while leading the development of AI agents and intelligent automation to accelerate platform operations and incident response. Your daily work involves establishing platform-wide standards like service-level objectives and capacity models, identifying systemic risks, and advancing observability through OpenTelemetry. The role requires deep expertise in Linux, Kubernetes, networking, and public cloud platforms such as AWS, Azure, or GCP. You will utilize programming languages including Python, Go, TypeScript, JavaScript, or Java, alongside infrastructure-as-code tools like Terraform and Crossplane to solve complex problems within the domain of large-scale AI-powered services and high-performance computing environments.
What does a Site Reliability Engineer earn in California?
Median $214000 from 54 postings across 15 companies.
Site Reliability Engineering
Kubernetes
Docker
Terraform
Ansible
Chef
Puppet
Python
Go
Java
JavaScript
Bash
Oracle Cloud Infrastructure
Microsoft Azure
Google Cloud Platform
infrastructure-as-code
Chaos Engineering
Site Reliability Engineering
Python
Java
Spring Boot
.Net
AI
CI/CD
Container Orchestration
AWS
Observability
Monitoring
Telemetry
Networking
System Architecture
SDLC
Site Reliability Engineering
Oracle Cloud Infrastructure
Linux
Windows Server
Python
PowerShell
Bash
Ansible
Chef
infrastructure-as-code
Networking
DNS
Firewalls
Load Balancing
Certificates
incident-management
Observability
Capacity Planning