Staff Site Reliability Engineer, AI Platform Runtime

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$168,000–$270,250 / yr
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $192k
This role $219k
$138k most similar roles pay here $284k

This role pays more than 72% of similar roles. Most pay $159,875–$224,250 — the shaded band above. At the midpoint, this role pays about $219k versus about $192k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Staff Site Reliability Engineer, AI Platform Runtime

Staff Site Reliability Engineer - AI Platform Runtime joins the SRE team to lead technical strategy and roadmaps for large-scale initiatives improving reliability, scalability, and developer productivity across enterprise systems. The role involves designing and building resilient distributed systems that power next-generation AI-driven products while architecting AI Agents and Skills to accelerate platform operations. Day-to-day responsibilities include driving automation and observability improvements using metrics and analytics, troubleshooting complex systems, and mentoring engineers to foster technical excellence. Candidates must possess expertise in Python, TypeScript, JavaScript, or Go, alongside experience with infrastructure-as-code tools like AWS CDK, CloudFormation, Terraform, or CrossPlane. The role requires deep knowledge of Kubernetes, networking, and public cloud services such as AWS, Azure, or GCP, while implementing OpenTelemetry for observability at scale to ensure high availability for AI infrastructure.

What does a Site Reliability Engineer earn in California?

Median $214000 from 54 postings across 15 companies.

See salary data

What you'll do

  • Lead the technical strategy and roadmap for large-scale SRE initiatives to improve reliability and developer productivity.
  • Design and build resilient distributed systems for AI-driven enterprise products and services.
  • Architect and develop AI Agents and Skills to accelerate platform operations.
  • Drive automation and observability improvements using metrics and analytics to enhance system performance.
  • Implement modern SRE components across Cloud, Platform, Security, and AI/ML teams to ensure high availability.
  • Analyze and troubleshoot complex systems while leading incident management and postmortem analysis.
  • Mentor engineers across different teams to foster technical excellence and a culture of reliability.

What we're looking for

  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
  • BS degree in Computer Science or a related technical field involving coding, or equivalent experience.
  • Strong proficiency in programming languages such as Python, TypeScript, JavaScript, or Go for automation and infrastructure-as-code.
  • Experience with infrastructure-as-code tools like AWS CDK, AWS CloudFormation, Terraform, or CrossPlane.
  • Solid understanding of OpenTelemetry or other observability implementations at scale.
  • Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP).
  • Outstanding problem-solving, communication, and teamwork skills to influence across technical and interpersonal boundaries.
  • Passion for and experience with Public Cloud or large-scale automation systems (preferred).

More like this

Similar roles

Principal Site Reliability Engineer

Nvidia

Santa Clara, CA 3 days ago $248,000$396,750
Kubernetes Distributed Systems Python Go Terraform AWS Azure GCP OpenTelemetry infrastructure-as-code Linux TypeScript JavaScript Java Crossplane AWS CDK CloudFormation AI/ML Platforms High-Performance Computing
10+ yrs exp Hybrid

Staff Site Reliability Engineer

TransUnion

Chicago, IL +4 140 days ago $112,500$187,500
GCP Kubernetes CI/CD Datadog Prometheus Grafana PagerDuty Linux PostgreSQL MySQL Cloud SQL Bigtable Firestore Redis Terraform Pulumi Python Bash Go Infrastructure-as-Code
5+ yrs exp Hybrid

Staff Site Reliability Engineer

CME Group

Chicago, IL 42 days ago $132,100$220,100
Python Go Kubernetes GCP GKE Kafka Terraform ArgoCD Node.js Gemini Distributed Systems GitOps SRE
10+ yrs exp Hybrid

Staff Site Reliability Engineer

MongoDB

Bengaluru, India 52 days ago
Kubernetes Python Go AWS GCP Azure Multi-cloud Distributed Systems Virtual Machines Capacity Planning Incident Response SLO Observability Alerting Automation Networking Infrastructure Architecture
10+ yrs exp Hybrid

Staff Site Reliability Engineer

Circle

Remote (San Francisco, CA) 37 days ago $195,000$257,500
Kubernetes Terraform Pulumi Go Python CI/CD Infrastructure as Code Blockchain Distributed Systems SQL Helm SRE Chaos Engineering Cloud Networking DNS
6+ yrs exp Remote

Staff Site Reliability Engineer

Anduril Industries

Costa Mesa, CA 31 days ago $191,000$253,000
SRE Kubernetes AWS GCP Azure Go Python Rust Prometheus Grafana OpenTelemetry Datadog Distributed Systems Canary Analysis Feature Flagging Capacity Planning Cost Optimization
10+ yrs exp