Senior Site Reliability Engineer, AIOPs

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$148,000–$235,750 / yr
Posted
122 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $182k
This role $192k
$133k most similar roles pay here $247k

This role pays more than 54% of similar roles. Most pay $149,580–$214,625 — the shaded band above. At the midpoint, this role pays about $192k versus about $182k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Site Reliability Engineer, AIOPs

As a Senior Site Reliability Engineer, AIOPs, you will join the team building an AI Data Center AIOps platform that transforms high-volume telemetry into reliable insights and automation for GPU fleets. You will be responsible for the platform's uptime, performance, data integrity, and safe change management rather than the compute cluster itself. Your daily work involves managing SLOs/SLIs, incident response, and postmortems for telemetry ingestion, processing, storage, and APIs. You will manage Kubernetes deployments using Helm and Terraform to ensure scalable environments while automating recurring checks and building comprehensive runbooks. The role requires expertise in Python, Bash, CI/CD pipelines, and infrastructure-as-code. You will also navigate complex distributed systems involving Kafka, Pulsar, Flink, Spark, ClickHouse, Elastic, and TSDBs to solve challenges related to backpressure, hotspots, and failure domains within the observability domain.

What does a Site Reliability Engineer earn in California?

Median $214000 from 54 postings across 15 companies.

See salary data

What you'll do

  • Monitor platform health via dashboards and logs while automating recurring checks to ensure reliability and resource efficiency.
  • Manage end-to-end Kubernetes deployments including runbooks, canary checks, post-deploy validation, and rollbacks.
  • Lead first-level incident triage by collecting diagnostics, identifying root causes, and providing actionable findings to engineering teams.
  • Own SLOs/SLIs for telemetry ingestion, processing, storage, and the associated APIs and dashboards.
  • Manage deployment infrastructure and packaging using Helm and Terraform to ensure scalable and reproducible environments.
  • Develop and maintain runbooks, SOPs, and checklists to drive continuous improvement through automation.
  • Translate platform signals into trustworthy alerts and automated actions in partnership with software and systems engineering teams.

What we're looking for

  • BS/MS in Computer Science, Computer Engineering, or equivalent experience.
  • 5+ years of experience operating production distributed systems as an SRE, DevOps, or Platform Ops engineer.
  • Proven ownership of reliability for observability or AIOps platforms including SLOs, SLIs, and incident response.
  • Deep experience with Kubernetes and containers for deploying, debugging, and scaling telemetry-heavy microservices.
  • Proficiency in automation using Python, Bash, CI/CD pipelines, and Infrastructure as Code (Terraform and Helm).
  • Strong Linux and networking fundamentals with experience in distributed systems and streaming stacks like Kafka or Pulsar.
  • Experience managing large-scale production deployments across multiple Kubernetes environments and clusters.
  • Ability to communicate clearly and translate ambiguous requirements into technical documentation and operational practices.

More like this

Similar roles

Senior Site Reliability Engineer

Okta Inc

Bellevue, WA +1 18 days ago $147,000$202,400
Terraform Kubernetes Spinnaker Flyway Snowflake CI/CD Infrastructure as Code Containerization SaaS AIOps Cloud Infrastructure Automation
Hybrid

Senior Site Reliability Engineer

The Federal Reserve

Boston, MA 15 days ago $140,000$210,900
AWS EKS Terraform Python Java Go Docker Ansible CI/CD IaC Linux Shell Scripting Prometheus Grafana CloudWatch OpenSearch Dynatrace Consul Vault S3 RDS Aurora Route 53 ELB ECR

Senior Site Reliability Engineer

Autodesk

Remote (ID) +1 28 days ago $117,000$209,330
Site Reliability Engineering AWS Kubernetes Python Go Java Infrastructure as Code CI/CD CloudWatch Splunk Datadog Dynatrace Bash PowerShell FedRAMP Distributed Systems Load Balancing DNS
7+ yrs exp Remote

Senior Site Reliability Engineer

Autodesk

San Francisco, CA 28 days ago $117,000$209,330
SRE Python Go Java Bash PowerShell AWS Kubernetes Infrastructure as Code CI/CD CloudWatch Splunk Datadog Dynatrace FedRAMP Distributed Systems Load Balancing DNS
7+ yrs exp

Senior Site Reliability Engineer

Salesforce

Remote (San Francisco, CA) 58 days ago $148,500$223,900
SRE Python Go Docker Kubernetes CI/CD Prometheus Grafana ELK Splunk Datadog Temporal Airflow Argo Workflows AWS GCP Linux Unix LLM Prompt Engineering
5+ yrs exp Remote

Senior Site Reliability Engineer

MongoDB

Gurugram, India 52 days ago
Kubernetes Python Go AWS Google Cloud Platform Azure Linux TCP/IP DNS TLS Istio Cilium Service Mesh Distributed Systems Multi-cloud Alerting Networking
6+ yrs exp Hybrid