Senior Site Reliability Engineer

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Employment
Full-time
Posted
15 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $190k
$139k most similar roles pay here $252k

This listing doesn't post a salary. Most similar roles pay $155,000–$225,250.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1150 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 912 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Site Reliability Engineer

Senior Site Reliability Engineer, DGX Cloud joins the production engineering team to build automation, tooling, and operational systems for large-scale GPU infrastructure supporting AI research and production workloads. The role involves building and operating automation for Kubernetes clusters across cloud partner and on-prem environments while developing tools for provisioning, validation, upgrades, monitoring, and repair. Key responsibilities include improving lifecycle workflows, defining SLOs and SLIs, reducing manual touches through GitOps and APIs, and participating in incident response. Candidates must possess strong programming skills in Python or Go, along with expert knowledge of Linux, Kubernetes, containers, and infrastructure automation. The role requires experience with observability stacks including Prometheus, Grafana, and the ELK Stack. Technical expertise may include Terraform, ArgoCD, and specialized AI inference workloads involving vLLM, PyTorch, TensorRT-LLM, CUDA, and NCCL.

What does a Site Reliability Engineer earn?

Median $187850 from 129 postings across 37 companies.

See salary data

What you'll do

  • Build and operate automation for large-scale Kubernetes clusters across cloud and on-prem environments.
  • Develop tools and services for provisioning, validation, upgrades, monitoring, and cluster lifecycle operations.
  • Improve Day 0, Day 1, and Day 2 workflows for cluster bringup and production handoff.
  • Define SLOs/SLIs, monitor error allowances, and streamline reporting for infrastructure reliability.
  • Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows.
  • Participate in on-call rotations, incident response, and debugging of distributed systems.
  • Build and operate comprehensive observability stacks including monitoring, logging, and tracing.

What we're looking for

  • 8+ years of experience building or operating production infrastructure.
  • Strong programming skills in Python, Go, or similar languages.
  • Expert-level knowledge with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
  • Solid grasp of SRE principles including SLOs, SLIs, error budgets, and incident management.
  • Ability to troubleshoot distributed systems in production environments.
  • Experience building and operating comprehensive observability stacks using tools like Prometheus, Grafana, or the ELK Stack.
  • BS/MS in Computer Science or equivalent experience.
  • Experience with GPU infrastructure, GitOps, Terraform, ArgoCD, or AI inference workloads (preferred).

More like this

Similar roles

NCX Senior Engineer

Nvidia

Remote (Santa Clara, CA) +1 41 days ago $184,000–$287,500
Kubernetes Python Go Prometheus Grafana InfiniBand RoCE CUDA Linux OpenTelemetry Shell Scripting DevOps Configuration Management Distributed Systems NVLink NVSwitch
8+ yrs exp Remote

Senior Network Reliability Engineer

Nvidia

Remote 5 days ago $136,000–$224,250
TCP/IP BGP OSPF MPLS IS-IS VxLAN EVPN QoS GRE IPsec DNS MACsec AWS Azure GCP OCI Arista Fortinet Juniper Python Shell Linux Prometheus Grafana Netbox Nautobot InfiniBand Cumulus OS
5+ yrs exp Remote

Senior Site Reliability Engineer

Oracle

Pleasanton, CA +2 101 days ago $81,100–$187,000
Terraform Chef Ansible Python Java Bash Kubernetes Helm Jenkins Grafana Prometheus CI/CD OCI DevOps
3+ yrs exp

Senior Site Reliability Engineer

Adobe

New York, NY 15 days ago $139,000–$257,550
AWS Kubernetes Python Terraform Docker CI/CD SageMaker Bedrock vLLM LangGraph PostgreSQL Aurora Memcached Ansible Chef Prometheus Grafana Splunk New Relic Fastly Node.js PHP Ruby Bash

Senior Site Reliability Engineer

The Federal Reserve

Boston, MA 37 days ago
AWS EKS Terraform Python Java Go Docker Ansible CI/CD IaC Linux Shell Scripting Prometheus Grafana CloudWatch OpenSearch Dynatrace Consul Vault S3 RDS Aurora Route 53 ELB ECR

Senior Site Reliability Engineer

Oracle

119 days ago $91,400–$187,000
OCI AWS Azure GCP Terraform Chef Ansible Jenkins Docker CI/CD RESTful APIs Infrastructure-as-a-Service Log Analysis Agile Source Control Management
3+ yrs exp