Staff Site Reliability Engineer, Ads

Reddit

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
San Francisco, CA
Salary
$217,000–$303,900 / yr
Posted
126 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $189k
This role $260k
$131k most similar roles pay here $322k

This role pays more than 94% of similar roles. Most pay $161,278–$216,250 — the shaded band above. At the midpoint, this role pays about $260k versus about $189k for comparable roles.

Based on 238 similar postings.

Employer

About Reddit

Reddit is a social news aggregation and discussion platform where users share content, vote on posts, and engage in community conversations across thousands of interest-based forums called subreddits.

Reddit currently has 77 open roles on FindRole.

Listed pay typically runs $217,000–$303,400 across 77 roles with salary data.

Most-posted roles

View all roles at Reddit

At a glance

TL;DR · Staff Site Reliability Engineer, Ads

Staff Site Reliability Engineer, Ads joins the Site Experience SRE team to lead reliability engineering initiatives for critical user-facing systems at internet scale. This technical leadership role involves partnering with product and infrastructure teams to improve availability, latency, scalability, and operational excellence across APIs, content delivery, feed generation, search, messaging, and real-time experiences. The successful candidate will architect for scale, reduce operational risk through proactive mitigation, drive automation to eliminate repetitive work, and lead complex incident response efforts. Key technical requirements include expertise in distributed systems, networking, Linux systems, and cloud native architectures. Candidates should possess strong programming skills in Go or Python, a deep understanding of observability systems like metrics and tracing, and experience with tools such as Kubernetes, Prometheus, Grafana, Envoy, Kafka, ClickHouse, Cassandra, and Redis to ensure high-performance delivery.

What does a Site Reliability Engineer earn in California?

Median $214000 from 54 postings across 15 companies.

See salary data

What you'll do

  • Drive reliability, scalability, and performance for critical user-facing systems including APIs, content delivery, and real-time experiences.
  • Design high-availability architectures involving failover, redundancy, traffic management, and capacity planning for massive global loads.
  • Identify systemic risks and bottlenecks to develop proactive mitigation strategies that reduce incidents and improve service health.
  • Develop automation tools and workflows to eliminate repetitive operational tasks and improve deployment safety.
  • Lead complex incident responses, conduct blameless postmortems, and implement sustainable long-term fixes.
  • Define and champion company-wide standards for SLIs, SLOs, capacity management, and operational maturity.
  • Provide technical leadership and mentorship to engineers to improve the organization's reliability culture.

What we're looking for

  • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large scale distributed systems.
  • Strong programming skills in languages such as Go, Python, or similar.
  • Deep understanding of distributed systems, networking, Linux systems, or cloud native architectures.
  • Experience designing highly available systems with strong operational and reliability practices.
  • Strong understanding of observability systems including metrics, logging, tracing, and alerting.
  • Experience improving reliability through SLOs, automation, incident management, and performance optimization.
  • Demonstrated ability to troubleshoot complex issues across applications, infrastructure, networking, and services.
  • Strong collaboration and communication skills with the ability to influence technical direction across teams.

More like this

Similar roles

Senior Site Reliability Engineer, Ads

Reddit

San Francisco, CA 14 days ago $190,800$267,100
Site Reliability Engineering Go Distributed Systems Cloud-native Architecture Monitoring Alerting Tracing Logging Incident Response Root Cause Analysis Automation Capacity Planning SLIs SLOs Infrastructure Engineering Networking
5+ yrs exp

Site Reliability Engineer

Balyasny Asset Management

Warsaw, Poland 78 days ago
Prometheus Grafana Loki Tempo OTEL Kubernetes Docker AWS Python Bash Go CI/CD DevOps SRE Agile
5+ yrs exp

Lead Principal Site Reliability Engineer

Oracle

Vienna, VA 53 days ago $96,300$264,100
Oracle Cloud Infrastructure (OCI) Kubernetes Terraform Python Bash Docker CI/CD Linux Administration Prometheus Grafana OpenSearch Splunk Datadog New Relic Infrastructure as Code Site Reliability Engineering Incident Management Root Cause Analysis Capacity Planning
10+ yrs exp

Staff Site Reliability Engineer

TransUnion

Chicago, IL +4 140 days ago $112,500$187,500
GCP Kubernetes CI/CD Datadog Prometheus Grafana PagerDuty Linux PostgreSQL MySQL Cloud SQL Bigtable Firestore Redis Terraform Pulumi Python Bash Go Infrastructure-as-Code
5+ yrs exp Hybrid

Staff Site Reliability Engineer

CME Group

Chicago, IL 42 days ago $132,100$220,100
Python Go Kubernetes GCP GKE Kafka Terraform ArgoCD Node.js Gemini Distributed Systems GitOps SRE
10+ yrs exp Hybrid

Staff Site Reliability Engineer

MongoDB

Bengaluru, India 52 days ago
Kubernetes Python Go AWS GCP Azure Multi-cloud Distributed Systems Virtual Machines Capacity Planning Incident Response SLO Observability Alerting Automation Networking Infrastructure Architecture
10+ yrs exp Hybrid