Staff Site Reliability Engineer

Okta Inc

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Bellevue, WAChicago, ILNew York, NYSan Francisco, CAWashington, DC
Salary
$194,000–$267,000 / yr
Posted
43 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $188k
This role $230k
$136k most similar roles pay here $281k

This role pays more than 85% of similar roles. Most pay $159,625–$217,225 — the shaded band above. At the midpoint, this role pays about $230k versus about $188k for comparable roles.

Based on 238 similar postings.

Employer

About Okta Inc

Okta, Inc. is an American identity and access management company based in San Francisco. It provides cloud software that helps companies manage and secure user authentication into applications, and for developers to build identity controls into applications, websites, web services, and devices.[

Okta Inc currently has 137 open roles on FindRole.

Listed pay typically runs $174,000–$239,000 across 137 roles with salary data.

Most-posted roles

View all roles at Okta Inc

At a glance

TL;DR · Staff Site Reliability Engineer

Staff Site Reliability Engineer - Splunk joins the Observability team to own and evolve the company's Splunk ecosystem. This role involves building a scalable, world-class Observability Platform by treating infrastructure as code to automate the deployment of agents and collectors across distributed systems. The engineer will manage log data collection, processing, and storage while optimizing for high reliability and low latency. Key responsibilities include participating in on-call rotations, leading post-incident reviews, and eliminating toil through automation. Required technical skills include proficiency in SPL, Go, Python, or Ruby, along with experience using Terraform for infrastructure management. The role requires deep knowledge of Linux internals, networking protocols like TCP/IP and DNS, and container orchestration via Kubernetes/EKS. Candidates must possess expertise in Splunk Cloud at scale, including Workload Management and HEC optimization to solve complex performance bottlenecks.

What does a Site Reliability Engineer earn in California?

Median $214000 from 54 postings across 15 companies.

See salary data

What you'll do

  • Design and maintain scalable observability infrastructure using Terraform as code.
  • Optimize log data collection, processing, and storage within the Splunk ecosystem.
  • Automate the deployment and scaling of observability agents and collectors to eliminate manual toil.
  • Create intuitive and actionable Splunk dashboards that correlate data across multiple sources.
  • Participate in on-call rotations and lead post-incident reviews to drive systemic improvements.
  • Develop internal tools and automate workflows using Go, Python, or Ruby.
  • Debug complex, cross-service performance bottlenecks using a data-driven approach.

What we're looking for

  • Candidates must have at least 5 years of experience in an SRE, DevOps, or Systems Engineering role focused on high-availability systems.
  • Candidates must have at least 5 years of experience scaling and managing Splunk Cloud at scale (1000+ SVCs).
  • Experience with Workload Management (WLM) and HEC optimization is required for Splunk management.
  • Expertise in creating actionable Splunk dashboards that correlate data across multiple sources is required.
  • Candidates must possess strong coding skills in SPL, Go, Python, or Ruby to build tools and automate workflows.
  • Proficiency in infrastructure as code using Terraform is required.
  • Deep understanding of Linux internals, networking (TCP/IP, DNS, Load Balancing), and container orchestration (Kubernetes/EKS) is required.
  • U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee.

More like this

Similar roles

Senior Staff Site Reliability Engineer

Cisco

Remote (Irvine, CA) +1 15 days ago $192,400$275,800
SRE Linux Administration AWS GCP Azure Python Go Distributed Systems Splunk SPL Indexer Clustering Search Head Clusters KVStore Monitoring Alerting Observability Root Cause Analysis
10+ yrs exp Remote

Staff Site Reliability Engineer, Kubernetes

Okta Inc

Bellevue, WA +4 15 days ago $194,000$267,000
Kubernetes AWS Helm Karpenter Istio Terraform CI/CD Python Bash Go Docker Prometheus Grafana CloudWatch ELK Stack S3 RDS EC2 IAM
Hybrid

Senior Site Reliability Engineer

Fiserv

Berkeley Heights, NJ 6 days ago $128,000$216,000
AWS Kubernetes Terraform CI/CD GitHub Actions Python Bash Ruby on Rails Docker Linux Unix RDBMS Document Storage New Relic Dynatrace Datadog DNS Load Balancing Virtual Networking

Staff Site Reliability Engineer

TransUnion

Chicago, IL +4 140 days ago $112,500$187,500
GCP Kubernetes CI/CD Datadog Prometheus Grafana PagerDuty Linux PostgreSQL MySQL Cloud SQL Bigtable Firestore Redis Terraform Pulumi Python Bash Go Infrastructure-as-Code
5+ yrs exp Hybrid

Staff Site Reliability Engineer

CME Group

Chicago, IL 42 days ago $132,100$220,100
Python Go Kubernetes GCP GKE Kafka Terraform ArgoCD Node.js Gemini Distributed Systems GitOps SRE
10+ yrs exp Hybrid

Staff Site Reliability Engineer

MongoDB

Bengaluru, India 52 days ago
Kubernetes Python Go AWS GCP Azure Multi-cloud Distributed Systems Virtual Machines Capacity Planning Incident Response SLO Observability Alerting Automation Networking Infrastructure Architecture
10+ yrs exp Hybrid