Senior Staff Engineer, SRE, Incident Prevention / Post Incident Correction of Errors

GEICO

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Bethesda, MDPalo Alto, CARichardson, TXSeattle, WA
Salary
$110,000–$260,000 / yr
Employment
Full-time
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $191k
This role $185k
$92k most similar roles pay here $278k

This role pays less than 56% of similar roles. Most pay $157,500–$225,300 — the shaded band above. At the midpoint, this role pays about $185k versus about $191k for comparable roles.

Based on 240 similar postings.

Employer

About GEICO

GEICO (Government Employees Insurance Company) is one of the largest auto insurers in the United States, offering affordable auto, home, renters, and other personal insurance products. Industry: Insurance

GEICO currently has 81 open roles on FindRole.

Listed pay typically runs $110,000–$230,000 across 79 roles with salary data.

Most-posted roles

View all roles at GEICO

At a glance

TL;DR · Senior Staff Engineer, SRE, Incident Prevention / Post Incident Correction of Errors

Senior Staff Engineer - SRE - Incident Prevention / Post Incident Correction of Errors joins the site reliability team to enhance the reliability and availability of distributed platforms. This role focuses on improving Correction of Error (COE) tooling, conducting root cause analysis, and building automation to eliminate recurring incident patterns. The engineer will design data pipelines, dashboards, and self-service tools while coaching teams to improve post-incident review quality. Key technologies include Go, Java, Python, C#, Kubernetes, and KNative across Azure and AWS environments. The role requires expertise in SQL, NoSQL, Spark, Trino, Grafana, Datadog, Splunk, and PagerDuty. Candidates must possess strong skills in OpenTelemetry, AI-assisted development tools like GitHub Copilot, and advanced incident forensics to solve complex reliability challenges within a high-scale production environment while ensuring continuous improvement through automated workflows and technical leadership.

What does a Engineer earn in California?

Median $209800 from 87 postings across 24 companies.

See salary data

What you'll do

  • Develop automation, self-service tools, dashboards, and data pipelines to scale Correction of Error (COE) workflows.
  • Lead and moderate weekly company-wide presentation sessions for high-severity incidents to ensure proper reporting and preparation.
  • Provide technical leadership in system design and architecture to improve incident tooling and reliability processes.
  • Coach engineering teams to identify true root causes and produce high-quality, actionable post-incident reports.
  • Perform deep root cause analysis on distributed systems using logs, metrics, traces, and observability data.
  • Identify infrastructure gaps such as missing alerts, weak monitoring, or inadequate runbooks following incidents.
  • Analyze incident patterns across the organization to identify and eliminate recurring reliability issues.
  • Manage high-severity production incidents and participate in a 24/7 on-call rotation for mission-critical platforms.

What we're looking for

  • Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience.
  • 10+ years of professional software engineering experience, preferably in platform, reliability, backend, or distributed systems.
  • 8+ years of experience with architecture, design, system reliability, scalability, and technical leadership for production systems.
  • 6+ years of experience with open-source frameworks, modern engineering practices, or platform technologies.
  • 4+ years of experience with Azure, AWS, GCP, or other cloud service providers in complex hybrid environments.
  • Hands-on proficiency in Go, Java, Python, and C# for building full-stack applications on Kubernetes and serverless technologies.
  • Experience with SQL/NoSQL, data pipelines (Spark, Trino), observability tools (Grafana, Datadog, Splunk), and incident management platforms like PagerDuty.
  • Proficiency in AI-assisted development tools such as Claude Code, Cursor, and GitHub Copilot (preferred).

More like this

Similar roles

Staff Engineer - SRE, Retail and Pharmacy

CVS Health

Remote (Woonsocket) 52 days ago $118,450–$260,590
SRE Kubernetes Terraform Prometheus Grafana OpenTelemetry Kafka CI/CD OPA Conftest LitmusChaos Istio Envoy Pulumi Ansible eBPF LLM
8+ yrs exp Remote

Senior Engineer

GEICO

Bethesda, MD +1 3 days ago $100,000–$215,000
SRE Kubernetes AWS Azure Python Go Java C# CI/CD Infrastructure as Code SQL NoSQL Spark Trino Grafana Datadog Splunk PagerDuty OpenTelemetry KNative
4+ yrs exp

Staff Engineer, Software Engineering

GEICO

Bethesda, MD +4 149 days ago $100,000–$230,000
Python C# Go Java C++ SQL NoSQL Docker Kubernetes Azure AWS GCP PowerShell REST APIs Microservices Infrastructure as Code SRE Active Directory SAML OAuth
10+ yrs exp

Senior Software Engineer - SRE, Retail and Pharmacy

CVS Health

Remote (Woonsocket) 53 days ago $92,700–$203,940
SRE DevOps Python Go Java Kubernetes AWS Azure GCP Terraform Ansible Prometheus Grafana OpenTelemetry Apache Kafka ClickHouse LitmusChaos Gremlin Istio Envoy
5+ yrs exp Remote

Senior Engineer, Incident Response Engineering

Target

Brooklyn Park, MN 158 days ago
SOAR Python JavaScript TypeScript React REST APIs Data Pipelines Software Engineering Automation Workflows Backend Development Incident Response Security Automation
5+ yrs exp Hybrid

Senior Staff Engineer, SoC Safety Architect

Qualcomm

San Diego, CA 15 days ago $186,700–$280,100
SoC Architecture ARM x86 RISC V MIPS ISO 26262 IEC 61508 AUTOSAR QNX RTOS RTL Design FMEA FTA DFA FMEDA JAMA Codebeamer DOORS PCIe DDR NoC SMMU
6+ yrs exp