Director, Site Reliability Engineering

Early Warning Services

Confirmed live today High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Scottsdale, AZChicago, ILSan Francisco, CA
Salary
$173,000–$230,000 / yr
Employment
Full-time
Posted
3 days ago
Freshness
Confirmed live today

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $187k
This role $202k
$136k $249k
below market most similar roles pay here above market

This role pays more than 53% of similar roles. Most pay $147,200–$226,350 — the blue band above. At the midpoint, this role pays about $202k versus about $187k for comparable roles.

Based on 240 similar postings.

Employer

About Early Warning Services

Early Warning Services is a fintech company that operates the Zelle person-to-person payments network, the Paze digital checkout wallet, and Certos fraud prevention and identity risk solutions for financial institutions.

Early Warning Services currently has 78 open roles on FindRole.

Listed pay typically runs $147,500–$180,500 across 40 roles with salary data.

Most-posted roles

View all roles at Early Warning Services

At a glance

TL;DR · Director, Site Reliability Engineering

The Director, Site Reliability Engineering leads the SRE function for an assigned product, platform, pillar, or business domain. This leader is accountable for the reliability, scalability, performance, and operability of business-critical production services. Key responsibilities include establishing Service Level Indicators and Objectives, managing error budgets, and championing automation to reduce operational toil. The role requires developing engineering talent, managing incident response, and ensuring production readiness for highly available distributed systems. The candidate must possess technical depth to develop code and contribute to engineering solutions while partnering with Product, Infrastructure, and Security teams. Core competencies include SRE practices, observability engineering, and capacity planning. This role solves critical reliability risks within a highly regulated financial services environment, ensuring that payment systems remain resilient, secure, and performant for millions of consumers.

What you'll do

  • Establish and monitor Service Level Indicators (SLIs) and Service Level Objectives (SLOs) aligned with business outcomes.
  • Identify and prioritize reliability risks using production data, error budgets, and failure analysis.
  • Develop and maintain reusable software, tooling, and automation to improve scalability and reduce manual toil.
  • Drive observability by establishing requirements for metrics, logs, traces, and actionable alerting.
  • Lead disciplined incident response, service restoration, and blameless post-incident reviews to drive systemic improvements.
  • Ensure production readiness by addressing systemic failure modes, capacity constraints, and recovery gaps.
  • Build and develop high-performing SRE teams while providing hands-on technical leadership and code contributions.
  • Translate business priorities into a multi-quarter reliability roadmap and measurable engineering outcomes.

What we're looking for

  • Candidates must independently possess the eligibility to work in the United States for any employer.
  • Typically 12+ years of relevant software engineering, site reliability engineering, production engineering, platform engineering, or closely related experience.
  • 5+ years of people leadership experience leading engineering teams and developing managers or senior technical leaders.
  • Experience operating highly available, business-critical distributed systems and leading reliability improvements across a major product or platform.
  • Strong understanding of SRE practices including SLOs/SLIs, error budgets, observability, incident management, automation, capacity, resilience, and production readiness.
  • Demonstrated hands-on technical capability to develop code, contribute to automation, and work directly with engineers.
  • Ability to communicate technical risk, tradeoffs, and investment needs clearly to engineering, product, and senior business stakeholders.
  • Experience in payments, financial services, or other highly regulated environments (preferred).

More like this

Similar roles

Principal Site Reliability Engineer

Early Warning Services

Scottsdale, AZ +3 20 days ago $194,000–$237,000
AWS Azure GCP OCI Linux Unix CI/CD Infrastructure as Code Containers Distributed Systems Observability Monitoring Logging Tracing Incident Management SLIs SLOs Error Budgets Capacity Management Disaster Recovery
10+ yrs exp Hybrid

Senior Site Reliability Engineer

Early Warning Services

Scottsdale, AZ +3 25 days ago $106,000–$130,000
AWS Azure GCP OCI Linux Unix CI/CD Infrastructure as Code Distributed Systems Observability Monitoring Logging Tracing Incident Management SLIs SLOs Error Budgets Capacity Management
5+ yrs exp Hybrid

Staff Site Reliability Engineer

Early Warning Services

San Francisco, CA +3 30 days ago $131,000–$160,000
AWS Azure GCP OCI Linux Unix CI/CD Infrastructure as Code Containers Distributed Systems Observability Monitoring Logging Tracing Incident Management SLIs SLOs Error Budgets
8+ yrs exp Hybrid

Principal Site Reliability Engineer

Early Warning Services

San Francisco, CA +3 24 days ago $173,000–$230,000
AWS Azure GCP OCI Linux Unix CI/CD Infrastructure as Code Containers Distributed Systems Observability Monitoring Logging Tracing Incident Management SLIs SLOs Error Budgets Capacity Management Networking
10+ yrs exp Hybrid

Principal Site Reliability Engineer

Early Warning Services

Scottsdale, AZ 11 days ago $194,000–$237,000
Python Go Java Docker Kubernetes AWS Microservices Kafka SQS JMS Oracle Dynamo DB Aurora Redis Memcached Linux CI/CD Git Chef Maven Jenkins Ruby JavaScript Swarm
10+ yrs exp Hybrid

Principal Site Reliability Engineer

Nvidia

Santa Clara, CA 33 days ago $248,000–$396,750
Kubernetes Distributed Systems Python Go Terraform AWS Azure GCP OpenTelemetry infrastructure-as-code Linux TypeScript JavaScript Java Crossplane AWS CDK AWS CloudFormation Capacity Management Observability
10+ yrs exp Hybrid