Principal Site Reliability Engineer

Early Warning Services

Confirmed live today High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Francisco, CAScottsdale, AZChicago, ILNew York, NY
Salary
$173,000–$230,000 / yr
Posted
4 days ago
Freshness
Confirmed live today

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $181k
This role $202k
$132k most similar roles pay here $241k

This role pays more than 62% of similar roles. Most pay $147,075–$215,000 — the shaded band above. At the midpoint, this role pays about $202k versus about $181k for comparable roles.

Based on 238 similar postings.

Employer

About Early Warning Services

Early Warning Services is a fintech company that operates the Zelle person-to-person payments network, the Paze digital checkout wallet, and Certos fraud prevention and identity risk solutions for financial institutions.

Early Warning Services currently has 55 open roles on FindRole.

Listed pay typically runs $132,000–$165,000 across 53 roles with salary data.

Most-posted roles

View all roles at Early Warning Services

At a glance

TL;DR · Principal Site Reliability Engineer

Principal Site Reliability Engineer The Principal Site Reliability Engineer joins a team focused on enhancing the reliability, resilience, scalability, and operational health of production services within the financial payment systems domain. This role involves applying software engineering and systems engineering practices to ensure observability, recoverability, and performance are integrated throughout the service lifecycle. The individual will build and improve CI/CD pipelines, Infrastructure as Code, automation, and monitoring tools while managing incident response and capacity planning. Key technical requirements include proficiency in modern programming languages for scripting, experience with distributed systems, and expertise in public cloud technologies like AWS, Azure, or GCP. Core competencies include Linux/Unix environments, networking, and the development of reusable patterns to reduce operational toil. The role solves critical reliability risks by translating production experiences into improved architecture, tooling, and engineering practices across the organization.

What does a Site Reliability Engineer earn in California?

Median $214000 from 58 postings across 16 companies.

See salary data

What you'll do

  • Apply software engineering and automation principles to improve service reliability, scalability, and operational health.
  • Use data and rigorous analysis to identify reliability risks and guide technical decision-making.
  • Define and implement SLIs, SLOs, error budgets, and other metrics to monitor service health.
  • Enhance observability through improved logging, tracing, monitoring, alerting, and dashboarding.
  • Drive continuous improvement across CI/CD pipelines, Infrastructure as Code, and deployment practices.
  • Translate recurring production issues into improvements in code, architecture, and engineering tools.
  • Lead incident response efforts and conduct blameless post-incident learning to improve system resilience.
  • Reduce operational toil by creating reusable automation patterns and improving engineering practices.

What we're looking for

  • Candidates must independently possess the eligibility to work in the United States at the date of hire.
  • Minimum 15 years of relevant professional experience in SRE, Software Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, or Architecture.
  • Experience with software development or scripting using one or more modern programming languages.
  • Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability.
  • Experience with public cloud technologies and architectures, including infrastructure, networking, Linux/Unix, and modern application architectures.
  • Hands-on experience with AWS is preferred, or comparable experience with other major cloud platforms like Azure, GCP, or OCI (preferred).
  • Experience developing, deploying, operating, or improving highly available production software or distributed systems (preferred).
  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience.

More like this

Similar roles

Principal Site Reliability Engineer

Early Warning Services

Scottsdale, AZ +3 124 days ago $194,000$237,000
Python Go Java Docker Kubernetes Microservices Kafka SQS JMS Oracle Dynamo DB Aurora Redis Memcached Linux CI/CD Git Chef Maven Jenkins AWS Ruby JavaScript
10+ yrs exp Hybrid

Principal Site Reliability Engineer

Nvidia

Santa Clara, CA 13 days ago $248,000$396,750
Kubernetes Distributed Systems Python Go Terraform AWS Azure GCP OpenTelemetry infrastructure-as-code Linux TypeScript JavaScript Java Crossplane AWS CDK CloudFormation AI/ML Platforms High-Performance Computing
10+ yrs exp Hybrid

Lead Site Reliability Engineer

The Federal Reserve

San Francisco, CA +1 24 days ago
AWS Terraform Kubernetes Docker GitLab Python Java Node.js CI/CD S3 DynamoDB RDS Lambda ECS Fargate CloudWatch X-Ray Grafana Datadog Splunk SAST DAST Microservices Agile
7+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

OH 2 days ago
Site Reliability Engineering Python Go Java C++ Rust Kubernetes Terraform CI/CD Prometheus Grafana Splunk Datadog Dynatrace Networking AI Prompt Engineering Agent Orchestration Infrastructure Automation
5+ yrs exp

Principal Site Reliability Engineer

Fidelity Financial Services

Westlake, TX 40 days ago
Site Reliability Engineering Kubernetes Terraform Python AWS Azure Datadog Splunk Grafana PowerBI Jenkins Azure DevOps Cloud Formation Lambda API Gateway infrastructure-as-code Linux Windows
5+ yrs exp