Principal Site Reliability Engineer

Early Warning Services

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Scottsdale, AZSan Francisco, CAChicago, ILNew York, NY
Salary
$194,000–$237,000 / yr
Posted
118 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $184k
This role $216k
$132k most similar roles pay here $248k

This role pays more than 76% of similar roles. Most pay $152,150–$214,875 — the shaded band above. At the midpoint, this role pays about $216k versus about $184k for comparable roles.

Based on 238 similar postings.

Employer

About Early Warning Services

Early Warning Services is a fintech company that operates the Zelle person-to-person payments network, the Paze digital checkout wallet, and Certos fraud prevention and identity risk solutions for financial institutions.

Early Warning Services currently has 64 open roles on FindRole.

Listed pay typically runs $143,500–$183,000 across 60 roles with salary data.

Most-posted roles

View all roles at Early Warning Services

At a glance

TL;DR · Principal Site Reliability Engineer

The Principal Site Reliability Engineer joins the engineering team to design availability and resiliency patterns for infrastructure and applications. This role involves building automation tools for deployment, configuration changes, and disaster recovery while designing observability systems to proactively detect issues. The engineer will identify performance bottlenecks, manage capacity planning, and serve as a technical liaison by providing runbooks to support teams. Key responsibilities include implementing microservice design patterns and mentoring team members on site reliability trends. The role requires proficiency in Python, Go, Java, or Ruby; experience with Docker, Kubernetes, and Swarm; and familiarity with messaging frameworks like Kafka, SQS, or JMS. Additionally, the candidate must possess skills in Linux administration, networking fundamentals, and CI/CD pipelines using Git, Chef, Maven, and Jenkins while managing database technologies such as Oracle, DynamoDB, and Aurora.

What does a Site Reliability Engineer earn in California?

Median $214000 from 60 postings across 17 companies.

See salary data

What you'll do

  • Design and implement software tools to improve application performance, availability, scalability, and latency.
  • Build automation and tooling for deployment, configuration changes, and disaster recovery scenarios.
  • Develop and manage observability and monitoring systems to proactively detect problems and identify root causes.
  • Analyze application capacity continuously to provide scaling recommendations to product and business teams.
  • Identify performance bottlenecks and collaborate with cross-functional teams to troubleshoot and resolve issues.
  • Create technical documentation and runbooks for Level 1 and Level 2 support teams.
  • Provide feedback on software development lifecycles by implementing microservice design patterns and modern reliability techniques.
  • Mentor team members and serve as a subject matter expert on site reliability engineering trends.

What we're looking for

  • Candidates must possess eligibility to work in the United States at the date of hire.
  • A Bachelor's degree in Business, Computer Science, or a related field is typically required.
  • 12+ years of experience managing complex projects in a technical or software development environment is required.
  • Proven ability to lead teams through high-priority incidents and improve the RCA process is required.
  • Experience designing and developing with Python, Go, Java, Docker, Microservices Architecture, and messaging frameworks like Kafka, SQS, or JMS is required.
  • Strong understanding of Linux administration, networking fundamentals, and CI/CD pipelines including GIT, Chef, Maven, and Jenkins is required.
  • Proven track record of designing and building complex end-to-end systems as a full stack developer is required.
  • Experience in 24/7 production environments and knowledge of AWS, Docker, Kubernetes, or Swarm are preferred.

More like this

Similar roles

Principal Site Reliability Engineer

Oracle

Reston, VA +1 49 days ago $84,900$209,500
Kubernetes Terraform Docker Python Bash Linux Unix Oracle Database RAC Chef Puppet DNS DHCP HTTP TCP/IP LLM VMware Cisco
6+ yrs exp

Lead Site Reliability Engineer

Mastercard

O Fallon, MO 5 days ago $122,000$207,000
Site Reliability Engineering Java Spring Framework Python Go DevOps CI/CD Configuration Management Distributed Systems Automation Observability ITSM Root Cause Analysis Capacity Planning Monitoring

Senior Site Reliability Engineer

Fiserv

Berkeley Heights, NJ 10 days ago $128,000$216,000
AWS Kubernetes Terraform CI/CD GitHub Actions Python Bash Ruby on Rails Docker Linux Unix RDBMS Document Storage New Relic Dynatrace Datadog DNS Load Balancing Virtual Networking

Lead Site Reliability Engineer

The Federal Reserve

San Francisco, CA +1 18 days ago
AWS Terraform Kubernetes Docker GitLab Python Java Node.js CI/CD S3 DynamoDB RDS Lambda ECS Fargate CloudWatch X-Ray Grafana Datadog Splunk SAST DAST Microservices Agile
7+ yrs exp

Principal Site Reliability Engineer

Oracle

Nashville, TN 61 days ago $84,900$209,500
Site Reliability Engineering Oracle Cloud Infrastructure Linux Windows Server Python PowerShell Bash Ansible Chef infrastructure-as-code Networking DNS Firewalls Load Balancing Certificates incident-management Observability Capacity Planning
3+ yrs exp

Principal Site Reliability Engineer

Oracle

Nashville, TN 61 days ago $84,900$209,500
Site Reliability Engineering Linux Windows Server Python Bash PowerShell Ansible Chef Oracle Cloud Infrastructure infrastructure-as-code Monitoring Logging Observability Capacity Planning incident-management Patching Citrix
3+ yrs exp