Lead Site Reliability Engineer

The Federal Reserve

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
San Francisco, CARichmond, VA
Posted
14 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

How this pay compares to similar roles

Similar $181k
$133k most similar roles pay here $235k

This listing doesn't post a salary. Most similar roles pay $147,480–$213,692.

Based on 238 similar postings.

Employer

About The Federal Reserve

The Federal Reserve is the central bank of the United States—one of the world's most influential, trusted and prestigious financial organizations.

The Federal Reserve currently has 23 open roles on FindRole.

Listed pay typically runs $133,350–$174,450 across 16 roles with salary data.

Most-posted roles

View all roles at The Federal Reserve

At a glance

TL;DR · Lead Site Reliability Engineer

Lead Site Reliability Engineer The Lead Site Reliability Engineer joins the engineering team to drive the reliability, scalability, and performance of critical systems. This role involves designing and maintaining highly available cloud infrastructure, establishing SLIs, SLOs, and SLAs, and leading incident response and root cause analysis. The candidate will build internal tools, manage automated deployment pipelines, and implement security controls like SAST and DAST into CI/CD workflows. Technical requirements include proficiency in Java, Python, and Node.js, along with experience in microservices architecture and distributed systems. The role requires expertise in AWS services, Terraform for infrastructure as code, GitLab for CI/CD, and containerization via Docker and Kubernetes. Additionally, the candidate will utilize monitoring tools like Grafana or Datadog to ensure operational excellence across complex financial and payment systems while mentoring junior team members on SRE practices.

What you'll do

  • Design, implement, and maintain highly available and scalable systems across cloud infrastructure.
  • Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance.
  • Lead incident response efforts, conduct root cause analysis, and implement preventive measures.
  • Architect and manage AWS infrastructure using Terraform for Infrastructure as Code.
  • Automate deployment pipelines, monitoring tools, and operational workflows.
  • Develop internal tools and observability frameworks to improve operational efficiency.
  • Integrate security practices into CI/CD pipelines and manage vulnerability assessments.
  • Mentor junior SRE team members and provide technical guidance on system design.

What we're looking for

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 7+ years of experience in Site Reliability Engineering, DevOps, or related roles.
  • 3+ years of experience in a lead or senior technical position.
  • Proficiency in Java, Python, and Node.js with experience in microservices and distributed systems.
  • Extensive experience with AWS services, including compute, storage, database, and networking components.
  • Expert-level knowledge of GitLab CI/CD pipelines, Terraform for infrastructure as code, and container orchestration like Kubernetes or ECS.
  • Experience with security tools (SAST/DAST), identity access management, and monitoring platforms like Grafana or Datadog.
  • Knowledge of LLMs, agentic applications, serverless architectures, and chaos engineering principles is a plus.

More like this

Similar roles

Lead Site Reliability Engineer

Mastercard

O Fallon, MO 16 days ago $122,000$207,000
Site Reliability Engineering Java Spring Framework Python Go DevOps CI/CD Configuration Management Distributed Systems Automation Observability ITSM Root Cause Analysis Capacity Planning Monitoring

Site Reliability Engineer, Lead

Booz Allen Hamilton

Chantilly, VA 50 days ago $99,000$225,000
Prometheus Grafana ELK Stack Linux AWS Python Terraform Terragrunt Kubernetes OpenTelemetry AWS CloudWatch AWS EKS Rancher Jenkins Git Docker Nessus JIRA Confluence SRE
8+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Site Reliability Engineer

Booz Allen Hamilton

McLean, VA 9 days ago $86,800$198,000
AWS Kubernetes Terraform Ansible CI/CD Python Bash PowerShell Docker Prometheus Grafana Loki Elasticsearch Kibana GitLab GitHub CloudFormation OpenShift Jenkins REST JSON YAML XML Agile
6+ yrs exp

Lead Site Reliability Engineer

Alloy

New York, NY 158 days ago $179,000$250,000
Kubernetes AWS Terraform Python Go Docker Infrastructure as Code Datadog CloudWatch ELK EFK Distributed Systems SRE
10+ yrs exp Hybrid

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 43 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp