Site Reliability Engineering Lead

US Bank

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Atlanta, GAChicago, ILIrving, TX
Salary
$111,605–$131,300 / yr
Posted
3 days ago
Freshness
Confirmed live yesterday
Closes
Sep 30, 2026

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $183k
This role $121k
$97k most similar roles pay here $244k

This role pays less than 95% of similar roles. Most pay $152,150–$214,000 — the shaded band above. At the midpoint, this role pays about $121k versus about $183k for comparable roles.

Based on 238 similar postings.

Employer

About US Bank

U.S. Bank (U.S. Bancorp) is the fifth-largest bank in the United States, providing retail banking, corporate and commercial banking, wealth management, and payment services to millions of customers. Industry: Banking & Financial Services

US Bank currently has 30 open roles on FindRole.

Listed pay typically runs $119,765–$140,900 across 29 roles with salary data.

Most-posted roles

View all roles at US Bank

At a glance

TL;DR · Site Reliability Engineering Lead

Site Reliability Engineering (SRE ) lead joins the team to oversee production reliability and infrastructure stability. This role involves leading the troubleshooting of complex incidents like application failures and cloud outages while performing root cause analysis and implementing corrective actions. The individual will design monitoring, observability tools, and automated runbooks using Infrastructure as Code and CI/CD pipelines to reduce manual effort. Key responsibilities include serving as an Incident Commander, mentoring engineers, and managing workloads based on metrics like MTTR and SLA compliance. Technical requirements include expertise in AWS, Azure, Kubernetes, Docker, and Python, along with proficiency in PowerShell, Shell Scripting, Terraform, Ansible, and SQL. The role utilizes tools such as Datadog, Splunk, Dynatrace, Grafana, Prometheus, CloudWatch, and Azure Monitor to ensure high availability for distributed systems and cloud-native infrastructure platforms.

What you'll do

  • Lead the troubleshooting and resolution of complex production incidents including application failures and cloud outages.
  • Conduct root cause analysis to identify impacts and implement permanent corrective actions.
  • Design and enhance monitoring, observability tools, dashboards, and operational runbooks for platform reliability.
  • Drive automation initiatives using scripting, Infrastructure as Code, and CI/CD pipelines to reduce manual effort.
  • Serve as the Incident Commander during major incidents to coordinate cross-functional response teams.
  • Provide leadership, coaching, and workload management for SRE, DevOps, and production support engineers.
  • Utilize operational metrics like MTTR and SLA compliance to drive continuous improvement.

What we're looking for

  • Bachelor's degree or equivalent work experience.
  • Six to eight years of relevant work experience in business/risk analysis, IT Service Management, production support, project management, or application development.
  • Expertise in Site Reliability Engineering (SRE), DevOps, Production Support, Platform Engineering, and Distributed Systems Operations (preferred).
  • Experience leading technical teams, incident response efforts, workload prioritization, and reliability improvement programs (preferred).
  • Advanced knowledge of Incident Management, Problem Management, Change Management, and Root Cause Analysis methodologies (preferred).
  • Hands-on experience with AWS, Azure, Kubernetes, Docker, and cloud-native infrastructure platforms (preferred).
  • Proficiency in Python, PowerShell, Shell Scripting, and automation frameworks for operational efficiency (preferred).
  • Experience with CI/CD pipelines, monitoring tools like Datadog or Splunk, and Infrastructure as Code tools like Terraform or Ansible (preferred).

More like this

Similar roles

Site Reliability Engineer, Lead

Booz Allen Hamilton

Chantilly, VA 50 days ago $99,000$225,000
Prometheus Grafana ELK Stack Linux AWS Python Terraform Terragrunt Kubernetes OpenTelemetry AWS CloudWatch AWS EKS Rancher Jenkins Git Docker Nessus JIRA Confluence SRE
8+ yrs exp

Senior Site Reliability Engineer

CVS Health

Woonsocket, RI +1 14 days ago $92,700$203,940
SRE DevOps Kubernetes OpenShift Docker CI/CD Splunk Dynatrace Datadog Prometheus Grafana Python Java AWS Microsoft Azure Google Cloud Rancher GitHub BitBucket Jenkins Microservices web API’s Apigee
5+ yrs exp Hybrid

Senior Manager Site Reliability Engineer

The Walt Disney Company

Remote 16 days ago $175,000$215,000
SRE AWS GCP Azure Kubernetes Terraform Ansible Harness GitLab CloudFormation CI/CD Observability Automation Serverless DevOps
10+ yrs exp Remote

Lead Site Reliability Engineer

Alloy

New York, NY 158 days ago $179,000$250,000
Kubernetes AWS Terraform Python Go Docker Infrastructure as Code Datadog CloudWatch ELK EFK Distributed Systems SRE
10+ yrs exp Hybrid

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 43 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 43 days ago
SRE CI/CD Jenkins GitLab Terraform Docker Kubernetes ECS AI Python Go JavaScript GraphQL Kafka OpenTelemetry Networking
5+ yrs exp