Lead Site Reliability Engineer, Public Cloud Operational Excellence

JPMorgan Chase

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Jersey City, NJ
Employment
Full-time
Posted
22 days ago
Freshness
Confirmed live today

Market check

Salary context

How this pay compares to similar roles

Similar $186k
$141k most similar roles pay here $240k

This listing doesn't post a salary. Most similar roles pay $151,771–$220,100.

Based on 240 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1164 open roles on FindRole.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Lead Site Reliability Engineer, Public Cloud Operational Excellence

As a Lead Site Reliability Engineer - Public Cloud Operational Excellence within the Public Cloud team, you will combine hands-on engineering with program leadership to enhance platform stability and ensure consistent execution across multiple SRE teams. You will be responsible for establishing gold standards for SLOs, on-call readiness, runbooks, and postmortems while partnering with Engineering and Product teams to integrate reliability outcomes into roadmaps. Your daily work involves managing risk governance, designing operating cadences, and identifying automation opportunities to reduce operational toil. You will utilize AWS, Azure, Python, Go, Bash, and Terraform to build automated solutions. Additionally, you will apply AI/LLM-enabled approaches for triage and incident summarization. The role focuses on solving complex reliability challenges across multi-cloud environments by improving availability metrics, reducing mean time to recovery, and managing the cost of failure.

What you'll do

  • Establish and scale a "gold standard" for SLOs, on-call readiness, runbooks, and postmortems across multi-cloud environments.
  • Integrate reliability outcomes into product roadmaps by converting incident data and support signals into prioritized backlogs.
  • Design and manage operating cadences for incident reviews, KPI tracking, and standardized intake processes.
  • Automate recurring remediations to address configuration drift and risk findings without sacrificing development velocity.
  • Identify and execute automation opportunities to reduce operational toil and deflect high volumes of support tickets.
  • Own cross-platform reporting for metrics including SLO attainment, MTTR, change failure rates, and cost of failure.
  • Implement AI and LLM technologies to improve incident triage, summarization, and automated symptom mapping.
  • Lead systemic remediation efforts and oversee major incident response improvements across the platform.

What we're looking for

  • 5+ years of experience in SRE, production engineering, platform reliability, or infrastructure operations at enterprise scale.
  • Demonstrated success driving cross-team standardization and measurable reliability outcomes through influence and operating mechanisms.
  • Deep knowledge of SLOs/SLIs, error budgets, observability, incident response, postmortems, and change reliability.
  • Strong experience partnering with Engineering and Product leadership to align priorities and deliver results.
  • Hands-on experience with AWS and/or Azure and automation/IaC fundamentals including Python, Go, Bash, Terraform, and CI/CD.
  • Experience using AI/LLM-enabled approaches in operations such as AIOps or agentic workflows with appropriate controls.
  • Proven ability to improve SLO attainment and reduce Sev1/Sev2 frequency (preferred).
  • Ability to reduce MTTR/MTTI, change failure rates, and ticket volume through scaled automation (preferred).

More like this

Similar roles

Lead Site Reliability Engineer

JPMorgan Chase

OH 19 days ago
Site Reliability Engineering Python Go Java C++ Rust Kubernetes Terraform CI/CD Prometheus Grafana Splunk Datadog Dynatrace Networking AI Prompt Engineering Agent Orchestration Infrastructure Automation
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 70 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Houston, TX today
Site Reliability Engineering Python Java Spring Boot .Net CI/CD Container Orchestration Observability Telemetry System Architecture Networking SDLC Automation
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 72 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 70 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp