Lead Site Reliability Engineer

JPMorgan Chase

Confirmed live today High trust

Quick summary

Work type
On-site
Location
OH
Posted
2 days ago
Freshness
Confirmed live today

Market check

Salary context

How this pay compares to similar roles

Similar $181k
$132k most similar roles pay here $235k

This listing doesn't post a salary. Most similar roles pay $146,633–$214,875.

Based on 238 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1151 open roles on FindRole.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Lead Site Reliability Engineer

As a Lead Site Reliability Engineer within the Enterprise technology Infrastructure Platforms team, you will hold a leadership role providing technical guidance, mentoring, and oversight for complex products. You will build reliability into the platform by developing production-grade software, including automation, control loops, and self-healing tools to eliminate manual operations. Your daily work involves managing end-to-end service ownership regarding performance, security, and cost while leading incident triage and blameless post-mortems. You will utilize declarative, intent-based systems and manage telemetry data to drive automated remediation. Key technical requirements include proficiency in Python, Go, Java, C++, or Rust, along with experience in Kubernetes, Terraform, and CI/CD. You will also integrate enterprise-authorized AI capabilities into reliability workflows while ensuring security compliance. The role focuses on solving large-scale infrastructure challenges through systems thinking and robust observability using tools like Grafana, Prometheus, Splunk, Datadog, or Dynatrace.

What you'll do

  • Build production-grade software including automation, control loops, and self-healing tools to eliminate manual operations.
  • Implement declarative, intent-based systems that reconcile actual state to a trusted source of truth via automation.
  • Instrument telemetry and state data at scale to drive automated detection, diagnosis, and remediation.
  • Define and operationalize SLIs and SLOs with stakeholders to create actionable, impact-based alerting.
  • Own services end-to-end by managing reliability, performance, security, cost, and leading major incident response.
  • Integrate enterprise-authorized AI capabilities into the SDLC to accelerate triage, troubleshooting, and post-incident analysis.
  • Lead the adoption of AI-assisted reliability workflows across CI/CD pipelines and operational readiness practices.
  • Conduct resiliency design reviews and mentor other engineers on technical and business challenges.

What we're looking for

  • Formal training or certification on site reliability engineering concepts and 5+ years of applied experience.
  • Demonstrated experience using enterprise-authorized AI capabilities to improve SRE workflows with strong validation habits and data sensitivity awareness.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk while defining guardrails for team usage.
  • Strong production software engineering experience in an industry-standard language such as Python, Go, Java, C++, or Rust.
  • Proven ownership of production systems at scale, including on-call responsibility, incident response, and designing for reliability.
  • Demonstrated SLI/SLO/error-budget practice and deep observability skills using platforms like Grafana, Prometheus, Splunk, Datadog, or Dynatrace.
  • Strong *nix fundamentals and infrastructure automation experience with tools such as Kubernetes and Terraform.
  • Networking depth or experience in regulated environments (preferred).

More like this

Similar roles

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 54 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 53 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 53 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 55 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 53 days ago
SRE CI/CD Jenkins GitLab Terraform Docker Kubernetes ECS AI Python Go JavaScript GraphQL Kafka OpenTelemetry Networking
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 5 days ago
SRE AI CI/CD Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
10+ yrs exp