Lead Site Reliability Engineer

JPMorgan Chase

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Jersey City, NJ
Employment
Full-time
Posted
5 days ago
Freshness
Confirmed live today

Market check

Salary context

How this pay compares to similar roles

Similar $181k
$135k $235k
below market most similar roles pay here above market

This listing doesn't post a salary. Most similar roles pay $145,000–$217,500.

Based on 240 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1219 open roles on FindRole.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Lead Site Reliability Engineer

As a Lead Site Reliability Engineer within the Asset and Wealth Management, Tech Production and Infrastructure Delivery team, you will improve reliability, resilience, and operational performance across a hybrid technology environment spanning modern distributed platforms and mainframe systems. You will lead the adoption of SRE practices, including SLIs, SLOs, error budgets, and blameless post-incident processes. Daily responsibilities involve designing monitoring and observability capabilities for metrics, logs, and traces, while driving automation for self-healing and CI/CD operational controls. You will use Python, Go, or Shell to reduce operational toil and manage incident response. The role requires expertise in Linux, networking, middleware, containers, and cloud platforms. You will partner with various teams to standardize reliability patterns and operational controls within a regulated, high-control environment.

What you'll do

  • Lead the adoption of SRE practices including SLIs, SLOs, error budgets, and blameless post-incident processes.
  • Design and improve monitoring and observability capabilities for metrics, logs, traces, and event telemetry.
  • Establish actionable alerting standards, dashboards, and runbooks to improve operational readiness.
  • Drive automation initiatives for self-healing, automated remediation, and CI/CD operational controls.
  • Identify and reduce operational toil through process optimization and platform improvements.
  • Improve incident management practices including response coordination and root cause analysis.
  • Support capacity planning, performance engineering, and resilience testing to strengthen service stability.
  • Standardize reliability patterns and operational controls across distributed and mainframe domains.

What we're looking for

  • Relevant experience in Site Reliability Engineering, Production Engineering, Infrastructure Engineering, or a similar reliability-focused role.
  • Experience leading technical initiatives.
  • Strong knowledge of operating and supporting distributed systems in production, including Linux, networking, middleware, containers, and cloud platforms.
  • Hands-on experience with monitoring and observability platforms, including metrics, logs, traces, dashboarding, and alert engineering.
  • Ability to automate operational workflows using scripting/programming languages like Python, Go, or Shell and infrastructure-as-code.
  • Experience supporting or integrating mainframe systems into enterprise operations.
  • Strong communication and stakeholder management skills to lead cross-team reliability improvements.
  • Experience establishing SLIs/SLOs, resilience patterns, and reliability testing (preferred).

More like this

Similar roles

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 73 days ago
Site Reliability Engineering SRE Python Shell PowerShell Ansible Terraform Kubernetes OpenShift AWS Azure GCP Microservices CI/CD Dynatrace Splunk Geneos Grafana ITIL Observability
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Houston, TX 3 days ago
Site Reliability Engineering Python Prometheus OpenTelemetry Datadog Dynatrace Splunk CI/CD SLOs SLIs Observability Incident Management Resilience Engineering Capacity Management Networking Athena AI Automation
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

OH 22 days ago
Site Reliability Engineering Python Go Java C++ Rust Kubernetes Terraform CI/CD Prometheus Grafana Splunk Datadog Dynatrace AI Networking SLO SLI
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 73 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Container Orchestration Networking FinOps Disaster Recovery Capacity Management ITIL AI Telemetry Infrastructure Optimization
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 75 days ago
Site Reliability Engineering Python Java Spring Boot .Net CI/CD AWS Observability Telemetry Networking System Architecture AI SDLC
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 25 days ago
Site Reliability Engineering SRE Python Shell PowerShell Ansible Terraform Kubernetes OpenShift AWS Azure GCP Microservices CI/CD Dynatrace Splunk Geneos Grafana ITIL Observability AI
10+ yrs exp