Senior Lead Site Reliability Engineer

JPMorgan Chase

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Jersey City, NJ
Posted
5 days ago
Freshness
Confirmed live today

Market check

Salary context

How this pay compares to similar roles

Similar $181k
$134k most similar roles pay here $234k

This listing doesn't post a salary. Most similar roles pay $146,633–$215,450.

Based on 238 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1151 open roles on FindRole.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Senior Lead Site Reliability Engineer

As a Senior Lead Site Reliability Engineer within the Audit Technology Production Management team, you will hold a leadership role overseeing stability, availability, and resiliency for business-critical Sales platforms. You will conduct resiliency design reviews, mentor other engineers, and break down complex problems into manageable tasks while acting as a technical lead for large products. Your daily responsibilities include managing major incidents, performing root-cause analysis, and driving service maturity through automation and standardized governance. You will integrate enterprise-authorized AI capabilities to accelerate incident triage and automate workflows within the software development lifecycle. Required skills include expertise in Dynatrace, Splunk, Geneos, Grafana, and ITIL frameworks. Technical proficiency includes experience with AWS, Azure, GCP, Python, Shell, PowerShell, Ansible, Terraform, containers, microservices, and Kubernetes or OpenShift to ensure high-quality operational performance across distributed systems.

What you'll do

  • Lead the Production Management team by setting direction, priorities, and performance expectations.
  • Own the stability, availability, and end-to-end operational performance of business-critical Sales platforms.
  • Act as a senior escalation point to lead triage, coordination, and recovery during critical incidents.
  • Drive service maturity through standardization, governance participation, and disciplined support model integration.
  • Improve reliability by implementing systems thinking, root-cause analysis, and advanced observability tools.
  • Manage incident, problem, and change processes to ensure high standards for testing and disaster recovery.
  • Integrate enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis.
  • Lead the adoption of AI-assisted reliability workflows across the software development life cycle and toolchain.

What we're looking for

  • Formal training or certification on site reliability engineering concepts and 10+ years of applied experience.
  • Experience using enterprise-authorized AI capabilities to improve SRE workflows while ensuring data security and validation.
  • Ability to evaluate AI-assisted operational recommendations for correctness, risk, and alignment with resiliency expectations.
  • Leadership experience in Production Support, Production Management, SRE, and Technology Operations teams.
  • Strong systems thinking and problem-solving skills to resolve complex, cross-domain production issues.
  • Demonstrated major incident leadership including triage coordination and root-cause remediation.
  • Expertise in observability tools (Dynatrace, Splunk, Geneos, Grafana) and ITIL service management processes.
  • Experience with cloud platforms, automation scripting, and modern distributed systems (preferred).

More like this

Similar roles

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 53 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 55 days ago
Site Reliability Engineering Observability Monitoring Telemetry Service Level Objectives Alerting AI SDLC Automation
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

OH 2 days ago
Site Reliability Engineering Python Go Java C++ Rust Kubernetes Terraform CI/CD Prometheus Grafana Splunk Datadog Dynatrace Networking AI Prompt Engineering Agent Orchestration Infrastructure Automation
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 53 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 54 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 9 days ago
SRE DevOps Kubernetes AWS Terraform Python Bash Go CI/CD Spinnaker Harness EKS ECS AWS Lambda DynamoDB S3 Linux Distributed Systems Infrastructure as Code Observability
5+ yrs exp