Lead Site Reliability Engineer, Operations Excellence for AI Platforms

JPMorgan Chase

Confirmed live yesterday Trusted

Quick summary

Work type
On-site
Location
Jersey City, NJ
Posted
24 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $182k
$139k most similar roles pay here $235k

This listing doesn't post a salary. Most similar roles pay $148,830–$214,625.

Based on 240 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1117 open roles on FindRole.

Listed pay typically runs $186,160–$215,000 across 7 roles with salary data.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Lead Site Reliability Engineer, Operations Excellence for AI Platforms

Lead Site Reliability Engineer - Operations Excellence for AI Platforms joins the AI Machine Learning and Data platform team to provide technical leadership, mentorship, and expert guidance on complex engineering challenges. This role involves leading resiliency design reviews, managing incident response as an Incident Commander, and overseeing root cause analysis and problem management. The engineer will improve operational readiness through better change management, establish service level objectives, and drive reliability using data-driven analytics to resolve bottlenecks. Key responsibilities include integrating enterprise-authorized AI capabilities into SRE workflows, including CI/CD quality checks and automated testing, while ensuring security and auditability. Required skills include over five years of experience in site reliability, proficiency in Python, and expertise in scalability, performance, and distributed systems. The role focuses on enhancing the stability of large-scale platforms within a high-stakes financial environment.

What you'll do

  • Own and improve the incident management process, including triage, escalation, stakeholder communication, and recovery.
  • Serve as Incident Commander for major incidents while leading root cause analysis and problem management.
  • Drive operational readiness by establishing delivery standards and strengthening change and release management practices.
  • Lead initiatives to improve platform reliability using data-driven analytics to identify and resolve technical bottlenecks.
  • Define service level indicators, objectives, and error budgets in collaboration with stakeholder partners.
  • Integrate enterprise-authorized AI capabilities into incident triage, troubleshooting, and post-incident analysis workflows.
  • Implement AI-assisted reliability workflows across the software development lifecycle while ensuring security and auditability.
  • Provide technical leadership, mentorship, and guidance to other engineers on complex technical and business issues.

What we're looking for

  • Formal training or certification on site reliability engineering concepts is required.
  • Candidates must have 5+ years of applied experience in site reliability engineering.
  • Proficiency in programming languages such as Python is required.
  • Demonstrated proficiency in reliability, scalability, performance, security, and enterprise system architecture is required.
  • Experience using enterprise-authorized AI capabilities to improve SRE workflows with strong validation habits is required.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk while ensuring security alignment is required.
  • Experience improving operational processes including incident management, root cause analysis, and problem management is preferred.
  • Proven ability to influence across teams and communicate clearly with stakeholders in high-urgency situations is preferred.

More like this

Similar roles

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 45 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 43 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 43 days ago
SRE CI/CD Jenkins GitLab Terraform Docker Kubernetes ECS AI Python Go JavaScript GraphQL Kafka OpenTelemetry Networking
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 43 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp