Senior Lead Site Reliability Engineer

JPMorgan Chase

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Jersey City, NJ
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $182k
$134k most similar roles pay here $235k

This listing doesn't post a salary. Most similar roles pay $148,740–$214,375.

Based on 238 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1138 open roles on FindRole.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Senior Lead Site Reliability Engineer

As a Senior Lead Site Reliability Engineer within the Commercial Investment Banking team of Fraud Prevention, you will manage production reliability outcomes by monitoring availability, latency, throughput, and error rates. You will build infrastructure as code using Terraform modules, automate operational tasks with Python, Bash, or Go, and manage Kubernetes workloads alongside AWS components like EKS, ECS, and Lambda. Your daily work involves defining SLIs/SLOs, improving observability through metrics and logs, and leading incident response and root cause analyses for distributed systems. You will also enhance deployment safety using Spinnaker and Harness while ensuring security controls are embedded into operations. The role focuses on solving complex business problems within fraud screening and payment flows by optimizing infrastructure, reducing operational toil, and ensuring the performance of mission-critical systems in a production environment.

What you'll do

  • Manage production reliability outcomes including availability, latency, throughput, and error rates.
  • Define and evolve SLIs, SLOs, and error budgets while tuning alerts to reduce noise.
  • Improve end-to-end observability and troubleshooting across Kubernetes and AWS environments.
  • Lead incident response, triage, mitigation, and root cause analysis (RCA) for distributed systems.
  • Operate and tune Kubernetes workloads including autoscaling, resource management, and resilience patterns.
  • Manage AWS container and serverless components such as EKS, ECS, and Lambda.
  • Build infrastructure as code using Terraform to ensure environment consistency and automation.
  • Improve release engineering by increasing deployment safety and repeatability using Spinnaker and Harness.

What we're looking for

  • Formal training or certification on software engineering concepts and 5+ years of applied experience.
  • Experience in SRE, DevOps, or production engineering.
  • Hands-on experience operating Kubernetes workloads including deployments, scaling, and debugging.
  • Practical experience with AWS services such as EKS, ECS, Lambda, DynamoDB, and S3 in production.
  • Proficiency with Terraform (IaC) and scripting/automation using Python, Bash, or Go.
  • Experience with CI/CD and release tooling such as Spinnaker and/or Harness.
  • Strong incident response skills, RCA writing, and ability to troubleshoot distributed systems on Linux and networking.
  • Working knowledge of using enterprise-authorized AI capabilities for SRE workflows while following data sensitivity requirements.
  • Experience implementing SLO programs, performance testing, and capacity planning (preferred).
  • Experience with fraud screening, payment flows, or secure operational practices (preferred).

More like this

Similar roles

Senior Lead Site Reliability Engineer

JPMorgan Chase

Palo Alto, CA 63 days ago
Site Reliability Engineering Java Go Python Terraform Kubernetes Docker CI/CD GitOps Grafana Prometheus Dynatrace Datadog Splunk Kafka RabbitMQ SQS Neo4j Pinecone Weaviate Chroma LangChain LangGraph AutoGen CrewAI GitHub Copilot Fluentd Logstash Vector RESTful APIs RAG TensorFlow PyTorch scikit-learn Hadoop Spark Flink MongoDB Cassandra DynamoDB InfluxDB TimescaleDB AWS Azure GCP
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 49 days ago
Site Reliability Engineering Observability Monitoring Telemetry Service Level Objectives Alerting AI SDLC Automation
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 48 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 47 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 49 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 47 days ago
SRE CI/CD Jenkins GitLab Terraform Docker Kubernetes ECS AI Python Go JavaScript GraphQL Kafka OpenTelemetry Networking
5+ yrs exp