Lead Site Reliability Engineer

JPMorgan Chase

Confirmed live yesterday Trusted

Quick summary

Work type
On-site
Location
New York, NY
Posted
43 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $181k
$133k most similar roles pay here $235k

This listing doesn't post a salary. Most similar roles pay $147,480–$213,692.

Based on 238 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1117 open roles on FindRole.

Listed pay typically runs $186,160–$215,000 across 7 roles with salary data.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Lead Site Reliability Engineer

Lead Site Reliability Engineer As a Lead Site Reliability Engineer within the Production Management team, you will lead efforts to ensure the stability, availability, and resiliency of business-critical Sales Execution platforms. You will manage technical priorities, conduct resiliency design reviews, mentor other engineers, and serve as a senior escalation point during critical incidents. Your daily work involves managing incident, problem, and change management processes while integrating AI-assisted reliability workflows into the software development lifecycle. You will partner with cross-functional teams to improve platform performance across Rates, Credit, FX, SPG, and Repo products. Required skills include expertise in Dynatrace, Splunk, Geneos, Grafana, and ITIL frameworks. Technical proficiency includes AWS, Azure, GCP, Python, Shell, PowerShell, Ansible, Terraform, containers, microservices, and Kubernetes or OpenShift to support high-availability trading environments and complex trade lifecycle processes.

What you'll do

  • Lead the Production Management team by setting direction, priorities, and performance expectations for Sales Execution platforms.
  • Own the stability, availability, and end-to-end operational performance of business-critical trading systems.
  • Act as a senior escalation point to lead triage, coordination, and recovery during critical incidents.
  • Drive reliability improvements through root-cause analysis, observability, monitoring, and automation.
  • Manage incident, problem, and change management processes to ensure system readiness and eliminate recurring issues.
  • Integrate enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis.
  • Implement AI-assisted reliability workflows across the software development lifecycle to improve testing and operational readiness.
  • Partner with cross-functional teams to improve platform stability and service maturity for Front Office stakeholders.

What we're looking for

  • Formal training or certification on site reliability engineering concepts and 5+ years of applied experience.
  • Demonstrated experience using enterprise-authorized AI capabilities to improve SRE workflows with strong validation habits and data sensitivity awareness.
  • Ability to evaluate AI-assisted operational recommendations for correctness, risk, and security alignment.
  • Leadership experience across Production Support, Production Management, SRE, and Technology Operations teams.
  • Proven background supporting Front Office Sales users and Sales Execution platforms in high-availability environments.
  • Strong knowledge of Rates, Credit, FX, Repo, or SPG workflows including RFQs, pricing, and trade lifecycles.
  • Expertise in observability tools (Dynatrace, Splunk, Geneos, Grafana) and ITIL incident, problem, and change management.
  • Experience with cloud platforms, automation/scripting (Python, Shell, Ansible, Terraform), and distributed systems (Kubernetes, containers). (preferred)

More like this

Similar roles

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Palo Alto, CA 59 days ago
Site Reliability Engineering Java Go Python Terraform Kubernetes Docker CI/CD GitOps Grafana Prometheus Dynatrace Datadog Splunk Kafka RabbitMQ SQS Neo4j Pinecone Weaviate Chroma LangChain LangGraph AutoGen CrewAI GitHub Copilot Fluentd Logstash Vector RESTful APIs RAG TensorFlow PyTorch scikit-learn Hadoop Spark Flink MongoDB Cassandra DynamoDB InfluxDB TimescaleDB AWS Azure GCP
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 43 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 45 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 43 days ago
SRE CI/CD Jenkins GitLab Terraform Docker Kubernetes ECS AI Python Go JavaScript GraphQL Kafka OpenTelemetry Networking
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 45 days ago
Site Reliability Engineering Observability Monitoring Telemetry Service Level Objectives Alerting AI SDLC Automation
5+ yrs exp