Lead Site Reliability Engineer

JPMorgan Chase

Confirmed live yesterday Low trust

Quick summary

Work type
On-site
Location
Plano, TX
Posted
44 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $181k
$133k most similar roles pay here $235k

This listing doesn't post a salary. Most similar roles pay $147,480–$213,692.

Based on 238 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1117 open roles on FindRole.

Listed pay typically runs $186,160–$215,000 across 7 roles with salary data.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Lead Site Reliability Engineer

Lead Site Reliability Engineer serves as a technical leader within the Corporate Technology team, where you will oversee resiliency design reviews, mentor other engineers, and break down complex problems into manageable tasks for the team. You will champion site reliability culture by using data-driven analytics to improve service levels, identifying bottlenecks, and establishing error budgets with stakeholders. The role involves managing major incidents, performing post-incident analysis, and integrating enterprise-authorized AI capabilities into SRE workflows like CI/CD quality checks and automated testing. Required skills include proficiency in Python, Java/Spring Boot, or .Net, along with experience in container orchestration, networking, and observability tools such as Grafana, Prometheus, and Splunk. You will manage distributed systems using infrastructure-as-code tools like Terraform to ensure reliability, security, and scalability across the organization's technical platforms.

What you'll do

  • Conduct resiliency design reviews and break down complex problems into manageable tasks for other engineers.
  • Use data-driven analytics to identify and resolve technology bottlenecks to improve application stability.
  • Establish service level indicators, objectives, and error budgets in collaboration with stakeholder partners.
  • Act as the primary point of contact during major incidents to quickly resolve issues and prevent financial loss.
  • Provide technical expertise, mentorship, and guidance to other engineers across multiple technical domains.
  • Integrate AI-assisted reliability workflows into CI/CD pipelines while ensuring security and auditability.
  • Evaluate AI-generated operational recommendations for accuracy and risk to ensure alignment with resiliency standards.

What we're looking for

  • Formal training or certification on site reliability engineering concepts and 5+ years of applied experience.
  • Proficiency in reliability, scalability, performance, security, enterprise system architecture, and toil reduction.
  • Fluency in at least one programming language such as Python, Java/Spring Boot, or .Net.
  • Experience using enterprise-authorized AI capabilities to improve SRE workflows while ensuring data sensitivity and security.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk while defining appropriate guardrails.
  • Proficiency in observability (white/black box monitoring, SLO alerting, telemetry), CI/CD practices, and container orchestration.
  • Experience troubleshooting common networking technologies and issues.
  • Advanced knowledge of software applications and technical processes with a commitment to self-education on new technologies.

More like this

Similar roles

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 43 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 43 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Palo Alto, CA 59 days ago
Site Reliability Engineering Java Go Python Terraform Kubernetes Docker CI/CD GitOps Grafana Prometheus Dynatrace Datadog Splunk Kafka RabbitMQ SQS Neo4j Pinecone Weaviate Chroma LangChain LangGraph AutoGen CrewAI GitHub Copilot Fluentd Logstash Vector RESTful APIs RAG TensorFlow PyTorch scikit-learn Hadoop Spark Flink MongoDB Cassandra DynamoDB InfluxDB TimescaleDB AWS Azure GCP
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 45 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 43 days ago
SRE CI/CD Jenkins GitLab Terraform Docker Kubernetes ECS AI Python Go JavaScript GraphQL Kafka OpenTelemetry Networking
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 45 days ago
Site Reliability Engineering Observability Monitoring Telemetry Service Level Objectives Alerting AI SDLC Automation
5+ yrs exp