Senior Lead Site Reliability Engineer

JPMorgan Chase

Confirmed live yesterday Low trust

Quick summary

Work type
On-site
Location
Palo Alto, CA
Posted
59 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $182k
$134k most similar roles pay here $235k

This listing doesn't post a salary. Most similar roles pay $150,000–$214,000.

Based on 238 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1117 open roles on FindRole.

Listed pay typically runs $186,160–$215,000 across 7 roles with salary data.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Senior Lead Site Reliability Engineer

Senior Lead Site Reliability Engineer joins the Infrastructure Platforms and Foundational Services team to define non-functional requirements, availability targets, and service level objectives for various product lines. This role involves creating high-quality designs, roadmaps, and program charters while mentoring technologists and championing site reliability adoption. The engineer will build observability and reliability designs for complex systems, manage infrastructure as code, and implement AI-assisted reliability workflows across the software development lifecycle. Key technical requirements include proficiency in Java, Go, Python, and Terraform, alongside experience with Docker, Kubernetes, and CI/CD pipelines. The role requires expertise in monitoring tools like Grafana and Prometheus, message queues like Kafka, and various databases including Neo4j and Pinecone. Additionally, the candidate will build AI agents using frameworks like LangChain and LangGraph to enhance reliability engineering within a distributed systems architecture.

What you'll do

  • Define non-functional requirements and availability targets for products and service lines.
  • Design and implement observability and reliability systems to ensure stable performance without increasing technical debt.
  • Integrate enterprise-authorized AI capabilities into reliability designs, incident analysis, and operational decision-making.
  • Lead the adoption of AI-assisted workflows across the software development lifecycle to automate testing and production readiness.
  • Develop high-quality engineering designs, roadmaps, and program charters for infrastructure platforms.
  • Debug and evolve critical components by analyzing application and platform interdependencies.
  • Mentor technologists on technical issues while promoting site reliability culture and practices.
  • Contribute to the firm's site reliability community through internal forums, guilds, and conferences.

What we're looking for

  • Formal training or certification in site reliability engineering concepts and 5+ years of applied experience.
  • Advanced knowledge of site reliability culture and principles to implement them within applications or platforms.
  • Expert proficiency in Java, Go (Golang), Python, and Terraform for building high-performance systems and infrastructure as code.
  • Experience using enterprise-authorized AI capabilities to improve reliability workflows while ensuring data security and auditability.
  • Hands-on experience building AI Agents and autonomous systems using frameworks like LangChain, LangGraph, AutoGen, or CrewAI.
  • Advanced knowledge of observability tools (Grafana, Dynatrace, Prometheus, Datadog, Splunk) and logging pipelines (Fluentd, Logstash, Vector).
  • Experience building production-grade RESTful APIs, message queue architectures (Kafka, RabbitMQ, SQS), and various database types including vector and graph databases.
  • Proficiency in containerization (Docker, Kubernetes), CI/CD pipelines, and GitOps workflows.

More like this

Similar roles

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 45 days ago
Site Reliability Engineering Observability Monitoring Telemetry Service Level Objectives Alerting AI SDLC Automation
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 43 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 45 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 43 days ago
SRE CI/CD Jenkins GitLab Terraform Docker Kubernetes ECS AI Python Go JavaScript GraphQL Kafka OpenTelemetry Networking
5+ yrs exp