Lead Site Reliability Engineer, Network

JPMorgan Chase

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
OH
Employment
Full-time
Posted
29 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $181k
$135k most similar roles pay here $240k

This listing doesn't post a salary. Most similar roles pay $145,000–$217,500.

Based on 240 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1164 open roles on FindRole.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Lead Site Reliability Engineer, Network

Lead Site Reliability Engineer - Network serves as a technical leader within the Enterprise technology Infrastructure Platforms team. This role involves leading resiliency design reviews, mentoring engineers, and managing end-to-end problem management including root cause analysis for mission-critical network services. The position focuses on ensuring availability, performance, and recoverability while driving automation to reduce toil. Key responsibilities include utilizing enterprise-authorized AI capabilities for incident triage and workflow automation across the software development lifecycle. The role requires expertise in SD-WAN, routing, switching, and network security technologies like firewalls and load balancers. Candidates must be proficient in Python, Shell, and Ansible to build self-healing systems. Technical requirements include experience with container orchestration, CI/CD pipelines, and observability tools. The work centers on solving complex networking bottlenecks and improving service levels through data-driven analytics and proactive reliability engineering practices.

What you'll do

  • Lead resiliency design reviews and break down complex technical problems into manageable tasks for other engineers.
  • Use data-driven analytics to identify bottlenecks and improve the reliability of applications and platforms.
  • Establish service level indicators, objectives, and error budgets in collaboration with stakeholder partners.
  • Act as the primary point of contact during major incidents to resolve issues quickly and prevent financial loss.
  • Manage end-to-end problem management by authoring RCAs and implementing durable remediations for systemic root causes.
  • Develop automation using Python, Shell, and Ansible to reduce toil and improve system reliability.
  • Provide technical leadership in network services including routing, switching, firewalls, and load balancers.
  • Integrate AI-assisted workflows into the SDLC to accelerate incident triage and automate operational readiness.

What we're looking for

  • Formal training or certification on site reliability engineering concepts and 5+ years of applied experience.
  • Demonstrated proficiency in reliability, scalability, performance, security, enterprise system architecture, and toil reduction.
  • Fluency in at least one programming language such as Python, Java/Spring Boot, or .Net.
  • Experience using enterprise-authorized AI capabilities to improve SRE workflows while ensuring data sensitivity and security compliance.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk while defining appropriate guardrails.
  • Proficiency in observability (white/black box monitoring, SLO alerting, telemetry), CI/CD practices, and container orchestration.
  • Experience troubleshooting common networking technologies and issues.
  • Extensive experience operating large-scale networks and advanced automation using Python, Shell, and Ansible (preferred).

More like this

Similar roles

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 70 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Houston, TX today
Site Reliability Engineering Python Java Spring Boot .Net CI/CD Container Orchestration Observability Telemetry System Architecture Networking SDLC Automation
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 72 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 70 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

OH 19 days ago
Site Reliability Engineering Python Go Java C++ Rust Kubernetes Terraform CI/CD Prometheus Grafana Splunk Datadog Dynatrace Networking AI Prompt Engineering Agent Orchestration Infrastructure Automation
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 22 days ago
SRE AI CI/CD Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
10+ yrs exp