Vice President, Senior Manager of Site Reliability Engineering

JPMorgan Chase

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Jersey City, NJ
Employment
Full-time
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $188k
$140k $240k
below market most similar roles pay here above market

This listing doesn't post a salary. Most similar roles pay $156,562–$219,343.

Based on 240 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1197 open roles on FindRole.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Vice President, Senior Manager of Site Reliability Engineering

The Vice President - Senior Manager of Site Reliability Engineering joins the Chief Data and Analytics Office AI/ML and Data Platforms team to lead a group of 8–10 engineers. This role serves as the non-functional requirement owner, defining availability targets and embedding reliability principles into product designs for large-scale data platforms and AI/ML workloads. Key responsibilities include overseeing infrastructure stability, managing stakeholders, and driving a culture of blameless post-mortems and continuous improvement. The position requires expertise in site reliability engineering, platform engineering, and distributed systems. Candidates must utilize tools such as Grafana, Dynatrace, Prometheus, Datadog, and Splunk, alongside Databricks, Spark, Docker, Kubernetes, and Terraform. Proficiency in Python is required to manage CI/CD pipelines, automation frameworks, and secure, scalable analytics infrastructure.

What you'll do

  • Lead, mentor, and develop a team of 8–10 site reliability and platform engineers through tailored growth plans.
  • Own non-functional requirements, availability targets, and service level objectives for large-scale data and AI/ML workloads.
  • Drive the design and implementation of observability frameworks using tools like Grafana, Dynatrace, and Prometheus.
  • Oversee the architecture and operational stability of Databricks, Spark-based pipelines, and big data infrastructure.
  • Establish and govern CI/CD pipelines, infrastructure as code practices using Terraform, and automation frameworks.
  • Lead the adoption of enterprise-authorized AI capabilities within site reliability workflows with human-in-the-loop validation.
  • Manage stakeholders to ensure project delivery aligns with compliance standards, risk requirements, and business objectives.
  • Conduct data-driven post-mortems and real-time feedback loops to foster a culture of continuous improvement.

What we're looking for

  • Formal training or certification in site reliability engineering concepts and 5+ years of applied experience.
  • 2+ years of experience leading technologists to manage and solve complex technical items within your domain of expertise.
  • Demonstrated experience managing and growing site reliability or platform engineering teams with direct responsibility for 8 or more engineers.
  • Advanced proficiency in site reliability culture, including implementing SLI/SLO/SLA frameworks, error budgets, and incident management.
  • Hands-on experience with observability tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk.
  • Experience leading platform engineering for large-scale data platforms using Spark and Databricks.
  • Proficiency in Python or similar programming languages for automation and platform development.
  • Experience with Docker, Kubernetes, Terraform, and CI/CD practices.
  • Experience with AWS platforms and cloud-native infrastructure for AI/ML workloads (preferred).
  • Background in AI/ML platform engineering, including infrastructure for model training and monitoring (preferred).

More like this

Similar roles

Lead Site Reliability Engineer

JPMorgan Chase

OH 22 days ago
Site Reliability Engineering Python Go Java C++ Rust Kubernetes Terraform CI/CD Prometheus Grafana Splunk Datadog Dynatrace AI Networking SLO SLI
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Houston, TX 3 days ago
Site Reliability Engineering Python Prometheus OpenTelemetry Datadog Dynatrace Splunk CI/CD SLOs SLIs Observability Incident Management Resilience Engineering Capacity Management Networking Athena AI Automation
5+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Jersey City, NJ +1 106 days ago
Python PySpark AWS Kubernetes Databricks Snowflake Java Spring Boot .Net Grafana Prometheus Datadog Splunk Dynatrace AI/ML Site Reliability Engineering SLI/SLO/SLA System Design
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Jersey City, NJ 41 days ago
SRE Python Bash Terraform Ansible Kubernetes Docker CI/CD Jenkins GitLab CI AWS Azure Prometheus Grafana Datadog Splunk CloudWatch Dynatrace infrastructure-as-code Service Mesh
3+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 73 days ago
Site Reliability Engineering SRE Python Shell PowerShell Ansible Terraform Kubernetes OpenShift AWS Azure GCP Microservices CI/CD Dynatrace Splunk Geneos Grafana ITIL Observability
5+ yrs exp