Site Reliability Engineer III, Production Management

JPMorgan Chase

Confirmed live 2 days ago Low trust

Quick summary

Work type
On-site
Location
New York, NY
Posted
59 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

How this pay compares to similar roles

Similar $179k
$133k most similar roles pay here $229k

This listing doesn't post a salary. Most similar roles pay $143,475–$213,692.

Based on 238 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1117 open roles on FindRole.

Listed pay typically runs $186,160–$215,000 across 7 roles with salary data.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Site Reliability Engineer III, Production Management

As a Site Reliability Engineer III- Production Management within the Commercial & Investment Bank, you will join the production management team to solve complex business problems through code and cloud infrastructure. You will be responsible for configuring, maintaining, monitoring, and optimizing applications while improving end-to-end operations, availability, reliability, and scalability. Your daily work involves building automated continuous integration and delivery pipelines, enhancing production observability, managing runbooks, and leading L1/L2 support in time-sensitive environments. You will utilize Python, Java/Spring Boot, or .Net to automate tasks and reduce toil while leveraging enterprise-authorized AI capabilities for incident triage and pattern recognition. The role requires expertise in AWS, Datadog, and distributed systems debugging. You will address critical infrastructure challenges within the domain of pricing, risk, and market data support for financial services.

What does a Site Reliability Engineer earn?

Median $186200 from 131 postings across 36 companies.

See salary data

What you'll do

  • Configure, maintain, monitor, and optimize applications and cloud infrastructure using code.
  • Design and implement deployment and reliability approaches using automated CI/CD pipelines.
  • Use enterprise-authorized AI tools to accelerate incident triage, troubleshooting, and post-incident analysis.
  • Monitor operational signals and coordinate rapid responses to service health issues.
  • Improve observability by enhancing instrumentation, dashboards, and alert quality to reduce MTTR.
  • Manage the support operating model including runbooks, shift coverage, and post-incident follow-through.
  • Partner with engineering teams to drive root-cause fixes and build self-service automation tools.
  • Lead L1/L2 production support by triaging incidents and providing stakeholder updates during resolution.

What we're looking for

  • Formal training or certification on site reliability engineering concepts and 3+ years of applied experience.
  • Proficiency in site reliability culture, principles, and implementation within an application or platform.
  • Proficiency in at least one programming language such as Python, Java/Spring Boot, or .Net.
  • Working knowledge of using enterprise-authorized AI capabilities to support SRE workflows with strong validation habits and data sensitivity awareness.
  • Ability to validate AI-assisted operational recommendations before applying changes while following security requirements.
  • Proficiency in software applications and technical processes within a specific discipline such as Cloud, AI, or Android.
  • Practical experience in AWS supporting production services including visibility, troubleshooting, deployments, and operational hygiene.
  • Strong debugging fundamentals across distributed systems involving logs, metrics, latency analysis, and dependency failures.
  • Experience in time-sensitive environments with strong incident discipline and stakeholder communication.
  • Prior Markets experience, especially pricing, risk, or market data support (preferred).
  • Experience with Datadog for metrics, logs, traces, alert tuning, dashboards, and basic SLO concepts (preferred).
  • Familiarity with messaging/streaming patterns like Kafka/MQ and data quality checks in execution pipelines (preferred).
  • Experience applying AI-assisted tooling to reduce support toil (preferred).

More like this

Similar roles

Site Reliability Engineer III

JPMorgan Chase

Chicago, IL 45 days ago
Site Reliability Engineering Python Java Spring Boot .NET Cloud Infrastructure Network as Code Observability Telemetry AI SRE Monitoring
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Houston, TX 8 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI Cloud Terraform Kubernetes Docker Prometheus Grafana Dynatrace Datadog Splunk Jenkins GitLab Observability
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Houston, TX 52 days ago
Site Reliability Engineering Python Java Spring Boot .NET Network as Code Cloud AI Observability Monitoring
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Jersey City, NJ 76 days ago
Site Reliability Engineering Python PySpark AWS Databricks Snowflake Kubernetes AI/ML Grafana Prometheus Dynatrace Datadog Splunk Java Spring Boot .Net System Design
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Houston, TX 35 days ago
Site Reliability Engineering Python Java Spring Boot PySpark Cloud Infrastructure Network as Code Container Orchestration Observability Telemetry AI SRE
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Irvine, CA 37 days ago
Site Reliability Engineering Python Java Spring Boot .Net EKS EC2 Terraform Kubernetes CloudWatch Route 53 Load Balancing DNS mTLS Monitoring Telemetry
3+ yrs exp