Site Reliability Engineer III

JPMorgan Chase

Confirmed live 2 days ago Low trust

Quick summary

Work type
On-site
Location
Jersey City, NJ
Posted
76 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

How this pay compares to similar roles

Similar $179k
$131k most similar roles pay here $230k

This listing doesn't post a salary. Most similar roles pay $144,423–$213,125.

Based on 238 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1117 open roles on FindRole.

Listed pay typically runs $186,160–$215,000 across 7 roles with salary data.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Site Reliability Engineer III

As a Site Reliability Engineer III within the Chief Data & Analytics Office AI/ML & Data Platforms team, you will manage and optimize complex infrastructure to ensure application reliability and scalability. You will build automated continuous integration and delivery pipelines, implement infrastructure as code, and develop AI/ML solutions for incident resolution and troubleshooting. Your daily work involves monitoring systems using tools like Grafana, Dynatrace, Prometheus, Datadog, and Splunk while managing production incident calls. You will utilize technologies including AWS, Databricks, Snowflake, and Kubernetes to improve service level objectives. To succeed, you must be proficient in Python, Java/Spring Boot, or .Net, with specific expertise in PySpark for automation. The role focuses on solving complex business problems by improving operational stability, reducing toil through intelligent automation, and managing high-availability systems within a data-centric environment.

What does a Site Reliability Engineer earn?

Median $186200 from 131 postings across 36 companies.

See salary data

What you'll do

  • Configure, maintain, and optimize application infrastructure using code-based configurations for network and systems.
  • Develop and implement automated CI/CD pipelines to support application development and production environments.
  • Utilize enterprise AI tools to accelerate incident triage, troubleshooting, and post-incident analysis while ensuring data security.
  • Build Python or PySpark tools to automate repetitive tasks and reduce operational toil in AI/ML workflows.
  • Monitor system health using observability tools like Grafana, Prometheus, and Datadog to manage SLIs and SLOs.
  • Manage production incident calls and coordinate cross-functional teams to resolve critical issues.
  • Design and implement scalable infrastructure solutions using technologies like AWS, Databricks, Snowflake, and Kubernetes.
  • Identify patterns in operational signals to proactively address reliability risks before they impact customers.

What we're looking for

  • Formal training or certification in site reliability engineering concepts and 3+ years of applied experience.
  • Proficiency in site reliability culture, principles, and a strong understanding of SLI/SLO/SLA and error budgets.
  • Proficiency in at least one programming language such as Python, Java/Spring Boot, or .Net.
  • Experience using Python or PySpark for AI/ML modeling and automation to reduce operational toil.
  • Working knowledge of enterprise-authorized AI capabilities to support SRE workflows while following data sensitivity requirements.
  • Hands-on experience in system design, resiliency, testing, operational stability, and disaster recovery.
  • Experience in observability using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk.
  • Experience with CI/CD tooling, container orchestration, and managing production incident calls.

More like this

Similar roles

Site Reliability Engineer III

JPMorgan Chase

Houston, TX 8 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI Cloud Terraform Kubernetes Docker Prometheus Grafana Dynatrace Datadog Splunk Jenkins GitLab Observability
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Irvine, CA 37 days ago
Site Reliability Engineering Python Java Spring Boot .Net EKS EC2 Terraform Kubernetes CloudWatch Route 53 Load Balancing DNS mTLS Monitoring Telemetry
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Chicago, IL 45 days ago
Site Reliability Engineering Python Java Spring Boot .NET Cloud Infrastructure Network as Code Observability Telemetry AI SRE Monitoring
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Houston, TX 52 days ago
Site Reliability Engineering Python Java Spring Boot .NET Network as Code Cloud AI Observability Monitoring
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Chicago, IL 8 days ago
SRE Python Ansible Terraform Kubernetes Docker AWS Prometheus Grafana Dynatrace Datadog Splunk Linux Windows Jenkins GitLab Network as Code
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Chicago, IL 8 days ago
SRE Python Ansible Terraform Kubernetes Docker AWS Prometheus Grafana Dynatrace Datadog Splunk Linux Windows Jenkins GitLab Network as Code
3+ yrs exp