Senior Lead Site Reliability Engineer, Manager AI/ML and Data Platforms

JPMorgan Chase

Confirmed live yesterday Low trust

Quick summary

Work type
On-site
Location
Jersey City, NJ
Posted
59 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $198k
$146k most similar roles pay here $257k

This listing doesn't post a salary. Most similar roles pay $164,500–$231,150.

Based on 240 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1117 open roles on FindRole.

Listed pay typically runs $186,160–$215,000 across 7 roles with salary data.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Senior Lead Site Reliability Engineer, Manager AI/ML and Data Platforms

As a Senior Lead Site Reliability Engineer within the Chief Data & Analytics Office AI/ML & Data Platforms team, you will define non-functional requirements and availability targets for large-scale data platforms and data lake ecosystems. You will design roadmaps, manage distributed systems initiatives, and implement observability designs to ensure stable performance for analytics and AI/ML workloads. Your daily work involves mentoring technologists, driving the evolution of critical components, and integrating AI-assisted reliability workflows into the software development lifecycle while ensuring security and auditability. Required skills include expertise in site reliability principles, SLIs, SLOs, and error budgets. You will utilize tools such as Grafana, Dynatrace, Prometheus, Datadog, and Splunk. Preferred technical competencies include AWS, Databricks, Spark, Docker, Kubernetes, Terraform, and Python to manage complex data pipelines and infrastructure in distributed environments.

What you'll do

  • Define non-functional requirements and availability targets for large-scale data platforms and AI/ML workloads.
  • Create high-quality designs, roadmaps, and program charters for distributed systems initiatives.
  • Implement observability and reliability designs to ensure stable performance without increasing technical debt.
  • Integrate enterprise-authorized AI capabilities into reliability workflows to accelerate design and operational decision-making.
  • Establish team practices for safe, secure, and auditable AI usage in production operations.
  • Debug and evolve critical components by analyzing application and infrastructure interdependencies.
  • Provide tools and guidance to support scalable data platform infrastructure and engineering best practices.
  • Contribute to the firm's site reliability community through internal forums and technical leadership.

What we're looking for

  • Formal training or certification in site reliability engineering concepts and 5+ years of applied experience.
  • Advanced understanding of site reliability culture, principles, SLI/SLO/SLA, and error budgets for large-scale data systems.
  • Advanced knowledge of observability tools such as Grafana, Dynatrace, Prometheus, Datadog, and Splunk.
  • Demonstrated experience using enterprise-authorized AI capabilities to improve reliability engineering workflows with strong validation habits.
  • Ability to establish team practices for safe AI usage in operations while maintaining security, auditability, and risk controls.
  • Advanced knowledge of distributed systems, system design, resiliency, testing, and disaster recovery.
  • Strong communication skills to report data-based solutions and mentor others on site reliability principles.
  • Experience with cloud platforms (AWS), containerization (Docker, Kubernetes), and CI/CD automation tools like Terraform.

More like this

Similar roles

Senior Lead Site Reliability Engineer

JPMorgan Chase

Palo Alto, CA 59 days ago
Site Reliability Engineering Java Go Python Terraform Kubernetes Docker CI/CD GitOps Grafana Prometheus Dynatrace Datadog Splunk Kafka RabbitMQ SQS Neo4j Pinecone Weaviate Chroma LangChain LangGraph AutoGen CrewAI GitHub Copilot Fluentd Logstash Vector RESTful APIs RAG TensorFlow PyTorch scikit-learn Hadoop Spark Flink MongoDB Cassandra DynamoDB InfluxDB TimescaleDB AWS Azure GCP
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 45 days ago
Site Reliability Engineering Observability Monitoring Telemetry Service Level Objectives Alerting AI SDLC Automation
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 43 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 45 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp