Lead Site Reliability Engineer

Morgan Stanley

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Alpharetta, GA
Salary
$125,000–$175,000 / yr
Posted
17 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $180k
This role $150k
$113k most similar roles pay here $237k

This role pays less than 74% of similar roles. Most pay $147,200–$213,692 — the shaded band above. At the midpoint, this role pays about $150k versus about $180k for comparable roles.

Based on 238 similar postings.

Employer

About Morgan Stanley

Morgan Stanley is a global financial services firm providing investment banking, securities, wealth management, and investment management services to corporations, governments, institutions, and individuals. Industry: Investment Banking & Financial Services

Morgan Stanley currently has 38 open roles on FindRole.

Listed pay typically runs $150,000–$210,000 across 31 roles with salary data.

Most-posted roles

View all roles at Morgan Stanley

At a glance

TL;DR · Lead Site Reliability Engineer

Lead Site Reliability Engineer is a Vice President level role within the Technology division's Reliability Operations team. This position focuses on ensuring the stability, performance, and reliability of Wealth Management Investment Management application platforms. You will champion SRE principles such as SLIs, SLOs, and error budgets while leading initiatives to reduce MTTR and MTTI. Daily responsibilities include proactive detection and resolution of production issues, managing change controls, automating manual tasks using AI-driven workflows, and developing comprehensive knowledge management frameworks. The role requires expertise in Python, Shell, or Perl scripting, along with experience in AWS, GCP, or Azure cloud environments. You will utilize tools like Grafana, Prometheus, Splunk, Kibana, and Autosys to monitor high-availability systems. Key technical competencies include database skills in DB2, Sybase, or Oracle, and deep analytical triage for complex infrastructure and software troubleshooting.

What you'll do

  • Champion SRE principles by defining and tracking metrics like SLIs, SLOs, and error budgets.
  • Lead the detection, triage, and resolution of production issues for business-critical applications.
  • Manage high-pressure incident responses while providing clear communication to senior leadership during outages.
  • Enforce production governance by ensuring all changes meet risk management and operational readiness standards.
  • Automate manual tasks and "toil" using scripting, self-healing capabilities, and AI-driven workflows.
  • Maintain a comprehensive knowledge management framework including runbooks, troubleshooting guides, and system diagrams.
  • Provide technical leadership on architecture reviews, capacity planning, and disaster recovery strategies.

What we're looking for

  • Bachelor’s or Master’s degree in a quantitative discipline like Computer Science or Computer Engineering.
  • 10+ years of experience in a production environment with a background in software development and performance tuning.
  • 5+ years of experience leading a small to medium team with similar skill sets.
  • 5+ years of experience driving SRE principles and Chaos Engineering.
  • Proficiency in scripting languages such as Shell, Python, or Perl and cloud-driven development.
  • Experience with AWS, GCP, or Azure cloud technologies and database systems like DB2, Sybase, or Oracle.
  • Knowledge of DevOps and observability tools including Grafana, Prometheus, Splunk, and Kibana.
  • Ability to manage 24/7 on-call rotations and coordinate high-pressure outage incidents.

More like this

Similar roles

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 43 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 45 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 43 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Principal Site Reliability Engineer

Oracle

Nashville, TN 57 days ago $96,300$264,100
Site Reliability Engineering Kubernetes Docker Terraform Ansible Chef Puppet Python Go Java JavaScript Bash Oracle Cloud Infrastructure Microsoft Azure Google Cloud Platform infrastructure-as-code Chaos Engineering
6+ yrs exp