Lead Systems Operations Engineer

Wells Fargo

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Irving, TXCharlotte, NCChandler, AZ
Posted
9 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $175k
$134k most similar roles pay here $217k

This listing doesn't post a salary. Most similar roles pay $141,625–$209,041.

Based on 240 similar postings.

Employer

About Wells Fargo

Wells Fargo & Company is one of the largest banks in the United States, providing banking, investment, mortgage, and consumer and commercial finance products and services nationwide. Industry: Banking & Financial Services

Wells Fargo currently has 33 open roles on FindRole.

Listed pay typically runs $159,000–$260,000 across 13 roles with salary data.

Most-posted roles

View all roles at Wells Fargo

At a glance

TL;DR · Lead Systems Operations Engineer

The Lead Systems Operations Engineer joins the Consumer Technology organization to provide technical leadership for platform reliability, observability, and support readiness across critical consumer-facing applications. This role focuses on driving site reliability engineering practices, establishing SLOs and SLIs, and implementing error budget protocols to improve service availability. Day-to-day responsibilities include leading major incident management, conducting root cause analyses, and developing automated recovery solutions to eliminate single points of failure. The candidate will manage infrastructure-as-code initiatives and oversee vendor dependencies for third-party integrations. Technical requirements include proficiency with Splunk, Grafana, AppDynamics, Dynatrace, and GCP Monitoring. The role addresses the challenge of maintaining highly resilient, scalable, and observable systems within a regulated financial services environment, specifically focusing on ensuring stable customer journeys through proactive monitoring, automated workflows, and robust infrastructure engineering for distributed, cloud-based platforms.

What you'll do

  • Establish and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budget practices for critical platforms.
  • Improve platform resilience through automation, self-healing capabilities, capacity planning, and fault-tolerant designs.
  • Serve as the technical lead during major production incidents to provide coordination and recovery leadership.
  • Lead Root Cause Analysis (RCA) efforts and ensure corrective actions are implemented and tracked.
  • Lead enterprise observability initiatives using tools like Splunk, Grafana, and AppDynamics to define monitoring standards and alerting strategies.
  • Develop automation strategies and Infrastructure-as-Code (IaC) practices to reduce manual effort and improve consistency.
  • Manage vendor relationships by evaluating performance and ensuring operational readiness for third-party integrations.
  • Provide technical leadership and mentorship to SREs and systems engineers while influencing architecture decisions.

What we're looking for

  • 5+ years of experience in Systems Engineering or Technology Architecture.
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, or Production Support.
  • 5+ years supporting mission-critical production applications in large enterprise environments.
  • 3+ years leading major incident management, operational support, or reliability engineering initiatives.
  • 3+ years of experience with observability and monitoring platforms like Splunk, Grafana, AppDynamics, or GCP Monitoring.
  • 2+ years of experience driving automation, operational improvements, and reliability initiatives.
  • 3+ years of experience supporting distributed systems, cloud-based platforms, infrastructure, networking, and application architectures.
  • 1+ year of experience supporting highly regulated or customer-facing financial services platforms.

More like this

Similar roles

Lead Systems Operations Engineer

Wells Fargo

Charlotte, NC +1 11 days ago
SRE Kubernetes CI/CD Spring Boot Spring WebFlux MongoDB Redis Resilience4J Distributed Tracing Chaos Engineering BlazeMeter Chaos Monkey SWIFT CHIPS FedNow RTP
5+ yrs exp Hybrid

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 43 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp

Lead Principal Site Reliability Engineer

Oracle

Vienna, VA 53 days ago $96,300$264,100
Oracle Cloud Infrastructure (OCI) Kubernetes Terraform Python Bash Docker CI/CD Linux Administration Prometheus Grafana OpenSearch Splunk Datadog New Relic Infrastructure as Code Site Reliability Engineering Incident Management Root Cause Analysis Capacity Planning
10+ yrs exp