Reliability Engineer III, Observability Specialist

US Bank

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Brookfield, WIAtlanta, GAHopkins, MNCupertino, CACharlotte, NC
Salary
$98,175–$115,500 / yr
Posted
16 days ago
Freshness
Confirmed live yesterday
Closes
Sep 21, 2026

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $172k
This role $107k
$84k most similar roles pay here $229k

This role pays less than 95% of similar roles. Most pay $136,493–$208,350 — the shaded band above. At the midpoint, this role pays about $107k versus about $172k for comparable roles.

Based on 240 similar postings.

Employer

About US Bank

U.S. Bank (U.S. Bancorp) is the fifth-largest bank in the United States, providing retail banking, corporate and commercial banking, wealth management, and payment services to millions of customers. Industry: Banking & Financial Services

US Bank currently has 30 open roles on FindRole.

Listed pay typically runs $119,765–$140,900 across 29 roles with salary data.

Most-posted roles

View all roles at US Bank

At a glance

TL;DR · Reliability Engineer III, Observability Specialist

Reliability Engineer 3 (Observability Specialist) serves as a senior-level technical lead responsible for enabling reliable, measurable, and supportable application operations across a broad portfolio of production applications. The role involves partnering with product owners and engineering teams to translate customer journeys into measurable reliability objectives by designing and governing SLIs, SLOs, error budgets, and telemetry standards. You will build and maintain executive dashboards, manage synthetic monitoring, and perform incident analysis to reduce alert fatigue while ensuring high service availability. Key technical requirements include expertise in distributed systems, microservices, and Kubernetes, alongside proficiency with tools such as Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry. The role focuses on solving complex reliability challenges by providing actionable visibility into service health, performance trends, and failure conditions to improve overall system stability.

What you'll do

  • Design and govern Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for enterprise applications.
  • Translate business requirements into scalable observability architectures including instrumentation standards and alerting models.
  • Develop and maintain executive dashboards to communicate service health, latency, and customer impact.
  • Identify observability gaps by analyzing telemetry data, incident trends, and alert performance.
  • Provide technical leadership and mentorship on distributed tracing, logging, metrics, and synthetic monitoring.
  • Establish governance frameworks for the lifecycle management of dashboards, alerts, and telemetry standards.
  • Partner with engineering teams to ensure applications are fully instrumented for production readiness.
  • Maintain an authoritative inventory of all observability assets and compliance documentation.

What we're looking for

  • Bachelor's degree or equivalent work experience.
  • Five to seven years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development.
  • Expertise in Observability Engineering, Site Reliability Engineering (SRE), or Reliability Engineering (preferred).
  • Strong knowledge of SLIs, SLOs, Error Budgets, and Customer Journey Monitoring (preferred).
  • Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry frameworks (preferred).
  • Proficiency with Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry (preferred).
  • Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes (preferred).
  • Excellent stakeholder management, communication, and technical leadership skills (preferred).

More like this

Similar roles

Reliability Engineer 3, Observability Specialist

US Bank

Brookfield, WI +4 8 days ago $98,175$115,500
Observability Site Reliability Engineering (SRE) SLIs SLOs Error Budgets Distributed Tracing Logging Metrics Synthetic Monitoring APM RUM Datadog Dynatrace Splunk Grafana Prometheus New Relic Elastic OpenTelemetry Kubernetes
5+ yrs exp

Senior Observability Engineer

RBC

Minneapolis, MN 29 days ago $90,000$140,000
ELK Stack Dynatrace Prometheus Grafana Open Telemetry Jaeger Python Java Go Linux Unix Kubernetes OpenShift Tableau Anthropic OpenAI Copilot SRE

Site Reliability Engineering Lead

US Bank

Atlanta, GA +2 3 days ago $111,605$131,300
SRE DevOps AWS Azure Kubernetes Docker Terraform Ansible Python PowerShell Shell Scripting CI/CD GitHub Actions Azure DevOps Jenkins GitLab Datadog Splunk Dynatrace Grafana Prometheus CloudWatch Azure Monitor OpenTelemetry ServiceNow Jira REST APIs SQL
6+ yrs exp

Senior Site Reliability Engineer

CVS Health

Woonsocket, RI +1 14 days ago $92,700$203,940
SRE DevOps Kubernetes OpenShift Docker CI/CD Splunk Dynatrace Datadog Prometheus Grafana Python Java AWS Microsoft Azure Google Cloud Rancher GitHub BitBucket Jenkins Microservices web API’s Apigee
5+ yrs exp Hybrid

Reliability Engineer

Anduril Industries

Atlanta, GA +1 79 days ago $126,000$167,000
FMEA FTA Weibull Analysis MIL-HDBK-217 MIL-HDBK-472 MIL-STD-810 MIL-STD-461 MIL-STD-516C MIL-STD-1629 HALT HASS HITL SITL Root Cause Analysis CAPA Technical Writing Data Acquisition Environmental Testing Mechanical Testing
5+ yrs exp