CaaS Private Site Reliability Lead Engineer, Vice President

Deutsche Bank

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Cary, NC
Salary
$125,000–$185,000 / yr
Posted
23 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $183k
This role $155k
$114k most similar roles pay here $231k

This role pays less than 73% of similar roles. Most pay $152,150–$213,817 — the shaded band above. At the midpoint, this role pays about $155k versus about $183k for comparable roles.

Based on 240 similar postings.

Employer

About Deutsche Bank

Deutsche Bank is a German multinational investment bank and financial services company offering corporate banking, investment banking, retail banking, asset management, and transaction banking worldwide. Industry: Investment Banking & Financial Services

Deutsche Bank currently has 18 open roles on FindRole.

Listed pay typically runs $125,000–$165,000 across 18 roles with salary data.

Most-posted roles

View all roles at Deutsche Bank

At a glance

TL;DR · CaaS Private Site Reliability Lead Engineer, Vice President

CaaS Private Site Reliability Lead Engineer - Vice President leads the reliability, resilience, and operational excellence for the CaaS Private platform in the US. This role involves bringing production engineering discipline to Kubernetes, observability, automation, and incident management while improving service health and platform readiness. The engineer will define SLO frameworks, manage complex troubleshooting, implement self-healing workflows, and oversee capacity planning and disaster readiness. Key responsibilities include mentoring junior engineers, refining alert thresholds, and ensuring platform changes are measurable and supportable. Required technical skills include extensive experience with Kubernetes, Linux, distributed systems reliability, and production platform operations. The candidate must demonstrate proficiency in monitoring, alerting, dashboarding, and infrastructure-as-code practices to solve critical problems regarding platform stability and automated recovery within a cloud-native environment while collaborating across engineering and application teams to meet enterprise standards.

What you'll do

  • Lead the reliability strategy for the CaaS Private platform including SLO frameworks and incident management maturity.
  • Drive resilience improvements across observability, capacity planning, upgrade safety, and disaster readiness.
  • Resolve complex production issues by leading troubleshooting, identifying root causes, and implementing preventive fixes.
  • Define and refine service indicators, alert thresholds, dashboard standards, and production readiness criteria.
  • Develop automation and self-healing workflows to reduce manual intervention and improve recovery times.
  • Influence platform architecture and roadmap decisions using data from incidents and reliability metrics.
  • Mentor junior and middle engineers while fostering a culture of blameless learning and operational excellence.
  • Communicate strategic insights and technical findings to both technical and non-technical stakeholders.

What we're looking for

  • Extensive experience with Bare Metal Kubernetes, Linux, distributed systems reliability, and production platform operations.
  • Strong hands-on experience with observability, monitoring, alerting, dashboarding, incident response, and root cause analysis.
  • Proven ability to design and implement automation, self-healing workflows, operational checks, and runbook improvements.
  • Experience defining SLOs, service indicators, alert quality standards, production readiness practices, and escalation models.
  • Ability to lead complex reliability improvements independently while partnering across engineering, operations, and application teams.
  • Proven ability to leverage AI tools to enhance productivity and optimize workflows with critical judgment.
  • Strong communication skills to explain technical findings to both technical and non-technical stakeholders.
  • Experience mentoring engineers and improving operational culture across platform or infrastructure teams.

More like this

Similar roles

Site Reliability Engineer, Lead

Booz Allen Hamilton

Chantilly, VA 50 days ago $99,000$225,000
Prometheus Grafana ELK Stack Linux AWS Python Terraform Terragrunt Kubernetes OpenTelemetry AWS CloudWatch AWS EKS Rancher Jenkins Git Docker Nessus JIRA Confluence SRE
8+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 45 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Lead Site Reliability Engineer

Mastercard

O Fallon, MO 16 days ago $122,000$207,000
Site Reliability Engineering Java Spring Framework Python Go DevOps CI/CD Configuration Management Distributed Systems Automation Observability ITSM Root Cause Analysis Capacity Planning Monitoring

Lead Principal Site Reliability Engineer

Oracle

Vienna, VA 53 days ago $96,300$264,100
Oracle Cloud Infrastructure (OCI) Kubernetes Terraform Python Bash Docker CI/CD Linux Administration Prometheus Grafana OpenSearch Splunk Datadog New Relic Infrastructure as Code Site Reliability Engineering Incident Management Root Cause Analysis Capacity Planning
10+ yrs exp