CaaS Private Site Reliability Engineer, Assistant Vice President

Deutsche Bank

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Cary, NC
Salary
$100,000–$153,000 / yr
Posted
23 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $179k
This role $126k
$85k most similar roles pay here $239k

This role pays less than 91% of similar roles. Most pay $144,975–$214,000 — the shaded band above. At the midpoint, this role pays about $126k versus about $179k for comparable roles.

Based on 240 similar postings.

Employer

About Deutsche Bank

Deutsche Bank is a German multinational investment bank and financial services company offering corporate banking, investment banking, retail banking, asset management, and transaction banking worldwide. Industry: Investment Banking & Financial Services

Deutsche Bank currently has 18 open roles on FindRole.

Listed pay typically runs $125,000–$165,000 across 18 roles with salary data.

Most-posted roles

View all roles at Deutsche Bank

At a glance

TL;DR · CaaS Private Site Reliability Engineer, Assistant Vice President

CaaS Private Site Reliability Engineer - Assistant Vice President will join the CaaS Private platform team to operate and improve an on-prem, multi-tenant Kubernetes platform running on bare metal. The role involves enhancing reliability, observability, scalability, and operational excellence for critical, low-latency, and regulated workloads. Key responsibilities include defining SLIs/SLOs, building monitoring dashboards, leading incident responses, and automating repetitive tasks to reduce toil. You will manage infrastructure components like ingress paths, service mesh, and node services while collaborating with network and security teams on capacity planning. Required skills include Kubernetes expertise, Linux administration, and scripting in Python, Ansible, and Bash. The role requires proficiency with Prometheus, Grafana, Splunk, and OpenTelemetry concepts. Candidates should understand distributed systems, networking, and containerization, with experience in Istio, Envoy, and supporting stateful services like PostgreSQL or Kafka.

What you'll do

  • Define and improve SLIs, SLOs, alerting standards, and error budgets for the CaaS Private platform.
  • Build and maintain observability across metrics, logs, alerts, and dashboards to monitor platform health and performance.
  • Lead incident response for platform events, including mitigation, communication, and conducting blameless postmortems.
  • Automate repetitive operational tasks and remediation workflows to reduce toil and improve recovery times.
  • Improve reliability and upgrade safety for Kubernetes clusters, ingress paths, service mesh components, and node services.
  • Perform capacity planning, troubleshooting, and technical documentation in coordination with network and security teams.

What we're looking for

  • Hands-on experience operating Kubernetes clusters on bare metal or private cloud environments.
  • Proven experience in Site Reliability Engineering, Production Engineering, DevOps, or related infrastructure roles.
  • Strong Linux system administration skills and infrastructure scripting using Python, Ansible, and Bash.
  • Practical knowledge of observability stacks including Prometheus, Grafana, Splunk, and Open Telemetry concepts.
  • Understanding of incident management, root cause analysis, networking, virtualization, and distributed systems.
  • Ability to leverage AI tools for productivity while ensuring responsible and ethical use of data.
  • Experience with Istio/Envoy, service mesh, OPA Gatekeeper, or policy-driven operational guardrails.
  • Ability to read and understand Golang code when troubleshooting platform components.

More like this

Similar roles

Site Reliability Engineer, Enterprise Technology Services

Apple Inc

Sunnyvale, CA 28 days ago $150,400$277,600
SRE DevOps Java Python Bash LUA Oracle MongoDB Prometheus Splunk Grafana CloudWatch Linux Networking TLS/SSL DNS Load Balancers Git CI/CD Kubernetes AWS GCP Nginx Envoy NetScaler
5+ yrs exp

Senior Site Reliability Engineer, AIOPs

Nvidia

Santa Clara, CA 122 days ago $148,000$235,750
Kubernetes Python Bash Terraform Helm CI/CD infrastructure-as-code Prometheus Grafana Kafka Pulsar Flink Spark ClickHouse Linux Distributed Systems Microservices SRE AIOps
5+ yrs exp