This role pays less than
96%
of similar roles. Most pay
$143,475–$213,692
— the shaded band above.
At the midpoint, this role pays about
$110k
versus about
$179k
for comparable roles.
Based on 238 similar postings.
Employer
About RBC
RBC (Royal Bank of Canada) is Canada''s largest bank by market capitalization, offering a broad range of personal and commercial banking, wealth management, insurance, and capital markets services. Industry: Banking & Financial Services
RBC currently has
11 open roles
on FindRole.
Listed pay typically runs
$80,000–$140,000
across 8 roles with salary data.
The Senior Site Reliability Engineer joins the Wealth Management SRE Team to ensure the performance, availability, and resilience of critical platforms supporting wealth management services. This role involves building a robust SRE product base by developing intelligent monitoring, automated remediation, and ML-based anomaly detection systems. The engineer will implement modern observability practices, standardize telemetry across distributed systems, and automate operational workflows using Ansible, GitHub Actions, and scripting in Bash, Python, or PowerShell. Key technologies include Elasticsearch, Dynatrace, Kubernetes, OpenShift, Kafka, PagerDuty, and Moosmoft. Responsibilities include defining SLIs and SLOs, managing incident response, and performing root cause analysis to improve system reliability at scale. The role addresses the challenge of transitioning from reactive to predictive operations while ensuring high-quality digital services for internal users and clients within a complex technical environment.
What does a Site Reliability Engineer earn?
Median $188000 from 130 postings across 40 companies.
Develop intelligent monitoring, alerting, reliability testing, and automated remediation capabilities for critical applications.
Implement modern observability practices including metrics, logs, traces, and dashboards across supported platforms.
Design ML-based anomaly detection and self-healing solutions to transition from reactive to predictive operations.
Standardize telemetry and instrumentation to improve visibility and correlation of operational signals.
Automate operational workflows using Ansible, GitHub Actions, and scripting in Bash, Python, or PowerShell.
Define and track service health metrics including SLIs, SLOs, error budgets, and automated runbooks.
Lead incident management by troubleshooting production issues, participating in on-call rotations, and performing root cause analysis.
Drive continuous improvement by identifying opportunities to simplify and modernize operations using AI-driven approaches.
What we're looking for
5+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or Systems Engineering roles with strong operational depth.
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
Strong experience with infrastructure automation and configuration management, particularly Ansible.
Strong scripting and automation skills in Bash, Python, PowerShell, or similar languages.
Hands-on experience with modern reliability and observability tooling such as Elasticsearch, Dynatrace, GitHub, Kubernetes, OpenShift, Kafka, PagerDuty, Moogsoft, or related platforms.
Strong understanding of production operations, incident management, root cause analysis, and reliability engineering practices.
Experience defining and operating SLIs, SLOs, alerting strategies, and service health metrics.
Knowledge of cloud-native and distributed systems concepts, including resiliency, scalability, fault isolation, and performance tuning.
Understanding of AIOps, AI/ML concepts, or intelligent automation as applied to observability and operations.
Experience in financial services, wealth management, banking, insurance, or other highly regulated environments (preferred).
Experience with OpenTelemetry and telemetry standardization across distributed systems (preferred).
Hands-on experience with Prometheus, Grafana, Splunk, Catchpoint, Azure Automation, or similar SRE and observability platforms (preferred).
Experience with CI/CD and developer platform tools such as Jenkins, Artifactory, and Vault (preferred).
Familiarity with containerization and cloud platform patterns, including Docker and Kubernetes-based deployments (preferred).
Experience building or operating anomaly detection, predictive alerting, or self-healing automation solutions (preferred).
Familiarity with AI governance, model validation, and operational controls in regulated environments (preferred).