Principal Site Reliability Engineering Manager

Microsoft

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Salary
$142,800–$274,800 / yr
Posted
15 days ago
Freshness
Confirmed live 2 days ago
Closes
Feb 23, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $184k
This role $209k
$127k most similar roles pay here $291k

This role pays more than 71% of similar roles. Most pay $152,971–$214,875 — the shaded band above. At the midpoint, this role pays about $209k versus about $184k for comparable roles.

Based on 238 similar postings.

Employer

About Microsoft

Microsoft Corporation is a global technology leader producing software, hardware, and cloud services including Windows, Office 365, Azure cloud platform, Xbox gaming, and Surface devices. Industry: Software & Cloud Computing

Microsoft currently has 598 open roles on FindRole.

Listed pay typically runs $119,800–$234,700 across 586 roles with salary data.

Most-posted roles

View all roles at Microsoft

At a glance

TL;DR · Principal Site Reliability Engineering Manager

As a Principal Site Reliability Engineering Manager, you will lead a team of site reliability engineers responsible for building and operating Substrate services within highly regulated environments. You will manage the operational health, reliability posture, and incident management for these critical systems while mentoring senior and principal engineers. Your daily work involves establishing SLIs, SLOs, and operational metrics to improve diagnosability and service resilience through automation and AI-assisted techniques. You will partner with feature engineering teams to embed security and compliance into early design phases and oversee disaster recovery exercises and playbooks. The role requires expertise in software engineering fundamentals, including clean code and robust telemetry, to solve complex reliability challenges for infrastructure providing shared identity, messaging, and storage capabilities. Candidates must possess a relevant degree and an active U.S. Government Secret Security Clearance.

What you'll do

  • Manage and coach a team of Site Reliability Engineers across senior and principal levels.
  • Oversee the operational health and reliability of Substrate services in regulated environments.
  • Establish and monitor SLIs, SLOs, and other metrics to improve service reliability and diagnosability.
  • Lead incident management and post-incident reviews to implement systemic fixes and automation.
  • Participate in an on-call rotation as a hands-on leader during critical incidents.
  • Manage disaster recovery by coordinating game days and maintaining auditable operational playbooks.
  • Drive engineering-led initiatives using automation and AI to reduce operational toil at scale.
  • Partner with product teams to embed security, compliance, and reliability into early service designs.

What we're looking for

  • A Doctorate degree with 2+ years of experience in software engineering, network engineering, or systems administration is required.
  • A Master's degree with 3+ years of experience in software engineering, network engineering, or systems administration is required.
  • A Bachelor's degree with 5+ years of experience in software engineering, network engineering, or systems administration is required.
  • Candidates must possess an active U.S. Government Secret Security Clearance.
  • Candidates must be able to provide verification of U.S. citizenship.
  • Preferred experience includes at least 7 years working with large-scale cloud or distributed systems.
  • Preferred experience includes at least 3 years of people management experience.
  • Preferred experience includes operating services in regulated, sovereign, or compliance-sensitive environments.

More like this

Similar roles

Principal Site Reliability Engineer

Microsoft

Redmond, WA 94 days ago $142,800$274,800
Site Reliability Engineering SRE Distributed Systems Automation Observability SLO Incident Response Software Engineering Network Engineering Systems Administration Microsoft Government Cloud GCC High CJIS Compliance-Sensitive Environments
10+ yrs exp

Senior Site Reliability Engineer

Microsoft

115 days ago $119,800$234,700
Site Reliability Engineering Distributed Systems C# Go Java Python CI/CD Automation Monitoring Alerting Debugging Troubleshooting Network Engineering Systems Administration Software Engineering
4+ yrs exp

Senior Site Reliability Engineering

Microsoft

Redmond, WA 175 days ago $119,800$234,700
Site Reliability Engineering Distributed Systems Automation Monitoring Alerting Debugging Incident Response Postmortems Network Engineering Systems Administration Software Engineering Microsoft Government Cloud GCC High GCCH CJIS
4+ yrs exp

Site Reliability Engineer II

Microsoft

100 days ago $102,100$202,200
Site Reliability Engineering Distributed Systems C# Go Java Python CI/CD Automation Incident Response Monitoring Observability Network Engineering Systems Administration
2+ yrs exp

Site Reliability Engineer II

Microsoft

14 days ago $102,100$202,200
Site Reliability Engineering Python C# Java C++ JavaScript C Cloud Engineering Distributed Systems Automation Telemetry Pipelines Monitoring Incident Management Observability Data-Driven Operations
2+ yrs exp

Director, Site Reliability Engineering

Anduril Industries

Costa Mesa, CA 25 days ago $253,000$336,000
Site Reliability Engineering Distributed Systems Cloud Infrastructure Container Orchestration Observability Incident Management Capacity Planning Disaster Recovery Networking Storage ERP MES WMS CRM AI Infrastructure
10+ yrs exp