Senior Manager, Site Reliability Engineering

Intuit

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Mountain View, CA
Salary
$222,000–$300,500 / yr
Posted
21 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $190k
This role $261k
$128k most similar roles pay here $319k

This role pays more than 95% of similar roles. Most pay $162,291–$217,668 — the shaded band above. At the midpoint, this role pays about $261k versus about $190k for comparable roles.

Based on 238 similar postings.

Employer

About Intuit

Intuit is a financial software company known for products like TurboTax, QuickBooks, Mint, and Credit Karma, helping consumers and small businesses manage their finances and taxes. Industry: Financial Software & Technology

Intuit currently has 174 open roles on FindRole.

Listed pay typically runs $202,500–$274,000 across 152 roles with salary data.

Most-posted roles

View all roles at Intuit

At a glance

TL;DR · Senior Manager, Site Reliability Engineering

As a Senior Manager, Site Reliability Engineering, you will lead a team of 10–15 engineers within the Fintech Platform Systems Engineering team to ensure the availability and performance of money-movement services. This player-coach role involves managing people, setting technical direction, and overseeing infrastructure for high-availability systems requiring 99.999% uptime. You will build self-healing infrastructure, manage incident response, and execute an AI Ops roadmap to automate detection and remediation while reducing developer toil. The role requires expertise in AWS services like EC2, EKS, and RDS, alongside tools such as Terraform, Datadog, and PagerDuty. You will solve complex reliability challenges within a regulated fintech environment involving payment systems and high-trust data integrity. Key competencies include distributed systems, container orchestration, infrastructure-as-code, chaos engineering, and experience in highly regulated domains like banking or payments.

What you'll do

  • Manage a team of 10–15 engineers by hiring, mentoring, and setting goals to develop technical leaders.
  • Execute an AI Ops roadmap to implement autonomous detection, diagnosis, and remediation in production systems.
  • Drive the strategy for achieving and maintaining 99.999% availability for fintech platform services.
  • Act as a player-coach by participating in architecture reviews, reviewing code/IaC, and joining incident bridges.
  • Lead incident management processes, including command, high-severity response, and conducting blameless postmortems.
  • Build and scale resilient AWS infrastructure with a focus on auto-remediation and multi-region failover.
  • Identify and replace high-toil workflows with automated agents to reduce manual engineering effort.
  • Define and report on SLOs, SLIs, and error budgets to prioritize reliability investments.

What we're looking for

  • 8+ years of experience in systems engineering, site reliability engineering, or infrastructure engineering.
  • 3+ years of experience directly managing engineering teams.
  • Proven hands-on experience operating production infrastructure in AWS at scale.
  • Track record of driving high-availability outcomes for mission-critical systems, preferably in fintech or regulated domains.
  • Deep technical foundation in distributed systems, networking, Kubernetes, and infrastructure-as-code.
  • Experience with observability tools, building SLO-driven operations, and incident management processes.
  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
  • Experience designing/scaling AI Ops capabilities (AIOps platforms, ML-based detection) to reduce toil or MTTR (preferred).

More like this

Similar roles

Senior Engineering Manager, Site Reliability

Upstart

Remote (Canada) 56 days ago $195,300$270,400
Site Reliability Engineering Distributed Systems Cloud Infrastructure Datadog Grafana Prometheus OpenTelemetry Kubernetes AWS Incident Management Observability Service Level Objectives Postmortem Cloud Native Architecture
7+ yrs exp Remote

Manager, Site Reliability Engineering

Okta Inc

San Francisco, CA 43 days ago $204,000$306,000
AWS Kubernetes Terraform CI/CD Grafana Splunk APM DevOps SaaS Cloud-native Architecture Containerization Edge Networking SDLC
3+ yrs exp Hybrid

Senior Site Reliability Engineer

Salesforce

Remote (San Francisco, CA) 58 days ago $148,500$223,900
SRE Python Go Docker Kubernetes CI/CD Prometheus Grafana ELK Splunk Datadog Temporal Airflow Argo Workflows AWS GCP Linux Unix LLM Prompt Engineering
5+ yrs exp Remote

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Senior Manager, Site Reliability Engineering

Oracle

Nashville, TN 31 days ago $121,500$264,100
AIOps Automation Business Intelligence Data Center Operations Incident Management Network Operations Network Routing Structured Cabling Deployment Telecommunications Networking
3+ yrs exp