Observability SRE Manager

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Seattle, WA
Salary
$225,600–$338,400 / yr
Posted
45 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $229k
This role $282k
$155k most similar roles pay here $358k

This role pays more than 95% of similar roles. Most pay $202,800–$254,750 — the shaded band above. At the midpoint, this role pays about $282k versus about $229k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Observability SRE Manager

The Observability SRE Manager, Apple Services Engineering, is a senior leadership role focused on the reliability and performance of the core observability platform. You will lead a team of Site Reliability Engineers to manage metrics, logging, tracing, and alerting infrastructure that serves as the central nervous system for cloud services. Responsibilities include defining strategic roadmaps, establishing SRE practices like SLOs and error budgets, driving automation to reduce toil, and integrating AI/ML tools for anomaly detection and root-cause analysis. The role requires expertise in multi-tenant distributed systems, Kubernetes, and full-stack troubleshooting across network, OS, and application layers. Key technologies include Prometheus, OpenTelemetry, Python, Go, Java, Scala, and infrastructure as code tools like Terraform. You will manage the lifecycle of monitoring systems to ensure high availability for critical services while mentoring a high-performing team.

What you'll do

  • Own the reliability, availability, and performance of the metrics, logging, tracing, and alerting infrastructure.
  • Define and execute a strategic roadmap for observability tools in partnership with engineering and product stakeholders.
  • Establish SRE practices including SLOs, error budgets, capacity planning, and disaster recovery protocols.
  • Drive automation to eliminate manual processes through tooling, self-service platforms, and APIs.
  • Execute an AI strategy to integrate machine learning into root-cause analysis and toil reduction.
  • Lead, mentor, and grow a team of Site Reliability Engineers while managing performance and career development.
  • Partner with cross-functional teams to influence system design for scalability and operational readiness.
  • Communicate technical strategies, trade-offs, and progress reports to executive leadership.

What we're looking for

  • Minimum of 5 years of engineering management experience leading SRE, infrastructure, or observability teams.
  • Experience hiring, growing, and mentoring a team of engineers.
  • Deep understanding of observability systems including metrics, logging, tracing, alerting, SLOs, and error budgets.
  • Strong systems background with ability to troubleshoot across network, OS, container runtime, and application layers.
  • Experience operating large-scale, multi-tenant distributed systems in production, including Kubernetes environments.
  • Proficiency in shell/bash scripting and at least one high-level language like Python, Go, Java, or Scala.
  • Demonstrated experience applying AI/ML tooling or LLM-based solutions to improve SRE or infrastructure operations.
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or equivalent experience.

More like this

Similar roles

Software Engineering Manager

Apple Inc

Cupertino, CA 27 days ago $237,600$356,400
SRE Kubernetes Infrastructure as Code Terraform Ansible Chef Salt Python Go CI/CD Prometheus Grafana OpenStack KVM Distributed Systems Capacity Planning Incident Management Batch Compute Machine Learning Distributed Tracing
5+ yrs exp

Senior SRE Software Engineer

Apple Inc

San Francisco, CA 38 days ago $184,700$324,800
Kubernetes Go SRE DevOps Prometheus Thanos Splunk Puppet Ansible AWS GCP Azure Bare-metal Container Runtime
8+ yrs exp

Software Engineering Manager, Apple Services Engineering

Apple Inc

New York, NY 79 days ago $206,200$356,400
Distributed Systems Microservices RESTful Web Services Object-Oriented Design NoSQL Hibernate JPA JUnit Mockito Solr Elastic Search Redis Memcached Cassandra Voldemort CI/CD DevOps Agile Sprint Planning