Evaluation Reliability SRE

Apple Inc

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Cupertino, CA
Salary
$216,200–$324,800 / yr
Posted
107 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $180k
This role $270k
$120k most similar roles pay here $347k

This role pays more than 96% of similar roles. Most pay $145,000–$215,225 — the shaded band above. At the midpoint, this role pays about $270k versus about $180k for comparable roles.

Based on 238 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 2000 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1613 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Evaluation Reliability SRE

As an Evaluation Reliability SRE within the Siri organization, you will join the Evaluation Reliability Engineering team to ensure the production backbone of machine learning evaluation infrastructure remains bulletproof. This senior role focuses on maintaining reliability across orchestration, capacity, and service health while managing resource management, session orchestration, and observability systems. You will be responsible for authoring high-quality runbooks, diagnosing complex failure modes in device orchestration and provisioning layers, and implementing automation to eliminate recurring failures. The role requires expertise in distributed systems, Kubernetes, and infrastructure reliability. You will also manage on-call rotations, lead incident investigations, and define SLOs and burn-rate alerting. This position addresses the critical challenge of ensuring that evaluation signals remain trustworthy for model and product decisions within a large-scale AI assistant ecosystem.

What you'll do

  • Manage reliability outcomes across the evaluation infrastructure stack, including orchestration, capacity, and service health.
  • Author high-quality runbooks for complex failure categories to guide team response standards.
  • Diagnose and resolve issues within device orchestration and provisioning layers, such as quota management and retry behaviors.
  • Instrument infrastructure components to improve observability and detect failures before they impact evaluation signals.
  • Develop automation and eliminate recurring failures to balance incident response with proactive reliability work.
  • Define SLOs and establish burn-rate alerting to turn reliability targets into measurable metrics.
  • Lead end-to-end incident investigations and command multi-team responses from declaration to close-out.
  • Influence the technical roadmap and mentor junior SREs on infrastructure reliability practices.

What we're looking for

  • 5+ years of site reliability, infrastructure, or platform engineering experience with direct on-call ownership in production systems.
  • Hands-on orchestration experience with Kubernetes or equivalent including cluster health, resource management, and failure diagnosis at scale.
  • Experience owning or operating a device or VM provisioning pipeline (preferred).
  • Track record of improving system reliability against measurable outcomes like uptime and MTTR (preferred).
  • Incident command discipline to lead multi-team incidents from declaration to close-out (preferred).
  • Depth in distributed systems, device management infrastructure, or ML platform operations (preferred).
  • Demonstrated cross-team technical influence to shape reliability practices beyond the immediate team (preferred).

More like this

Similar roles

Site Reliability Engineering Manager

Apple Inc

Cupertino, CA 45 days ago $267,800$401,700
SRE Distributed Systems Linux Kubernetes AWS GCP AI/ML LLM AIOps Capacity Planning Performance Engineering Infrastructure Architecture Networking Automation
10+ yrs exp

Senior Site Reliability Engineer

Apple Inc

Cupertino, CA 160 days ago $184,700$277,600
SRE Python Go Java Kubernetes Linux Distributed Systems micro-services Automation Networking Security Encryption Monitoring Alerting Capacity Planning Disaster Recovery
5+ yrs exp

Service Reliability Engineer

Apple Inc

Seattle, WA 52 days ago $142,300$263,300
SRE DevOps AWS Kubernetes Java Spark Flink Hadoop Big Data bare-metal Automation Scalability
5+ yrs exp

Site Reliability Engineering Lead

US Bank

Atlanta, GA +2 6 days ago $111,605$131,300
SRE DevOps AWS Azure Kubernetes Docker Terraform Ansible Python PowerShell Shell Scripting CI/CD GitHub Actions Azure DevOps Jenkins GitLab Datadog Splunk Dynatrace Grafana Prometheus CloudWatch Azure Monitor OpenTelemetry ServiceNow Jira REST APIs SQL
6+ yrs exp

Senior Site Reliability Engineer

CVS Health

Woonsocket, RI +1 17 days ago $92,700$203,940
SRE DevOps Kubernetes OpenShift Docker CI/CD Splunk Dynatrace Datadog Prometheus Grafana Python Java AWS Microsoft Azure Google Cloud Rancher GitHub BitBucket Jenkins Microservices web API’s Apigee
5+ yrs exp Hybrid