Site Reliability Engineer, Apple Data Platform & Multi-Cloud Infrastructure

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Austin, TX
Posted
43 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $199k
$138k most similar roles pay here $262k

This listing doesn't post a salary. Most similar roles pay $167,350–$231,150.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Site Reliability Engineer, Apple Data Platform & Multi-Cloud Infrastructure

Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure joins the Apple Data Platform SRE team to manage a massive multi-cloud platform supporting internal engineers building data and AI products. You will operate, monitor, and triage production environments for big data pipelines and machine learning services across AWS, GCP, and on-premise Kubernetes clusters. Key responsibilities include managing infrastructure as code using Terraform and Crossplane, maintaining GitOps workflows with Flux, and providing technical support via Slack. You will debug complex issues involving IAM policies, storage quotas, and cluster disruptions while building automation to reduce manual toil. The role requires proficiency in Python, experience with Golang, and deep knowledge of Kubernetes administration, Prometheus, Grafana, and Splunk. You will ensure the reliability of core services like Spark, Flink, Airflow, and Ray across diverse cloud environments.

What does a Site Reliability Engineer earn?

Median $186200 from 131 postings across 36 companies.

See salary data

What you'll do

  • Operate, monitor, and triage production environments across data processing, ML/AI, and multi-cloud infrastructure.
  • Participate in a rotating on-call schedule to provide coverage for supported services.
  • Serve as a subject matter expert for AWS services, EKS clusters, and cross-cloud networking.
  • Provide Slack-based support to internal customers by screening and resolving service-related issues.
  • Debug production incidents involving IAM permissions, storage limits, and cluster-wide disruptions.
  • Design monitoring, alerting, and dashboards using Prometheus, Grafana, and Splunk for new services.
  • Maintain and evolve Infrastructure-as-Code and GitOps workflows to manage state drift and reconciliation.
  • Build automation and self-healing tooling to reduce manual toil and increase operational capacity.

What we're looking for

  • Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years of experience in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.
  • Proficiency in Python and working knowledge of Golang is preferred.
  • Extensive hands-on AWS experience including IAM, EKS, RDS, S3, VPC networking, autoscaling, and EBS.
  • Experience in Kubernetes administration covering RBAC, scheduling, autoscalers, and troubleshooting cluster disruptions.
  • Solid grounding in SRE principles with prior on-call or production support experience.
  • Strong communication skills and the ability to remain composed under pressure during incidents.
  • Experience with Infrastructure-as-Code (Terraform/Crossplane) and GitOps workflows (Flux).

More like this

Similar roles

Site Reliability Engineer, Apple Data Platform

Apple Inc

Austin, TX 43 days ago
Kubernetes Go Python Java AWS GCP Ali Cloud Linux Flink Hive Hadoop HDFS Trino Druid Containers Virtualization Site Reliability Engineering Distributed Systems
5+ yrs exp

Senior Site Reliability Engineer

Apple Inc

San Francisco, CA 148 days ago $184,700$277,600
SRE Site Reliability Engineering Python Go Java Kubernetes Linux micro-services Distributed Systems Automation Networking Security Encryption Monitoring Alerting Capacity Planning Disaster Recovery
5+ yrs exp