Site Reliability Engineer, Apple Data Platform / Big Data Platform

Apple Inc

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Austin, TX
Posted
43 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

How this pay compares to similar roles

Similar $202k
$138k most similar roles pay here $266k

This listing doesn't post a salary. Most similar roles pay $173,175–$231,300.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Site Reliability Engineer, Apple Data Platform / Big Data Platform

Site Reliability Engineer, Apple Data Platform / Big Data Platform joins the Apple Data Platform SRE team to manage a multi-cloud infrastructure supporting internal engineers building data and AI products. You will operate, monitor, and triage production environments while serving as a subject matter expert for big data services including Spark, Flink, Airflow, Trino, Notebooks, REST Catalog services, and data governance tools. Day-to-day responsibilities involve managing incident response, providing technical support to internal customers via Slack, and building automation and self-healing tooling to reduce manual toil. The role requires proficiency in Python, knowledge of Golang, and experience with Kubernetes and major cloud providers like AWS or GCP. You will also utilize observability tools such as Prometheus, Grafana, and Splunk to design monitoring and alerting for complex data processing and ML/AI platform services.

What does a Site Reliability Engineer earn?

Median $186200 from 131 postings across 36 companies.

See salary data

What you'll do

  • Operate, monitor, and triage production environments across data processing, ML/AI, and multi-cloud infrastructure.
  • Serve as a subject matter expert for big data services including Spark, Flink, Airflow, Trino, and Notebooks.
  • Act as the primary point of contact for internal customers to diagnose and resolve service issues via Slack.
  • Manage and resolve customer-reported support tickets by prioritizing based on impact and urgency.
  • Design monitoring, alerting, and dashboards using Prometheus, Grafana, and Splunk for newly onboarded services.
  • Build automation and self-healing tooling to reduce manual toil and increase operational capacity.
  • Identify and escalate production issues to maintain platform reliability and improve customer experience.

What we're looking for

  • Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years of experience in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.
  • Proficiency in Python and working knowledge of Golang is preferred.
  • Deep understanding of one or more Big Data technologies including Spark, Flink, Airflow, Trino, or Notebooks.
  • Experience with Kubernetes and at least one major cloud provider such as AWS or GCP.
  • Solid grounding in SRE principles with experience in on-call, production support, or customer-facing roles.
  • Excellent written and verbal communication skills to explain technical issues to non-expert customers.
  • Familiarity with observability tools like Prometheus, Grafana, Splunk, and PagerDuty.

More like this

Similar roles

Site Reliability Engineer, Apple Data Platform

Apple Inc

Austin, TX 43 days ago
Kubernetes Go Python Java AWS GCP Ali Cloud Linux Flink Hive Hadoop HDFS Trino Druid Containers Virtualization Site Reliability Engineering Distributed Systems
5+ yrs exp

Senior Site Reliability Engineer

Apple Inc

San Francisco, CA 148 days ago $184,700$277,600
SRE Site Reliability Engineering Python Go Java Kubernetes Linux micro-services Distributed Systems Automation Networking Security Encryption Monitoring Alerting Capacity Planning Disaster Recovery
5+ yrs exp