Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Austin, TX
Posted
39 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $209k
$145k most similar roles pay here $276k

This listing doesn't post a salary. Most similar roles pay $177,050–$241,750.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Site Reliability Engineer, Apple Data Platform - AI/ML Platform joins the Apple Data Platform SRE team to manage a multi-cloud infrastructure supporting internal engineers building data and AI products. You will operate, monitor, and triage production environments for big data pipelines and ML/AI services including Spark, Flink, Airflow, Ray, LangGraph agents, RAG architectures, embeddings platforms, and vector store platforms. Responsibilities include providing Slack-based support, designing monitoring and alerting using Prometheus, Grafana, and Splunk, and building automation to reduce manual toil. The role requires proficiency in Python, knowledge of Golang, experience with Kubernetes, and familiarity with AWS or GCP. You will serve as a subject matter expert for ML infrastructure, ensuring the reliability of production systems while collaborating with development teams to onboard new services and maintain high availability across complex data processing environments.

What does a Site Reliability Engineer earn?

Median $186200 from 131 postings across 36 companies.

See salary data

What you'll do

  • Operate, monitor, and triage production environments across data processing, ML/AI, and multi-cloud infrastructure.
  • Participate in a rotating on-call schedule to provide coverage for supported services.
  • Serve as the subject matter expert for ML/AI platform services including Ray, LangGraph, RAG pipelines, and vector stores.
  • Provide Slack-based support to internal customers by screening and resolving service-related issues.
  • Design monitoring, alerting, and dashboards using Prometheus, Grafana, and Splunk for new services.
  • Build automation and self-healing tooling to reduce manual toil and increase operational capacity.
  • Identify, escalate, and resolve production incidents to maintain platform reliability.

What we're looking for

  • Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years of experience in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.
  • Proficiency in Python and working knowledge of Golang is preferred.
  • Experience with Kubernetes and at least one major cloud provider (AWS or GCP).
  • Exposure to operating or supporting ML pipelines, model-serving infrastructure, or LLM-based systems in production.
  • Solid grounding in SRE principles with prior on-call or production-support experience.
  • Strong communication skills and composure under pressure during incidents.
  • Experience with big data technologies (Spark, Flink, Airflow) and observability tools like Prometheus, Grafana, and Splunk.

More like this

Similar roles

Site Reliability Engineer, Apple Data Platform

Apple Inc

Austin, TX 43 days ago
Kubernetes Go Python Java AWS GCP Ali Cloud Linux Flink Hive Hadoop HDFS Trino Druid Containers Virtualization Site Reliability Engineering Distributed Systems
5+ yrs exp

Senior Site Reliability Engineer

Apple Inc

San Francisco, CA 148 days ago $184,700$277,600
SRE Site Reliability Engineering Python Go Java Kubernetes Linux micro-services Distributed Systems Automation Networking Security Encryption Monitoring Alerting Capacity Planning Disaster Recovery
5+ yrs exp