Site Reliability Engineer, AiDP Production Engineering

Apple Inc

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Austin, TX
Posted
3 days ago
Freshness
Confirmed live today

Market check

Salary context

How this pay compares to similar roles

Similar $180k
$133k most similar roles pay here $234k

This listing doesn't post a salary. Most similar roles pay $145,750–$215,000.

Based on 238 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 2321 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1891 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Site Reliability Engineer, AiDP Production Engineering

Site Reliability Engineer, AiDP Production Engineering joins the AI and Data Platform team to manage real-time, near real-time, and batch analytical solutions supporting core business functions like sales, finance, and marketing. This role involves configuring, tuning, and ensuring the resilience of multi-tiered systems across bare-metal and cloud environments. You will build automation for self-healing systems, develop monitoring tools for low-latency applications, and triage incidents to ensure high availability. The work focuses on managing complex data pipelines and infrastructure at scale to provide predictable performance for data analytics. Required technical skills include experience with Apache Spark, Flink, Kafka, Kubernetes, AWS, GCP, Snowflake, Cassandra, SingleStore, and SAP HANA. You will utilize Python or Java to support applications while leveraging tools like Prometheus, Grafana, CloudWatch, and various data visualization platforms to solve critical infrastructure challenges.

What does a Site Reliability Engineer earn?

Median $188000 from 129 postings across 39 companies.

See salary data

What you'll do

  • Assess and select appropriate services and topologies across AWS, Baremetal, and Kubernetes based on performance and security requirements.
  • Build automation to create self-healing systems and improve overall platform reliability.
  • Develop monitoring tools to track high-performance metrics and alert for low-latency applications.
  • Troubleshoot complex network, system, and application-specific issues in large-scale distributed environments.
  • Triage production incidents based on business impact and implement immediate mitigation steps.
  • Conduct root cause analysis (RCA) and partner with engineering teams to prioritize and fix defects.
  • Support Java-based applications and Spark/Flink jobs across multi-cloud and bare-metal infrastructures.
  • Participate in an on-call rotation to provide 24/7 support for mission-critical services.

What we're looking for

  • BS/MS in computer science or equivalent experience.
  • 4+ years of programming experience in Python or Java.
  • 4+ years of experience in cloud-native services, including ETL frameworks like Apache Spark and Flink.
  • 4+ years of experience in messaging systems (Kafka) and cloud infrastructure & services, AWS, GCP, and Kubernetes.
  • 4+ years of experience in modern and distributed databases such as Snowflake, Cassandra, SingleStore, and SAP HANA.
  • Solid understanding of system design, data structures, and incident management best practices (preferred).
  • Experience with observability tools like Prometheus, Grafana, or CloudWatch (preferred).
  • Experience using GenAI or automation tools for issue detection, alerting, or remediation (preferred).

More like this

Similar roles

Site Reliability Engineer

Morgan Stanley

Alpharetta, GA 31 days ago
Python Shell Scripting Perl Ruby Java C# AWS Azure Jenkins Splunk DB2 Oracle Sybase Autosys Linux Unix Windows Agile Scrum Web Services MQ
5+ yrs exp

Site Reliability Engineer

Berkeley Research Group

Remote 86 days ago $130,000$160,000
Azure Kubernetes CI/CD GitHub Actions GitLab CI Golang Ruby Python AWS GCP Infrastructure as Code Datadog OpsGenie PagerDuty SRE Incident Management
5+ yrs exp Remote

Site Reliability Engineer

Balyasny Asset Management

Warsaw, Poland 87 days ago
Prometheus Grafana Loki Tempo OTEL Kubernetes Docker AWS Python Bash Go CI/CD DevOps SRE Agile
5+ yrs exp

Site Reliability Engineer

F5 Inc

San Jose, CA 10 days ago $149,800$224,600
SRE AWS Python Kubernetes Terraform Prometheus Grafana Linux PostgreSQL REST APIs JSON HTTP Ticketing Systems Networking SaaS Automation
1+ yrs exp Hybrid

Site Reliability Engineer

Corepay

Brentwood, TN 10 days ago $157,000$190,000
AWS Azure SRE DevOps CI/CD Infrastructure-as-Code Monitoring Observability Git Generative AI Automation
5+ yrs exp