Site Reliability Engineer, AiDP Production Engineering

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Austin, TX
Posted
48 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $181k
$133k most similar roles pay here $235k

This listing doesn't post a salary. Most similar roles pay $147,075–$214,000.

Based on 238 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Site Reliability Engineer, AiDP Production Engineering

Site Reliability Engineer, AiDP Production Engineering joins the Production Engineering team within the AI and Data Platform organization to manage real-time, near real-time, and batch analytical solutions. This role involves configuring, tuning, and ensuring the resilience of multi-tiered systems across bare-metal and cloud environments to support critical business functions like sales, finance, and marketing. You will build automation for self-healing systems, develop monitoring tools for low-latency applications, triage incidents, and perform root cause analysis. The work involves managing data pipelines and infrastructure using technologies including Kafka, Spark, Flink, Iceberg, Airflow, Kubernetes, AWS, GCP, Snowflake, Cassandra, SingleStore, and SAP HANA. Candidates should possess proficiency in Python or Java and experience with observability tools like Prometheus and Grafana to maintain high-performance systems while solving complex infrastructure challenges at scale.

What does a Site Reliability Engineer earn?

Median $186200 from 131 postings across 36 companies.

See salary data

What you'll do

  • Assess and select appropriate services and topologies across AWS, Baremetal, and Kubernetes based on performance and security requirements.
  • Build automation to create self-healing systems and improve overall infrastructure reliability.
  • Develop tools to monitor high-performance applications and alert on low-latency issues.
  • Troubleshoot complex network, system, and application-specific performance issues in distributed environments.
  • Triage production incidents based on business impact and implement immediate mitigation steps.
  • Conduct root cause analysis (RCA) and partner with engineering teams to prioritize and fix defects.
  • Support Java-based applications and Spark/Flink jobs across multi-cloud and bare-metal infrastructures.
  • Participate in an on-call rotation to provide 24/7 support for mission-critical services.

What we're looking for

  • BS/MS in computer science or equivalent experience.
  • 4+ years of programming experience in Python or Java.
  • 4+ years of experience in cloud-native services, including ETL frameworks like Apache Spark and Flink.
  • 4+ years of experience in messaging systems (Kafka) and cloud infrastructure & services, AWS, GCP, and Kubernetes.
  • 4+ years of experience in modern and distributed databases such as Snowflake, Cassandra, SingleStore, and SAP HANA.
  • Ability to troubleshoot complex production issues and perform root cause analysis for large-scale distributed systems.
  • Experience with observability tools like Prometheus, Grafana, or CloudWatch.
  • Experience using GenAI or automation tools for issue detection, alerting, or remediation.

More like this

Similar roles

Site Reliability Engineer

Morgan Stanley

Alpharetta, GA 22 days ago
Python Shell Scripting Perl Ruby Java C# AWS Azure Jenkins Splunk DB2 Oracle Sybase Autosys Linux Unix Windows Agile Scrum Web Services MQ
5+ yrs exp

Site Reliability Engineer

Berkeley Research Group

Remote 77 days ago $130,000$160,000
Azure Kubernetes CI/CD GitHub Actions GitLab CI Golang Ruby Python AWS GCP Infrastructure as Code Datadog OpsGenie PagerDuty SRE Incident Management
5+ yrs exp Remote

Site Reliability Engineer

Balyasny Asset Management

Warsaw, Poland 78 days ago
Prometheus Grafana Loki Tempo OTEL Kubernetes Docker AWS Python Bash Go CI/CD DevOps SRE Agile
5+ yrs exp

Site Reliability Engineer

Booz Allen Hamilton

McLean, VA 9 days ago $86,800$198,000
AWS Kubernetes Terraform Ansible CI/CD Python Bash PowerShell Docker Prometheus Grafana Loki Elasticsearch Kibana GitLab GitHub CloudFormation OpenShift Jenkins REST JSON YAML XML Agile
6+ yrs exp

Site Reliability Engineer

Corepay

Brentwood, TN 2 days ago $157,000$190,000
AWS Azure Generative AI Git Agile Scrum Cloud Architecture Quality Engineering Source Control
6+ yrs exp

Site Reliability Engineer

Cisco

Raleigh, NC 3 days ago $128,600$184,900
Splunk ServiceNow Python Ansible Terraform Kubernetes AWS Linux VMware Cisco UCS HyperFlex ITIL Virtualization
4+ yrs exp Hybrid