Site Reliability Engineer - ML

Apple Inc

Confirmed live today High trust

Quick summary

Work type
On-site
Location
New York, NY
Salary
$150,400–$225,300 / yr
Posted
3 days ago
Freshness
Confirmed live today

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $214k
This role $188k
$137k most similar roles pay here $277k

This role pays less than 65% of similar roles. Most pay $174,437–$253,812 — the shaded band above. At the midpoint, this role pays about $188k versus about $214k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 2000 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1613 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Site Reliability Engineer - ML

Site Reliability Engineer - ML, Apple Ads joins the Site Reliability Engineering team to ensure the reliability, performance, and availability of machine learning platform services at scale. This role focuses on building and evolving the next generation of machine learning infrastructure for transactional and analytical workloads within AWS-based environments. You will build automation to eliminate manual processes, develop internal tooling for cost-efficiency, and manage Infrastructure as Code using Terraform. The work involves operating distributed systems utilizing EKS, ElasticCache, Ray over Kubernetes, and NVIDIA Triton Inference Server. Candidates must possess strong programming skills in Python, Java, Rust, or Go, along with deep knowledge of Linux internals. You will solve complex problems regarding ML training, inference, and serving workloads while managing infrastructure for products like the App Store and Apple News.

What does a Site Reliability Engineer earn?

Median $187850 from 131 postings across 38 companies.

See salary data

What you'll do

  • Build and operate distributed systems using AWS managed services, Kubernetes, and ML technologies like Ray and Triton.
  • Develop internal tooling and automation frameworks to improve infrastructure reliability and cost-efficiency.
  • Design and manage Infrastructure as Code using Terraform for secure and scalable deployments.
  • Manage the health, performance, and scalability of large-scale infrastructure for ML training and inference workloads.
  • Lead incident response efforts and conduct postmortems to drive continuous improvement and reduce risk.
  • Troubleshoot complex issues within distributed systems under real-world load.
  • Create automation that eliminates manual processes and improves operational visibility across the platform.

What we're looking for

  • 3+ years of experience in internet-facing backend production systems, SRE, or ML Operations roles on large scale distributed cloud infrastructure.
  • Proven expertise with AWS-managed infrastructure.
  • Familiarity with the ML lifecycle and technologies such as NVIDIA Triton, AnyScale Ray, and Apache Airflow.
  • Strong programming skills in Python, Java, Rust, Go, or similar languages.
  • Hands-on experience with Linux systems and deep knowledge of its internals.
  • Demonstrated experience with Infrastructure as Code, specifically using Terraform.
  • Strong foundation in SRE concepts including monitoring, alerting, observability, incident response, and SLAs/SLOs.
  • Experience with Kubernetes at scale, GPU hardware architectures, or high-performance networking (preferred).

More like this

Similar roles

Site Reliability Engineer

Apple Inc

Cupertino, CA 3 days ago $150,400$225,300
AWS Kubernetes Terraform Python Go Java EKS MSK ElastiCache GitOps Argo CD Flux Helm Crossplane Linux Infrastructure as Code SRE Monitoring Observability
3+ yrs exp

Reliability Engineer - Data

Apple Inc

Austin, TX 15 days ago
AWS Kubernetes Apache Spark Flink Kafka Iceberg EMR EKS MSK Python Java Scala Kotlin Helm Infrastructure as Code GenAI Machine Learning Distributed Systems
2+ yrs exp

Reliability Engineer - Data

Apple Inc

Austin, TX 12 days ago
AWS Kubernetes Apache Spark Flink Kafka Iceberg EMR EKS MSK Python Java Scala Kotlin Helm Infrastructure as Code GenAI Machine Learning Distributed Systems
2+ yrs exp

Site Reliability Engineer, Apple Data Platform

Apple Inc

Austin, TX 45 days ago
Kubernetes Go Python Java AWS GCP Ali Cloud Linux Flink Hive Hadoop HDFS Trino Druid Containers Virtualization Site Reliability Engineering Distributed Systems
5+ yrs exp