Technical Operations & Site Reliability Engineer, Customer Systems

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Sunnyvale, CA
Salary
$150,400–$277,600 / yr
Posted
148 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $180k
This role $214k
$126k most similar roles pay here $294k

This role pays more than 78% of similar roles. Most pay $146,633–$213,692 — the shaded band above. At the midpoint, this role pays about $214k versus about $180k for comparable roles.

Based on 238 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Technical Operations & Site Reliability Engineer, Customer Systems

The Technical Operations & Site Reliability Engineer, Customer Systems joins the Customer Systems Operations team to maintain the reliability, availability, and performance of business-critical, globally distributed systems. This role involves managing large-scale production outages, leading incident responses, and designing automation solutions to streamline monitoring and operational workflows. The engineer will develop tools and software to reduce manual intervention and improve system stability while collaborating with multi-functional teams to drive operational metrics. Key technologies include Java/JEE, REST, Swift/Objective C, Python, Go, and Bash, alongside experience with AI and LLM models for operational excellence. Candidates must possess expertise in networking protocols like HTTP, DNS, and TCP/IP, as well as monitoring tools such as Hubble, ExtraHop, and Splunk. The role focuses on solving complex reliability challenges within large-scale mission-critical applications in a 24x7 environment.

What you'll do

  • Manage large-scale production outages and lead incident response efforts.
  • Design, build, and maintain automation solutions to streamline the management of distributed systems.
  • Develop software tools using languages like Java, Python, or Go to automate repetitive operational tasks.
  • Integrate AI and LLM models to improve application support and operational efficiency.
  • Monitor system health and manage incident communication across critical global applications.
  • Track and align operational metrics and KPIs for business-critical systems.
  • Create technical documentation regarding architecture, infrastructure configuration, and standard operating procedures.
  • Train users on complex topics and provide training materials for internal teams.

What we're looking for

  • Bachelor of Science in Computer Science, Computer Engineering, or equivalent related work experience.
  • Experience interpreting operational data from monitoring tools like Hubble, ExtraHop, and Splunk.
  • Proficiency in scripting and automation using Java, JEE, REST, Swift/Objective C, Python, Go, or Bash.
  • Experience using AI and Large Language Models (LLMs) to enhance operational efficiency and model optimization.
  • Understanding of standard networking protocols including HTTP, DNS, TCP/IP, ICMP, the OSI Model, Subnetting, and Load Balancing.
  • Knowledge of Linux Operating Systems, including Kernel, Memory, Process, Threads, and IPC.
  • Understanding of distributed systems concepts such as Microservices, Messaging Brokers, and Versioning.
  • Experience managing large-scale mission-critical applications in a 24x7 global environment.

More like this

Similar roles

Site Reliability Engineer, Customer Systems

Apple Inc

Sunnyvale, CA 115 days ago $150,400$225,300
Kubernetes Helm Python Shell Scripting Ansible Splunk Grafana Prometheus Alertmanager CI/CD ArgoCD GitOps Java DNS TCP HTTP/HTTPS Infrastructure as Code GenAI

Site Reliability Engineer, Customer Systems

Apple Inc

Sunnyvale, CA 98 days ago $150,400$225,300
Kubernetes Helm Python Shell Scripting Ansible Splunk Grafana Prometheus Alertmanager CI/CD ArgoCD GitOps Java DNS TCP HTTP HTTPS GenAI

Site Reliability Engineer, Enterprise Technology Services

Apple Inc

Sunnyvale, CA 28 days ago $150,400$277,600
SRE DevOps Java Python Bash LUA Oracle MongoDB Prometheus Splunk Grafana CloudWatch Linux Networking TLS/SSL DNS Load Balancers Git CI/CD Kubernetes AWS GCP Nginx Envoy NetScaler
5+ yrs exp