Site Reliability Engineer - Kafka

Seattle, Washington, USA Posted 20 days ago

$139,500 - $258,100/year

Role Details

The Data Service SRE team develops applications and tooling that are safe, reliable, scalable, and fast. This work requires an innovative spirit and an extraordinary degree of care and difficulty in engineering. Team members contribute to all major components of Kafka deployment infrastructure, including maintenance automation, control plane enhancements, monitoring and alerting tooling/dashboards, advanced deployment architecture, focused on safety, stability, performance, and scaling. Come join us at Apple Services Engineering and help us deliver services and applications that are fluid and responsive. You will collaborate with engineers from across Apple to define the metrics, set targets, uncover optimization opportunities, and ship a service that will delight our customers. This role is for engineers who enjoy deep technical engineering that spans large cross-organizational projects. Your openness to learning and implementing new technologies will contribute to the continuous evolution of our organization. Good ideas are valued and rewarded. Understanding of core SRE concepts - Monitoring, Alerting, Incident management Deep and wide performance engineering (design concepts, profile-guided optimization) Service lifecycle mangement across bare metal, and virtualized (EC2), kubernetes platforms Prepare alert handling procedures, run-books, and collaborate with other SRE team members. Excellent communication and a high degree of customer focus when engaging with internal platform customers As a distributed team, ability to work optimally with colleagues based in other locations is essential Prior experience with development or maintenance of Kafka infrastructure or similar data service is highly recommended 5 or more years of experience in support of internet-facing production services and distributed systems via deployments, On Call and Incident Management. 5 or more years of experience running large scale infrastructure with a heavy reliance on automation tooling 5 or more years of experience troubleshooting and performance deep dive analysis Real operational experience managing services at scale on Kubernetes Proficient in one or more of the following programming languages: Java, Go (golang), Python Operational experience deploying in and running on Datacenter and Cloud architectures (networking topologies, host placement strategies, and failure modes); design of multi-datacenter systems; failure domains; and wide-area networking. Self motivated, inquisitive with an aptitude to learn new technologies quickly and effectively. Demonstrated expertise developing and troubleshooting distributed systems and database storage engines. Experience developing critical internet services and/or platform infrastructure. Experience with AWS, GCP and IaC such as Terraform Experience managing messaging services such as Kafka or other Data services Proficient in Java, Go (golang) & Python

For more details click Job Post.

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

View All Jobs →