Staff Machine Learning Platform Engineer, AI Evaluation

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Seattle, WA
Salary
$205,400–$308,500 / yr
Posted
142 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $226k
This role $257k
$168k most similar roles pay here $324k

This role pays more than 82% of similar roles. Most pay $197,631–$254,750 — the shaded band above. At the midpoint, this role pays about $257k versus about $226k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · Staff Machine Learning Platform Engineer, AI Evaluation

Staff Machine Learning Platform Engineer, AI Evaluation joins the Apple Services Engineering team to lead the architectural design and development of high-availability services and internal tools for self-service evaluation at scale. This role involves productionizing machine learning research by transforming complex workflows into developer-first platforms, building APIs, SDKs, and orchestration services. The engineer will manage technical direction, make strategic decisions on infrastructure versus software improvements, and improve the developer experience for evaluating generative AI and agent systems. Key responsibilities include defining operational standards for testing, CI/CD, and monitoring. Required skills include extensive software engineering experience, Python proficiency with FastAPI and Pydantic, and familiarity with Ray, Docker, Kubernetes, and distributed compute frameworks like Dask. The role addresses the technical challenges of non-deterministic outputs, multi-step agent reasoning, and scoring drift in large-scale AI systems.

What you'll do

  • Design and build APIs, SDKs, and orchestration services to turn research methodologies into self-service building blocks.
  • Partner with research engineers to productionize code by creating reusable Python services and infrastructure improvements.
  • Make high-level architectural decisions to distinguish scalable platform features from one-off requests.
  • Define the technical roadmap and strategy for org-wide AI evaluation systems.
  • Improve developer experience by supporting complex evaluation patterns like multi-turn agent trajectories and tool-use chains.
  • Establish standards for testing, CI/CD, monitoring, and reliability across the evaluation platform.
  • Manage infrastructure requirements such as GPU compute, distributed scheduling, and job orchestration.

What we're looking for

  • Must have 8+ years of software engineering experience with a track record of owning platform-level technical direction.
  • Proven ability to build "0-to-1" products and design for scale while making deliberate architectural trade-offs.
  • Deep understanding of machine learning research code and the ability to distinguish between software and infrastructure problems.
  • Experience in AI/Agent evaluation, including handling non-deterministic outputs, multi-step reasoning, and judge model reliability.
  • Proficiency in Python with experience in FastAPI, Pydantic, and job orchestration frameworks like Temporal.io.
  • Experience with operational requirements including CI/CD, containerization (Docker/K8s), and monitoring for production services.
  • Strong communication skills to write design docs, decision records, and platform roadmaps for diverse stakeholders.
  • Preferred experience with distributed compute frameworks (Ray, Dask) and LLM token economics or cost management.

More like this

Similar roles

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 65 days ago $184,700$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Machine Learning Evaluation Engineer

Apple Inc

Sunnyvale, CA 13 days ago $150,400$277,600
Machine Learning Computer Vision Python Data Analysis Statistical Modeling Deep Learning Data Pipelines Visualization Synthetic Data Data Augmentation Model Monitoring Generative AI Root-Cause Analysis Evaluation Frameworks
3+ yrs exp

Senior Staff Machine Learning Engineer, ML Platform

Apple Inc

New York, NY 107 days ago $216,200$394,000
Machine Learning Deep Learning Transformers LLMs TensorFlow PyTorch Distributed Training Model Pruning Quantization Distillation Agentic AI Federated Learning Differential Privacy SFT Agile
10+ yrs exp

Staff Machine Learning Engineer, AI Agent Platform

GEICO

Remote (New York, NY) +3 98 days ago $115,000$260,000
Python Java Go Kubernetes AKS FastAPI Docker Prometheus OpenTelemetry PostgreSQL Redis Neo4j OpenSearch Temporal TensorFlow PyTorch LangGraph CrewAI AutoGen RAG LLM MCP A2A
6+ yrs exp Remote

Machine Learning Platform Engineer

Apple Inc

Seattle, WA 115 days ago $175,000$263,300
Python FastAPI Pydantic LLM Docker Kubernetes CI/CD Ray Airflow Temporal.io MLflow Weights & Biases Promptfoo LangSmith Braintrust Inspect
4+ yrs exp

Staff Software Engineer, Machine Learning Platform

Stripe

Seattle, WA +1 101 days ago $203,600$305,400
MLOps Machine Learning LLM Distributed Systems AWS SageMaker Bedrock Databricks OpenAI Model Serving Feature Stores Retrieval-Augmented Generation System Architecture
10+ yrs exp