AI Evaluations Engineer, Decision Intelligence

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Cupertino, CA
Salary
$184,700–$277,600 / yr
Posted
2 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $206k
This role $231k
$135k most similar roles pay here $293k

This role pays more than 65% of similar roles. Most pay $162,000–$250,362 — the shaded band above. At the midpoint, this role pays about $231k versus about $206k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 2321 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1891 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · AI Evaluations Engineer, Decision Intelligence

The AI Evaluations Engineer, US Decision Intelligence joins the US Decision Intelligence team to own the end-to-end evaluation pipeline for AI products and agentic workflows. This role involves architecting comprehensive frameworks to trace agent responses, tool calling, and skill levels while implementing rubric-based evaluations for correctness, relevance, and grounding. The engineer will build and operate workflows to measure LLM outputs across chat, summarization, and recommendations, while managing the platform-wide evaluation gate to ensure quality standards are met before release. Key technical requirements include proficiency in Python and SQL, experience with LLM ecosystems like OpenAI and Anthropic, RAG pipelines, vector databases such as Pinecone or Milvus, and observability tools like Langfuse. The role addresses the challenge of improving accuracy and performance for AI systems that reduce time to insight and catalyze decision making.

What you'll do

  • Architect and maintain comprehensive evaluation frameworks to measure accuracy, performance, and latency across AI systems.
  • Build and operate workflows to evaluate LLM outputs for chat, summarization, recommendations, and agentic actions.
  • Implement rubric-based evaluations to score outputs for correctness, relevance, grounding, and consistency.
  • Develop advanced "agent-as-a-judge" patterns and harness-based evaluation techniques in sandbox environments.
  • Instrument workflows to capture traces, prompts, responses, metadata, and user feedback for observability.
  • Act as the gatekeeper for product releases by defining and enforcing pass/fail criteria for all tracks.
  • Define specific metrics for agent performance, including task completion, tool correctness, and error recovery.
  • Translate complex product requirements into measurable evaluation criteria and technical rubrics.

What we're looking for

  • 5+ years of experience in data and AI-related fields such as AI engineering, software development, ML engineering, data science, or QA roles.
  • B.S. degree in Computer Science/Engineering, or equivalent work experience.
  • Strong Python skills.
  • Hands-on experience with AI evaluation techniques, including Golden datasets, LLM-as-a-Judge, and rubric-based scoring.
  • Experience with LLM ecosystems (OpenAI, Anthropic, Gemini), RAG pipelines, and vector databases like Pinecone or FAISS.
  • Proficiency in SQL and experience with at least one major data analytics platform such as Hadoop, Spark, or Snowflake.
  • Experience with CI/CD or release validation workflows.
  • Hands-on experience with Langfuse or similar tools for LLM observability.
  • Familiarity with embeddings, retrieval algorithms, agents, and data modeling for vector and graph databases (preferred).
  • Advanced degree in Economics, Electrical Engineering, Statistics, Data Science, or a similar quantitative field (preferred).

More like this

Similar roles

Evaluation & Insights Machine Learning Engineer

Apple Inc

Cupertino, CA 74 days ago $184,700$324,800
Python PyTorch JAX Hugging Face LLMs RAG MLOps CI/CD vLLM Ray Fine-Tuning Prompt Engineering Vector Databases RLHF DPO MLflow Weights & Biases NLP Embedding-based Clustering
8+ yrs exp

Software Engineer, Evaluation

Apple Inc

Cupertino, CA 68 days ago $150,400$277,600
AI Machine Learning Distributed Systems React JavaScript RESTful API SQL NoSQL Benchmarking Software Engineering
3+ yrs exp

Lead Forward Deployed Engineer, AI Evaluation Platform

Apple Inc

Seattle, WA 122 days ago $175,000$263,300
AI Evaluation LLMs Agentic Systems Machine Learning DeepEval Ragas TruLens LangSmith SDKs ML Infrastructure Product Management Technical Program Management
5+ yrs exp Hybrid

AI Engineer, Algorithm Evaluation & Agentic Systems

Apple Inc

Sunnyvale, CA 38 days ago $150,400$277,600
Python PyTorch Computer Vision Machine Learning LLM VLM Vision Transformers Agentic Systems Statistics Evaluation Frameworks Benchmarking Reinforcement Learning Regression Testing Embeddings AI Safety
3+ yrs exp