Manager, AI Evaluation Engineering

Caterpillar

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Chicago, ILPeoria, ILWestminster, COIrving, TX
Salary
$147,760–$240,110 / yr
Posted
10 days ago
Freshness
Confirmed live yesterday
Closes
Sep 21, 2026

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $222k
This role $194k
$134k most similar roles pay here $277k

This role pays less than 66% of similar roles. Most pay $183,675–$260,000 — the shaded band above. At the midpoint, this role pays about $194k versus about $222k for comparable roles.

Based on 240 similar postings.

Employer

About Caterpillar

Caterpillar Inc. is the world''s largest manufacturer of construction and mining equipment, diesel and natural gas engines, industrial gas turbines, and diesel-electric locomotives. Industry: Heavy Equipment & Manufacturing

Caterpillar currently has 62 open roles on FindRole.

Listed pay typically runs $131,789–$192,710 across 61 roles with salary data.

Most-posted roles

View all roles at Caterpillar

At a glance

TL;DR · Manager, AI Evaluation Engineering

Manager; AI Evaluation Engineering joins the Cat Digital AI Engineering team to lead a group dedicated to evaluating and validating advanced generative AI solutions, including intelligent agents and digital assistants. The role involves providing technical direction, overseeing team performance, ensuring product quality, and implementing engineering best practices. You will manage the evaluation of non-deterministic systems by designing metrics for RAG quality, output reliability, and safety while utilizing tools like LangChain, LangGraph, Langfuse, Arize, and SageMaker. Key responsibilities include managing software development lifecycles, test automation, and CI/CD pipelines across hybrid cloud and edge environments. The work focuses on solving complex challenges in GenAI architecture, including LLMs, prompt engineering, vector databases, and fine-tuning techniques like LoRA to ensure robust, reliable AI-driven experiences for customers and dealers.

What you'll do

  • Lead a team dedicated to evaluating and validating advanced generative AI solutions including intelligent agents and digital assistants.
  • Provide technical support and direction to ensure the team aligns with company goals for AI projects.
  • Oversee individual and team performance while identifying and addressing training and development needs.
  • Ensure all AI engineering products are robust, reliable, and meet high quality standards.
  • Establish and supervise the implementation of engineering best practices across development processes.
  • Manage the evaluation of non-deterministic systems using metric design, RAG quality assessment, and safety evaluations.
  • Implement modern AI observability and monitoring practices including tracing, telemetry, and automated test generation.
  • Scale software quality, test automation, and validation processes across multiple products and enterprise initiatives.

What we're looking for

  • Proven experience leading software quality, test automation, validation, and AI evaluation teams.
  • Experience delivering enterprise-scale software and GenAI solutions across hybrid cloud and embedded/edge environments.
  • Expertise in modern software engineering practices including CI/CD, automated testing, incident management, and progressive deployment strategies.
  • Deep expertise in GenAI architecture, including LLMs, SLMs, prompt engineering, agentic systems, RAG architectures, and fine-tuning techniques like LoRA.
  • Hands-on experience with AI development ecosystems such as Azure, AWS, GCP, LangChain, LangGraph, and various evaluation tools.
  • Deep understanding of AI evaluation methodologies for non-deterministic systems, including metric design, safety evaluation, and human-in-the-loop validation.
  • Expertise in AI observability and monitoring practices, including tracing, telemetry, and automated test generation.
  • Knowledge of software development life cycle (SDLC) and software quality assurance processes.

More like this

Similar roles

Manager, AI Evaluation Engineering

Caterpillar

Broomfield, CO +3 10 days ago $147,760$240,110
Generative AI LLMs SLMs RAG Prompt Engineering LangChain LangGraph Vector Databases Azure AWS GCP SageMaker Bedrock Snowflake Cortex CI/CD Automated Testing LoRA A/B Testing Feature Flags

Manager, AI Engineering

Caterpillar

Chicago, IL +2 10 days ago $147,760$240,110
Generative AI LLMs RAG LangChain LangGraph Semantic Kernel CrewAI Python Go Java AWS Azure Google Cloud CI/CD infrastructure-as-code Agile Azure DevOps Prompt Engineering

Manager, AI Engineering

Caterpillar

Chicago, IL +2 10 days ago $147,760$240,110
Generative AI LLMs RAG LangChain LangGraph Semantic Kernel CrewAI Python Go Java AWS Azure Google Cloud CI/CD Infrastructure-as-Code Agile Azure DevOps Prompt Engineering

Manager, AI Engineering

Mastercard

O Fallon, MO 8 days ago $140,000$231,000
Generative AI LLM RAG Python SQL Databricks AWS LangChain LangGraph Hugging Face MLflow SageMaker Grafana Datadog CloudWatch CI/CD RAGAS TruLens DeepEval PromptFlow

Manager, AI Engineering

Rockwell Automation

Milwaukee, WI 51 days ago
Python Azure AI GenAI LLM RAG Vector Databases MLOps LLMOps CI/CD Microsoft 365 SharePoint Graph API Power Platform GitHub Actions API Data Engineering
10+ yrs exp Hybrid