Manager, AI Evaluation Engineering

Caterpillar

Confirmed live 3 days ago High trust

Quick summary

Work type
On-site
Location
Broomfield, COChicago, ILPeoria, ILIrving, TX
Salary
$147,760–$240,110 / yr
Posted
10 days ago
Freshness
Confirmed live 3 days ago
Closes
Sep 20, 2026

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $222k
This role $194k
$134k most similar roles pay here $277k

This role pays less than 66% of similar roles. Most pay $183,675–$260,250 — the shaded band above. At the midpoint, this role pays about $194k versus about $222k for comparable roles.

Based on 240 similar postings.

Employer

About Caterpillar

Caterpillar Inc. is the world''s largest manufacturer of construction and mining equipment, diesel and natural gas engines, industrial gas turbines, and diesel-electric locomotives. Industry: Heavy Equipment & Manufacturing

Caterpillar currently has 62 open roles on FindRole.

Listed pay typically runs $131,789–$192,710 across 61 roles with salary data.

Most-posted roles

View all roles at Caterpillar

At a glance

TL;DR · Manager, AI Evaluation Engineering

Manager; AI Evaluation Engineering joins the Cat Digital AI Engineering team to lead a group dedicated to evaluating and validating advanced generative AI solutions, including intelligent agents and digital assistants. The manager provides technical direction, oversees team performance, and ensures the quality of robust, reliable AI products while implementing engineering best practices. This role focuses on solving challenges related to non-deterministic systems through metric design, RAG quality assessment, safety evaluation, and human-in-the-loop validation. Candidates must possess expertise in LLMs, SLMs, prompt engineering, agentic systems, and RAG architectures. Required technical skills include experience with Azure, AWS, GCP, LangChain, LangGraph, and vector databases. The role involves managing complex software development lifecycles, CI/CD pipelines, and advanced evaluation ecosystems to ensure high-quality outputs across hybrid cloud and embedded environments for innovative AI-driven customer experiences.

What you'll do

  • Lead a team dedicated to evaluating and validating advanced generative AI solutions like intelligent agents and digital assistants.
  • Provide technical support and direction to align the team with corporate goals for AI project evaluation.
  • Oversee individual and team performance while identifying and addressing specific training and development needs.
  • Ensure all AI engineering products are robust, reliable, and meet high quality standards.
  • Establish and supervise the implementation of engineering best practices across development processes.
  • Scale processes, tools, metrics, and engineering practices across multiple products and enterprise initiatives.
  • Manage the deployment of GenAI solutions across hybrid cloud and embedded environments using CI/CD and automated testing.
  • Implement AI evaluation methodologies to manage non-deterministic systems, including RAG quality assessment and safety evaluations.

What we're looking for

  • Proven experience leading software quality, test automation, validation, and AI evaluation teams.
  • Experience delivering enterprise-scale software and GenAI solutions across hybrid cloud and embedded/edge environments.
  • Deep expertise in GenAI architecture including LLMs, SLMs, prompt engineering, agentic systems, RAG architectures, and fine-tuning techniques like LoRA.
  • Hands-on experience with AI development ecosystems such as Azure, AWS, GCP, LangChain, LangGraph, and various evaluation tools.
  • Deep understanding of AI evaluation methodologies for non-deterministic systems, including metric design, safety evaluation, and human-in-the-loop validation.
  • Expertise in modern AI observability and monitoring practices including tracing, telemetry, and automated test generation.
  • Knowledge of software development life cycle (SDLC) and engineering best practices like CI/CD and progressive deployment strategies.
  • Knowledge of software quality assurance and testing processes to ensure high-quality computer software products.

More like this

Similar roles

Manager, AI Evaluation Engineering

Caterpillar

Chicago, IL +3 10 days ago $147,760$240,110
Generative AI LLMs SLMs RAG LangChain LangGraph Prompt Engineering Vector Databases Azure AWS GCP SageMaker Bedrock Snowflake Cortex CI/CD Automated Testing LoRA Feature Flags

Manager, AI Engineering

Caterpillar

Chicago, IL +2 10 days ago $147,760$240,110
Generative AI LLMs RAG LangChain LangGraph Semantic Kernel CrewAI Python Go Java AWS Azure Google Cloud CI/CD Infrastructure-as-Code Agile Azure DevOps Prompt Engineering

Manager, AI Engineering

Caterpillar

Chicago, IL +2 10 days ago $147,760$240,110
Generative AI LLMs RAG LangChain LangGraph Semantic Kernel CrewAI Python Go Java AWS Azure Google Cloud CI/CD infrastructure-as-code Agile Azure DevOps Prompt Engineering

Manager, AI Engineering

Mastercard

O Fallon, MO 8 days ago $140,000$231,000
Generative AI LLM RAG Python SQL Databricks AWS LangChain LangGraph Hugging Face MLflow SageMaker Grafana Datadog CloudWatch CI/CD RAGAS TruLens DeepEval PromptFlow

AI Engineer

Global Payments (TSYS)

Alpharetta, GA +1 59 days ago
Generative AI Machine Learning Deep Learning Python LangChain LangGraph AgentSpace GCP Vertex AI AWS Bedrock SageMaker Snowflake Cortex PGVector Transformers PyTorch TensorFlow RAG RLHF Prompt Engineering MLOps CI/CD Apache Spark Kafka BigQuery Snowflake
4+ yrs exp

Manager, AI Engineering

Rockwell Automation

Milwaukee, WI 51 days ago
Python Azure AI GenAI LLM RAG Vector Databases MLOps LLMOps CI/CD Microsoft 365 SharePoint Graph API Power Platform GitHub Actions API Data Engineering
10+ yrs exp Hybrid