Lead Data Engineer, Physical AI Platform

Caterpillar

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Irving, TX
Salary
$128,470–$208,770 / yr
Posted
8 days ago
Freshness
Confirmed live today
Closes
Oct 29, 2026

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $175k
This role $169k
$118k most similar roles pay here $224k

This role pays less than 52% of similar roles. Most pay $135,675–$214,000 — the shaded band above. At the midpoint, this role pays about $169k versus about $175k for comparable roles.

Based on 240 similar postings.

Employer

About Caterpillar

Caterpillar Inc. is the world''s largest manufacturer of construction and mining equipment, diesel and natural gas engines, industrial gas turbines, and diesel-electric locomotives. Industry: Heavy Equipment & Manufacturing

Caterpillar currently has 76 open roles on FindRole.

Listed pay typically runs $128,470–$190,176 across 74 roles with salary data.

Most-posted roles

View all roles at Caterpillar

At a glance

TL;DR · Lead Data Engineer, Physical AI Platform

Lead Data Engineer – Physical AI Platform As a Lead Data Engineer, you will join the team to design, build, and maintain scalable data pipelines, microservices, and cloud-based platforms that provide high-quality data for business and engineering teams. You will collaborate with architects to define solution designs, optimize Python-based pipelines for real-time and batch processing, and develop cloud-native ingestion solutions using AWS services like Kinesis, S3, DynamoDB, and EventBridge. Your role involves establishing automated testing, ensuring data quality across distributed systems, and performing root-cause analysis using tools like CloudWatch. You will work within an agile environment to solve complex problems in construction autonomy by integrating massive volumes of real-world data from connected assets. The core technical stack includes Python, Java, SQL, and CI/CD tools such as Azure DevOps, Jira, and Jenkins to support large-scale infrastructure.

What you'll do

  • Design and optimize scalable data pipelines and microservices using Python for real-time and batch processing.
  • Develop cloud-native data ingestion and streaming solutions using AWS services like Kinesis, S3, and DynamoDB.
  • Build and maintain data integration frameworks to support CI Autonomy initiatives.
  • Translate complex business requirements into technical architectures, workflows, and system designs.
  • Establish automated testing, data quality controls, and validation frameworks for distributed data ecosystems.
  • Perform performance tuning and root-cause analysis on production platforms using tools like CloudWatch.
  • Lead the design of event-driven data systems to ensure high availability and reliability.

What we're looking for

  • Bachelor’s degree in Computer Science, Computer Engineering, or a related field (preferred).
  • 8+ years of experience in data engineering or related disciplines with increasing responsibility (preferred).
  • Extensive experience on modern, large-scale, and complex data platforms (preferred).
  • Strong foundation developing and deploying Python solutions to production environments (preferred).
  • Experience leading teams to build high-throughput, scalable data pipelines (preferred).
  • Hands-on experience with AWS data services including Kinesis, S3, DynamoDB, and EventBridge at scale (preferred).
  • Proficiency in SQL, including data quality and validation practices (preferred).
  • Experience deploying software using CI/CD tools such as Azure DevOps, Jira, or Jenkins (preferred).

More like this

Similar roles

Data Engineer - Physical AI Platform

Caterpillar

Irving, TX 8 days ago $97,530–$158,480
Python Java AWS Kinesis S3 DynamoDB EventBridge SQL Microservices CI/CD Azure DevOps Jenkins Jira NoSQL API OAuth 2.0 CloudWatch Data Pipelines Agile

Lead Data Engineer

Capital One Financial

Plano, TX 109 days ago
Python Java TypeScript React Databricks Snowflake Spark Kafka AWS Redshift SQL Scala Flink DynamoDB OpenSearch UNIX Linux CI/CD Agile Microservices
4+ yrs exp

Lead Data Engineer

Capital One Financial

McLean, VA +3 106 days ago
Python SQL Scala Java AWS Azure Redshift Snowflake Spark Kafka Hadoop Hive EMR NoSQL Cassandra MySQL UNIX Linux Shell Scripting Agile Machine Learning
4+ yrs exp

Lead Data Engineer

Capital One Financial

McLean, VA +3 102 days ago
Python SQL Scala Java AWS Azure Redshift Snowflake Spark Kafka Hadoop Hive EMR NoSQL Cassandra MySQL UNIX Linux Shell Scripting Agile Machine Learning
4+ yrs exp

Lead Data Engineer

Capital One Financial

Plano, TX +2 80 days ago
Python SQL Scala Java AWS Snowflake Redshift Spark Kafka Hadoop Hive EMR NoSQL Cassandra MySQL Linux Shell Scripting Agile
4+ yrs exp

Lead Data Engineer

Capital One Financial

Chicago, IL +1 21 days ago
Python SQL Scala Java AWS Snowflake Redshift Spark Kafka Hadoop Hive EMR NoSQL Cassandra MySQL Linux Shell Scripting Agile
4+ yrs exp