Associate Data Engineer, AI & Data Analytics

IBM

Confirmed live yesterday Trusted

Quick summary

Work type
On-site
Location
Posted
26 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $169k
$117k most similar roles pay here $221k

This listing doesn't post a salary. Most similar roles pay $126,800–$211,200.

Based on 240 similar postings.

Employer

About IBM

IBM is a US-based global technology company providing hybrid cloud, AI, consulting, enterprise software, and IT infrastructure products and services.

IBM currently has 506 open roles on FindRole.

Listed pay typically runs $174,632–$197,165 across 5 roles with salary data.

Most-posted roles

View all roles at IBM

At a glance

TL;DR · Associate Data Engineer, AI & Data Analytics

Associate Data Engineer 2027 - AI & Data Analytics is a role within the IBM Consulting team focused on building and improving data platforms, products, and services that support analytics, machine learning, generative AI, and agentic AI solutions. You will manage the full data lifecycle, including gathering, ingestion, transformation, storage, and real-time processing while ensuring data quality and governance. The work involves solving complex problems related to database integration and creating reliable systems for large datasets. Key technologies include Python, SQL, Java, Spark, Kafka, Airflow, dbt, Databricks, Snowflake, and Delta Lake. You will also utilize tools like LangChain, LlamaIndex, and various vector databases such as Pinecone or Weaviate. The role addresses the technical challenge of creating trustworthy, observable pipelines and secure interfaces for enterprise systems using RAG, vector search, and agent orchestration patterns.

What does a Data Engineer earn?

Median $183138 from 214 postings across 58 companies.

See salary data

What you'll do

  • Design and implement scalable data architectures and management systems for cloud environments and AI use cases.
  • Optimize data pipelines, retrieval indexes, and services to improve performance, reliability, and data quality.
  • Collect and analyze structured and unstructured data to provide actionable insights for client business practices.
  • Troubleshoot data quality issues, processing failures, and inconsistencies affecting generative AI or agentic workflows.
  • Create dashboards and reports to communicate pipeline health and AI system insights to various stakeholders.
  • Ensure data integrity and security through rigorous cleaning, validation, cataloging, and governance practices.
  • Translate client requirements into technical specifications for data products, APIs, and operational models.
  • Resolve complex business problems using technologies like RAG, vector search, and agent orchestration patterns.

What we're looking for

  • High School Diploma or GED is required.
  • A Bachelor's degree in a related quantitative field is preferred.
  • Familiarity with programming or query languages such as Python, SQL, Java, Scala, or JavaScript.
  • Foundational understanding of data engineering concepts including pipelines, databases, APIs, and ETL/ELT processes.
  • Basic understanding of cloud computing environments like AWS, Azure, Google Cloud, or IBM Cloud.
  • Ability to apply foundational statistical, machine learning, or information retrieval concepts to data work.
  • Experience with tools such as Spark, Hadoop, Kafka, Airflow, dbt, Databricks, Snowflake, or Linux (preferred).
  • Familiarity with LLMs, vector databases, RAG, and AI agent architectures (preferred).

More like this

Similar roles

Associate Data Engineer, AI & Data Analytics

IBM

26 days ago
Python SQL Java Scala JavaScript AWS Azure Google Cloud IBM Cloud Spark Hadoop Kafka Airflow dbt Databricks Snowflake Delta Lake Linux LangChain LlamaIndex Pinecone Weaviate Kubernetes Git CI/CD RAG LLM

Intern Data Engineer, AI & Data Analytics

IBM

26 days ago
Python SQL Java Scala JavaScript Spark Kafka Airflow dbt Databricks Snowflake Delta Lake Linux RAG LangChain LangGraph GitHub Copilot ETL/ELT APIs Vector Databases

Associate Data Engineer

IBM

26 days ago
Python SQL Java Scala Spark Hadoop Snowflake BigQuery Airflow dbt ETL ELT AWS Azure Linux Data Modeling Machine Learning GenAI CI/CD

Associate Data Engineer

IBM

26 days ago
Python SQL Java Scala Spark Hadoop Snowflake BigQuery Airflow dbt ETL ELT Data Modeling AWS Azure Linux CI/CD Machine Learning GenAI RAG

Associate Data Engineer

IBM

26 days ago
Python SQL Java Scala Spark Hadoop Linux ETL ELT Snowflake Big Query Airflow dbt AWS Azure Data Modeling CI/CD Machine Learning GenAI RAG

Associate Data Engineer

IBM

26 days ago
Python SQL Java Scala Spark Hadoop Snowflake BigQuery dbt Airflow AWS Azure Linux ETL ELT Data Modeling CI/CD Machine Learning GenAI RAG