AI Data Engineer

Howard Hughes Medical Institute (HHMI)

Confirmed live today High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Chevy Chase, MD
Employment
Full-time
Posted
5 days ago
Freshness
Confirmed live today

Market check

Salary context

How this pay compares to similar roles

Similar $161k
$111k $206k
below market most similar roles pay here above market

This listing doesn't post a salary. Most similar roles pay $126,800–$195,975.

Based on 240 similar postings.

Employer

About Howard Hughes Medical Institute (HHMI)

Howard Hughes Medical Institute (HHMI) is one of the largest private biomedical research organizations in the world, funding basic research and science education to advance human health and knowledge. Industry: Biomedical Research & Science Education

Howard Hughes Medical Institute (HHMI) currently has 3 open roles on FindRole.

Most-posted roles

View all roles at Howard Hughes Medical Institute (HHMI)

At a glance

TL;DR · AI Data Engineer

The AI Data Engineer joins the AI Accelerator team to build the data-engineering foundation for institutional AI applications. This role focuses on implementing pipelines, transformation patterns, and orchestration frameworks that convert institutional content into governed, AI-ready data. Key responsibilities include building AI-facing pipelines, implementing medallion architecture, managing workflow orchestration, and maintaining retrieval-supporting infrastructure like embedding pipelines and vector stores. The successful candidate will utilize Databricks, Spark, Delta Lake, Delta Live Tables, Unity Catalog, Python, and SQL. They must also possess skills in Git, Terraform, and AWS foundations. This position solves the technical challenge of ensuring high-quality, governed data flows for AI systems, connecting source-system data to downstream AI developers while managing data quality, drift detection, and sensitivity classification within a complex knowledge-management environment.

What you'll do

  • Build data pipelines to ingest, transform, and serve governed, AI-ready content from raw to gold layers.
  • Implement medallion architecture patterns using Delta Lake tables with partitioning and optimization.
  • Own the workflow-orchestration framework using Databricks Workflows and Delta Live Tables.
  • Execute governance patterns including Unity Catalog structure, sensitivity classification, and access control.
  • Build and operate retrieval-supporting infrastructure such as embedding pipelines and vector store maintenance.
  • Define source-system contracts with Operations Capabilities regarding schemas, cadence, and quality thresholds.
  • Design and operate data quality checks, freshness monitoring, and drift detection alerts.
  • Contribute to and evolve a shared library of reusable pipeline patterns and code templates.

What we're looking for

  • Must have a bachelor’s degree or equivalent.
  • Must have at least four years of hands-on experience designing, building, and operating production data pipelines.
  • Must have experience building data foundations for AI use cases, including embedding pipelines, vector stores, and retrieval evaluation.
  • Must possess depth in Databricks and Spark, including Delta Lake, Delta Live Tables, Workflows, and Unity Catalog.
  • Must be fluent in Python and SQL, with experience in PySpark, ETL patterns, Git, CI/CD, and Terraform.
  • Must have experience with workflow orchestration tools like Databricks Workflows or Airflow.
  • Must have foundational AWS knowledge, including IAM, S3, and KMS.
  • Must have experience with knowledge graphs, entity resolution, or semantic data models (preferred).

More like this

Similar roles

Data Engineer, AI Enablement

AbbVie

North Chicago, IL 25 days ago $84,500–$162,000
Python SQL ETL ELT Airflow AWS Databricks Spark Snowflake Neo4j Vector Databases RAG Knowledge Graphs Data Modeling Data Governance Jira Agile data, lineage Metadata Management
5+ yrs exp

Data Engineer

SHI International

NY +1 46 days ago $120,000–$160,000
Python SQL C# Java ETL Microsoft Azure Infrastructure as Code TDD ORM PaaS IaaS SaaS Virtualization Data Engineering Data Pipelines Version Control
5+ yrs exp

Manager, Data and AI Engineering

Nvidia

Santa Clara, CA 4 days ago $200,000–$322,000
Databricks AWS Apache Spark PySpark Delta Lake Kafka LLM RAG Unity Catalog Amazon S3 EC2 Lambda API Gateway ETL ELT CDC Vector Search Knowledge Graphs Star Schema Snowflake Schema Data Vault
10+ yrs exp

Senior Data Engineer

CoStar Group

Arlington, VA 52 days ago $128,000–$190,000
Python PySpark SQL Databricks Snowflake AWS BigQuery Spark Terraform CI/CD Kafka DynamoDB Amazon S3 Power BI Machine Learning NoSQL Object-Oriented Programming Medallion Architecture LLMs
6+ yrs exp Hybrid

Data Engineer II

CoStar Group

Arlington, VA 25 days ago $103,000–$153,000
Python PySpark SQL Databricks Snowflake AWS BigQuery Spark ETL CI/CD Terraform Kafka DynamoDB Amazon S3 Power BI Machine Learning NoSQL Object-Oriented Programming Medallion Architecture Data Governance
3+ yrs exp Hybrid

AI Data Platform Engineer

Apple Inc

Cupertino, CA 72 days ago $150,400–$277,600
Python SQL Spark PySpark Kafka Airflow Kubeflow MLflow Ray Pandas Kubernetes Docker AWS Azure GCP CI/CD RAG Vector Databases Iceberg Delta Lake Java Scala
5+ yrs exp