Principal Scientist Data Pipeline Engineer

Adobe

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
San Jose, CASeattle, WASan Francisco, CA
Salary
$268,000–$388,000 / yr
Posted
56 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $202k
This role $328k
$115k most similar roles pay here $417k

This role pays more than 96% of similar roles. Most pay $173,200–$230,700 — the shaded band above. At the midpoint, this role pays about $328k versus about $202k for comparable roles.

Based on 240 similar postings.

Employer

About Adobe

Adobe Inc. is a global software company known for creative and multimedia software products including Photoshop, Illustrator, Acrobat, and its cloud-based Creative Cloud and Document Cloud suites. Industry: Creative & Digital Experience Software

Adobe currently has 218 open roles on FindRole.

Listed pay typically runs $187,100–$270,950 across 216 roles with salary data.

Most-posted roles

View all roles at Adobe

At a glance

TL;DR · Principal Scientist Data Pipeline Engineer

Principal Scientist - Data Pipeline Engineer As a Principal Scientist - Data Pipeline Engineer, you will join the team responsible for architecting and scaling multimodal data processing pipelines and infrastructure for Adobe Firefly’s foundation models. You will build distributed, GPU-accelerated systems to transform billions of raw image, video, and audio assets into training-ready data. Your daily work involves optimizing inference throughput, identifying bottlenecks in ingestion and delivery, and designing scalable storage and indexing systems. The role requires expertise in Python and a systems-level language like C++, Rust, Go, or Java, alongside experience with frameworks such as Ray or Spark. You will also focus on data curation for training generative models and optimizing GPU inference pipelines for VLMs and LLMs. This position solves the critical challenge of ensuring high-quality, high-volume data flows to improve model learning outcomes in a multimodal environment.

What you'll do

  • Architect and optimize distributed pipelines to process billions of multimodal assets into training-ready data.
  • Increase inference throughput by optimizing batching, parallelism, and hardware utilization across the pipeline.
  • Identify and eliminate bottlenecks in storage, I/O, and compute scheduling for large-scale data processing.
  • Design scalable infrastructure to store, index, and serve massive datasets using distributed systems like Ray.
  • Make high-level architectural decisions regarding database selection, storage, and GPU cluster utilization.
  • Partner with modeling teams to translate training requirements into specific pipeline and curation specifications.
  • Perform hands-on technical leadership to bridge the gap between data engineering and applied machine learning.

What we're looking for

  • 10+ years of experience in data engineering, ML infrastructure, or distributed systems at scale.
  • Strong software engineering background with expertise in distributed frameworks like Ray or Spark.
  • Proficiency in Python and a systems-level language such as C++, Rust, Go, or Java.
  • Deep knowledge of large-scale databases, storage systems, indexing, and retrieval for billions of data points.
  • Expertise in optimizing GPU inference pipelines for VLMs, LLMs, and other large models.
  • Experience with data curation specifically for training generative or multimodal models.
  • Ability to navigate the full stack from low-level systems and GPU optimization to high-level data strategy.
  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, Machine Learning, or a related field.

More like this

Similar roles

Senior Machine Learning Engineer, AI Platform

Adobe

San Jose, CA 8 days ago $211,800$306,625
Python Go C++ Rust Java Kubernetes Distributed Systems GPU PyTorch FSDP DeepSpeed vLLM TensorRT-LLM Triton Ray Serve Cloud Infrastructure
7+ yrs exp

Principal Data Scientist

Autodesk

San Francisco, CA 29 days ago $135,000$242,000
LLMs RAG Vector Databases Predictive Modeling A/B Testing Causal Inference Data Architecture Telemetry Observability Data Modeling Sequence Models Survival Analysis LTV Propensity Frameworks 2D/3D Geometric Data
8+ yrs exp

Principal Data Scientist

AbbVie

Mettawa, IL 44 days ago $124,500$236,500
Python R SQL PySpark Scikit-learn NumPy Pandas PyTorch TensorFlow MLOps Databricks MLflow Azure ML A/B Testing Causal Inference LLM
8+ yrs exp

Principal Data Scientist

Northrop Grumman

Palmdale, CA 39 days ago $125,300$187,900
SQL Python Tableau PowerBI Apache Airflow ETL Gitlab Agile Scrum Data Mining Machine Learning Prescriptive Analytics Data Visualization
5+ yrs exp