Principal Scientist, Data Science (Data Products, Integration & Analysis)

Johnson & Johnson

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Spring House, PAHorsham, PACambridge, MARaritan, NJTitusville, NJ
Salary
$117,000–$201,250 / yr
Posted
30 days ago
Freshness
Confirmed live yesterday
Closes
Sep 26, 2026

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $169k
This role $159k
$104k most similar roles pay here $238k

This role pays less than 65% of similar roles. Most pay $136,546–$202,352 — the shaded band above. At the midpoint, this role pays about $159k versus about $169k for comparable roles.

Based on 240 similar postings.

Employer

About Johnson & Johnson

Johnson & Johnson is a multinational corporation operating in three main segments: consumer health products, pharmaceuticals, and medical devices, known for brands like Tylenol, Band-Aid, and Janssen. Industry: Pharmaceuticals & Medical Devices

Johnson & Johnson currently has 46 open roles on FindRole.

Listed pay typically runs $117,000–$201,250 across 42 roles with salary data.

Most-posted roles

View all roles at Johnson & Johnson

At a glance

TL;DR · Principal Scientist, Data Science (Data Products, Integration & Analysis)

Principal Scientist, Data Science (Data Products, Integration & Analysis) leads the design, implementation, and evolution of scientific data products to support AI-enabled drug discovery and development. This role focuses on creating scalable, interoperable, and AI-ready data assets that connect discovery, preclinical, clinical, safety, and real-world evidence domains. The individual will establish data architecture, metadata frameworks, and integration strategies to enable semantic reasoning, knowledge graphs, GraphRAG, and agentic AI applications. Key responsibilities include building predictive AIML models for translational safety decision making and developing curated datasets and feature stores. Technical requirements include expertise in AWS-based platforms, data modeling, and governance, alongside familiarity with standards like SEND, SDTM, ADaM, MedDRA, FHIR, OMOP, DICOM, and omics data. The role solves complex problems in the drug development lifecycle by harmonizing heterogeneous scientific data to improve patient safety and research outcomes.

What you'll do

  • Lead the design and implementation of scientific data products to support AI-enabled drug discovery and development.
  • Create scalable, interoperable, and AI-ready data assets across preclinical, clinical, safety, and real-world evidence domains.
  • Establish data architecture, integration strategies, and metadata frameworks to support knowledge graphs and agentic AI applications.
  • Develop predictive AI/ML models to support data-driven translational safety decision-making.
  • Design integration frameworks to harmonize heterogeneous scientific data sources including SEND, SDTM, ADaM, and MedDRA.
  • Define and implement data quality standards, lineage, provenance, and FAIR data principles for scientific products.
  • Manage the implementation of scientific data products on AWS platforms in collaboration with engineering teams.
  • Translate complex scientific questions from stakeholders into scalable technical solutions and data products.

What we're looking for

  • Master’s or PhD in Computer Science, Data Engineering, Bioinformatics, Biomedical Informatics, Information Systems, Computational Biology, or a related scientific discipline.
  • 5+ years of experience in scientific data engineering, data architecture, data products, or life sciences informatics.
  • Demonstrated experience designing and delivering enterprise-scale scientific data products.
  • Experience supporting drug discovery, development, clinical research, or pharmacovigilance organizations.
  • Experience developing predictive models within drug discovery, development, clinical research, or pharmacovigilance environments.
  • Expertise in data architecture, modeling, product design, metadata management, and governance.
  • Proficiency with AWS-based data platforms, including data lakes, lakehouses, and distributed processing.
  • Familiarity with scientific data standards such as SEND, SDTM, ADaM, and MedDRA.

More like this

Similar roles

Principal Data Scientist

Autodesk

San Francisco, CA 29 days ago $135,000$242,000
LLMs RAG Vector Databases Predictive Modeling A/B Testing Causal Inference Data Architecture Telemetry Observability Data Modeling Sequence Models Survival Analysis LTV Propensity Frameworks 2D/3D Geometric Data
8+ yrs exp

Principal Data Scientist

AbbVie

Mettawa, IL 44 days ago $124,500$236,500
Python R SQL PySpark Scikit-learn NumPy Pandas PyTorch TensorFlow MLOps Databricks MLflow Azure ML A/B Testing Causal Inference LLM
8+ yrs exp

Principal Data Scientist

Northrop Grumman

Palmdale, CA 39 days ago $125,300$187,900
SQL Python Tableau PowerBI Apache Airflow ETL Gitlab Agile Scrum Data Mining Machine Learning Prescriptive Analytics Data Visualization
5+ yrs exp

Principal Data Scientist

Microsoft

36 days ago $130,900$277,200
Causal Inference Machine Learning Python R SQL LLMs AI Agents DoWhy EconML CausalML Matplotlib Seaborn Plotly Statistics Model Validation
10+ yrs exp