CM4AI Dataset Documentation

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

DescriptionIDName
Deliver machine-readable hierarchical maps of cell architecture as AI-Ready data from multimodal int...purpose-001AI-Ready Cell Architecture Maps for Biomedical AI
Address the grand challenge of interpretable genotype-phenotype learning in genomics and precision m...purpose-002Interpretable Genotype-Phenotype Learning
Establish standards, best practices, and guidelines for ethical AI-readiness in biomedical data. Thi...purpose-003Establishing AI-Readiness Standards for Biomedical Data
  • ID
    funder-001
    Name
    NIH Common Fund Bridge2AI Program
    Description
    Funded through National Institutes of Health grant 1OT2OD032742-01 (Bridge2AI Functional Genomics) and 5U54HG012513-02 (Bridge2AI Bridge Center), administered by NIH Office of the Director. Opportunity Number: OTA-21-008. Project dates: September 1, 2022 to August 31, 2026. FY 2025 funding: $5,289,382 (Direct: $4,632,095, Indirect: $657,287). Additional funding from the Frederick Thomas Fund of the University of Virginia.
📊

Composition

What do the instances represent?

DescriptionIDInstance TypeName
Triple negative breast cancer cell line (RRID:CVCL_0419) established from a metastatic site pleural ...instance-001Cultured cell line from de-identified human tissue, ethically sourced from ATCCMDA-MB-468 Breast Cancer Cell Line
Human induced pluripotent stem cell (iPSC) line (RRID:CVCL_B5P3) derived from a healthy male Norther...instance-002Cultured cell line from de-identified human tissue, ethically sourced from HipSciKOLF2.1J Induced Pluripotent Stem Cells
Near-comprehensive set of chromatin regulators encoded by the human genome analyzed via AP-MS, SEC-M...instance-003Protein targets for multi-modal analysis100 Chromatin Regulators
Set of metabolic enzymes involved in cancer, neuropsychiatric, and cardiac disorders analyzed via mu...instance-004Protein targets for multi-modal analysis100 Metabolic Enzymes
DescriptionIDName
MDA-MB-468 breast cancer cells in control/untreated conditionsubpop-001MDA-MB-468 Untreated
MDA-MB-468 breast cancer cells treated with paclitaxel chemotherapysubpop-002MDA-MB-468 Paclitaxel-Treated
MDA-MB-468 breast cancer cells treated with vorinostat chemotherapysubpop-003MDA-MB-468 Vorinostat-Treated
KOLF2.1J induced pluripotent stem cells in undifferentiated/naive statesubpop-004KOLF2.1J Undifferentiated iPSCs
KOLF2.1J iPSCs differentiated into neuronssubpop-005KOLF2.1J iPSC-Derived Neurons
KOLF2.1J iPSCs differentiated into neural progenitor cellssubpop-006KOLF2.1J iPSC-Derived Neural Progenitor Cells (NPCs)
KOLF2.1J iPSCs differentiated into cardiomyocytessubpop-007KOLF2.1J iPSC-Derived Cardiomyocytes
Access UrlsDescriptionIDName
https://doi.org/10.18130/V3/DXWOS5, https://doi.org/10.18130/V3/B35XWX, ... (+2 more)All CM4AI output data packaged as Research Object Crate (RO-Crate) packages containing datasets, met...format-001RO-Crate Packages with Provenance
MassIVE Repository (human iPSCs), MassIVE Repository (human cancer cells)Mass spectrometry data deposited to MassIVE Repository (Proteomics community-supported repository). ...format-002Mass Spectrometry Data in MassIVE
NCBI BioProject, Sequence Read Archive (SRA)Raw sequence data from CRISPR perturbation screens deposited to NCBI BioProject/Sequence Read Archiv...format-003Sequence Data in NCBI SRA
https://www.ndexbio.orgCell maps shared via Network Data Exchange (NDEx) for visualization and access. Maps can be visualiz...format-004Hierarchical Cell Maps in NDEx
https://doi.org/10.18130/V3/DXWOS5, LibraData University of VirginiaArchived RO-Crates available in University of Virginia's LibraData data archive (instance of Harvard...format-005University of Virginia Dataverse
🔍

Collection Process

How was the data acquired?

CM4AI
Cell Maps for Artificial Intelligence (CM4AI)
CM4AI is the Functional Genomics Data Generation Project in the U.S. National Institutes of Health's (NIH) Bridge to Artificial Intelligence (Bridge2AI) program. Its overarching mission is to produce ethical, AI-ready datasets of cell architecture, inferred from multimodal data collected for human cell lines, to enable transformative biomedical AI research. The project delivers machine-readable hierarchical maps of cell architecture as AI-Ready data produced from multimodal interrogation of 100 chromatin modifiers and 100 metabolic enzymes involved in cancer, neuropsychiatric, and cardiac disorders in disease-relevant cell lines under perturbed and unperturbed conditions. Data streams include immunofluorescence (IF) subcellular microscopy for spatial proteomics, affinity purification mass spectroscopy (AP-MS) and size exclusion mass spectroscopy (SEC-MS) for protein-protein interaction (PPI) data, and single-cell CRISPR-Cas perturbation screens by cell type. Input data streams are integrated via the Multi-Scale Integrated Cell (MuSIC) software pipeline employing deep learning models and community detection algorithms, and output cell maps are packaged with provenance graphs and rich metadata as AI-Ready datasets in RO-Crate format using the FAIRSCAPE framework.
en
  • Cell Maps
  • Artificial Intelligence
  • AI-Ready Data
  • Bridge2AI
  • Functional Genomics
  • Protein-Protein Interactions
  • Spatial Proteomics
  • CRISPR Perturbation
  • Hierarchical Cell Maps
  • MDA-MB-468
  • KOLF2.1J
  • iPSC
  • Breast Cancer
  • Chromatin Modifiers
  • Metabolic Enzymes
  • Immunofluorescence
  • Mass Spectrometry
  • AP-MS
  • SEC-MS
  • Perturb-Seq
  • FAIR Principles
  • RO-Crate
  • FAIRSCAPE
  • Visible Neural Networks
  • Deep Learning
DescriptionIDName
Address the limitation that machine learning models in genomics and precision medicine are typically...gap-001Black Box AI Models in Genomic Medicine
Provide fully provenanced, ethically validated, and FAIR-compliant AI-ready datasets with machine-re...gap-002Lack of AI-Ready Biomedical Datasets with Provenance
Create integrated datasets combining protein localization (spatial proteomics), protein-protein inte...gap-003Integration of Multimodal Cellular Data
RoleNameORCIDAffiliation
ContributorTrey Idekercreator-001-
ContributorJean-Christophe Bélisle-Piponcreator-002-
ContributorTimothy Clarkcreator-003-
ContributorJake Yue Chencreator-004-
ContributorNevan J Krogancreator-005-
ContributorEmma Lundbergcreator-006-
ContributorPrashant Malicreator-007-
ContributorSarah J Ratcliffecreator-008-
ContributorVardit Ravitskycreator-009-
ContributorAndrej Salicreator-010-
ContributorWade Loren Schulzcreator-011-
ContributorYing Dingcreator-012-
ContributorSamah Fodehcreator-013-
ContributorCynthia Brandtcreator-014-
ContributorPamela Payne-Fostercreator-015-
DescriptionIDName
Immunofluorescence-based staining (ICC-IF) and confocal microscopy images displaying spatial localiz...subset-001Spatial Proteomics IF Images - MDA-MB-468
Affinity purification mass spectrometry (AP-MS) data on endogenously tagged cell lines mapping prote...subset-002Protein-Protein Interaction AP-MS Data
Size exclusion chromatography coupled to mass spectrometry (SEC-MS) for proteome-wide complex/ inter...subset-003Protein-Protein Interaction SEC-MS Data
Genome-scale CRISPRi perturbation cell atlas in undifferentiated KOLF2.1J human induced pluripotent ...subset-004CRISPR Perturbation Cell Atlas
Integrated hierarchical cell maps produced by Multi-Scale Integrated Cell (MuSIC) pipeline from fusi...subset-005Hierarchical Cell Maps via MuSIC
  1. ID
    sampling-001
    Name
    Disease-Relevant Cell Line Selection
    Description
    Purposive selection of two disease-relevant cell lines: MDA-MB-468 triple negative breast cancer cell line for cancer research, and KOLF2.1J iPSCs for neuropsychiatric and cardiac disorder research. Both cell lines ethically sourced and well-characterized in the literature.
    Is Sample
    • False
    Is Random
    • False
    Is Representative
    • False
    Strategies
    • Selection of commercially available, ethically sourced cell lines
    • MDA-MB-468 chosen for triple-negative breast cancer research applications
    • KOLF2.1J chosen as reference iPSC line for large-scale collaborative studies
    • Both cell lines have extensive existing characterization data
DescriptionIDName
Automated fixation and permeabilization protocols using pipetting robot for MDA-MB-468 and KOLF2.1J ...collection-001Immunofluorescence Spatial Proteomics Imaging
Endogenous tagging of genes in cell lines followed by affinity purification mass spectrometry (AP-MS...collection-002Affinity Purification Mass Spectrometry
Size exclusion chromatography coupled to mass spectrometry (SEC-MS) for proteome-wide complex/ inter...collection-003Size Exclusion Chromatography Mass Spectrometry
Single-cell CRISPR screens using CRISPR lentiviral library targeting 100 chromatin factors with 6 gu...collection-004CRISPR Perturbation Screens
DescriptionIDName
High-resolution confocal microscopy of immunofluorescence-stained cells capturing four channels: DAP...acquisition-001Confocal Microscopy for Subcellular Imaging
State-of-the-art mass spectrometry-based proteomics including AP-MS on endogenously tagged cell line...acquisition-002Mass Spectrometry for Protein Interactions
Single-cell RNA sequencing using 10x Genomics 3'HT kit to capture transcriptional states following C...acquisition-003Single-Cell RNA Sequencing for Perturbation Mapping
Multi-Scale Integrated Cell (MuSIC) pipeline integrates PPI embeddings and image embeddings using co...acquisition-004MuSIC Pipeline Integration
DescriptionIDNamePreprocessing Details
PPI networks processed using node2vec deep learning model to reduce dimensionality and produce PPI e...preproc-001Deep Learning Embedding Generationnode2vec deep learning for PPI network dimensionality reduction, Human Protein Atlas deep learning model for image embedding, ... (+2 more)
Community detection performed on co-embedding space using multiscale community detection algorithms ...preproc-002Hierarchical Community DetectionAll-by-all similarity computation in co-embedding space, Multiscale community detection using Cytoscape algorithms, ... (+2 more)
Two-pronged annotation approach: (1) Alignment to known protein function and pathway resources inclu...preproc-003Cell Map AnnotationAlignment to Gene Ontology for functional annotation, Alignment to Reactome pathways for pathway annotation, ... (+2 more)
Standardized automated fixation and permeabilization protocols using pipetting robot. Consistent sta...preproc-004Quality Control for Imaging DataAutomated protocols for consistency, Standardized staining across all conditions, ... (+2 more)
Quality control and validation of AP-MS and SEC-MS data. Data currently undergoing quality checks be...preproc-005Quality Control for Mass Spectrometry DataQC procedures for AP-MS data, QC procedures for SEC-MS data, ... (+2 more)
Cleaning DetailsDescriptionIDName
RO-Crate packaging with metadata and provenance, JSON-Schema validation of all datasets, ... (+3 more)All datasets packaged using FAIRSCAPE framework which creates RO-Crate packages with datasets, metad...cleaning-001FAIRSCAPE AI-Readiness Packaging
GO mapping for protein functions, Reactome mapping for pathways, ... (+3 more)Data mapped to applicable standards and ontologies including Gene Ontology (GO), Reactome, Protein D...cleaning-002Data Standard Mapping
PDB structural information integration, AlphaFoldDB structure integration, ... (+3 more)Bioinformatics pipeline developed for annotating MuSIC communities with available structural informa...cleaning-003Integrative Structure Modeling Annotation
  1. ID
    maintainer-001
    Name
    CM4AI Consortium
    Description
    Multidisciplinary consortium managing dataset maintenance including University of California San Diego (lead), University of California San Francisco, Stanford University, University of Virginia, Yale University, University of Alabama at Birmingham, Simon Fraser University, and The Hastings Center. Data Governance Committee led by Jillian Parker. Ethical Review by Vardit Ravitsky and Jean-Christophe Belisle-Pipon.
    Maintainer Details
    • University of California San Diego (lead institution, Ideker Lab)
    • UCSF (Krogan Lab - protein interactions, Sali Lab - structure modeling)
    • Stanford University (Lundberg Lab - spatial proteomics)
    • University of Virginia (Clark Lab - standards and FAIRSCAPE)
    • Yale University (Schulz Lab - workforce development)
    • University of Alabama at Birmingham (Chen Lab - teaming, U-BRITE platform)
    • Simon Fraser University (Bélisle-Pipon - ethics)
    • The Hastings Center (Ravitsky - ethics)
    • Data Governance Committee (Jillian Parker)
ID
retention-001
Name
Long-Term Preservation Plan
Description
Digital data maintained according to NIH data sharing policies with long-term preservation in University of Virginia's LibraData repository supported by committed institutional funds. No planned sunset for data availability. Archived RO-Crates with persistent identifiers (ARK, future DOIs) ensure long-term accessibility and citability.
Retention Details
  • NIH data sharing policy compliance
  • UVA Dataverse institutional commitment
  • Persistent identifiers (ARK, future DOIs)
  • No planned data sunset
  • Machine-readable metadata for long-term discoverability
  1. ID
    sensitive-001
    Name
    Cell Line Origin Metadata
    Description
    While cell lines are de-identified and cannot be matched to specific individuals, metadata about cell line origins (age, sex, race of original donor) is retained for scientific context. This metadata does not constitute identifiable human subjects data under current knowledge.
    Sensitive Elements Present
    False
    Sensitivity Details
    • De-identified commercial cell lines
    • Donor demographic metadata for scientific context only
    • Cannot be matched to individuals with current knowledge
    • Ethically sourced from ATCC and HipSci
DescriptionExternal ResourcesIDName
Official project website and data portal using U-BRITE platformhttps://www.cm4ai.orgresource-001CM4AI Project Website
Federal grant information and project details for Bridge2AI Functional Genomicshttps://reporter.nih.gov/project-details/11211616resource-002NIH RePORTER Project Details
LibraData repository with archived RO-Crates and data releaseshttps://doi.org/10.18130/V3/DXWOS5, https://doi.org/10.18130/V3/B35XWX, https://doi.org/10.18130/V3/F3TD5Rresource-003University of Virginia Dataverse
Clark T, et al. Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from...https://doi.org/10.1101/2024.05.21.589311resource-004bioRxiv Publication
AI-readiness framework documentation, tutorial, and installation instructionshttps://fairscape.github.ioresource-005FAIRSCAPE Framework Documentation
Open source IMP package for integrative structure modelinghttp://integrativemodeling.orgresource-006Integrative Modeling Platform
Repository and visualization platform for cell maps and networkshttps://www.ndexbio.orgresource-007Network Data Exchange (NDEx)
Parent NIH Common Fund program supporting AI-ready biomedical datasetshttps://commonfund.nih.gov/bridge2airesource-008Bridge2AI Program
Collaboration partner for data curation and integrationhttps://www.nih-cfde.orgresource-009NIH Common Fund Data Ecosystem (CFDE)
Mass spectrometry data repository for iPSC and cancer cell dataMassIVE Repositoryresource-010MassIVE Proteomics Repository
Repository for CRISPR perturbation screen raw sequence dataNCBI BioProject, Sequence Read Archive (SRA)resource-011NCBI Sequence Read Archive
Nourreddine S, et al. A PERTURBATION CELL ATLAS OF HUMAN INDUCED PLURIPOTENT STEM CELLS. bioRxiv. 20...https://doi.org/10.1101/2024.11.03.621734resource-012Perturbation Cell Atlas Publication
🚀

Uses

What (other) tasks could the dataset be used for?

DescriptionIDName
Integrate multimodal data streams (spatial proteomics via IF imaging, protein-protein interactions v...task-001Multi-Scale Cell Mapping via MuSIC Pipeline
Enable development of visible neural networks (VNNs) and visible machine learning tools that use hie...task-002Visible Neural Network Development
Characterize cell architecture and protein interactions in disease-relevant cell lines including tre...task-003Disease-Relevant Cell Line Characterization
Develop ethical AI frameworks and governance structures for biomedical data, including Value-Sensiti...task-004Ethical AI Framework Development
DescriptionIDName
Primary intended use is training and development of artificial intelligence and machine learning mod...use-001AI Model Training for Functional Genomics
Development of visible neural networks (VNNs) that use hierarchical cell maps as interpretable model...use-002Visible Neural Network Development
Research into interpretable genotype-phenotype learning using multi-scale cell maps. Enables underst...use-003Genotype-Phenotype Mapping Research
Analysis of cellular responses to drug treatments (paclitaxel, vorinostat) to predict drug response ...use-004Drug Response and Synergy Prediction
Study of disease mechanisms in cancer, neuropsychiatric disorders, and cardiac disorders through ana...use-005Disease Mechanism Research
Use as exemplar for future AI-ready biomedical dataset development, demonstrating best practices in ...use-006Model for AI-Ready Biomedical Dataset Development
DescriptionIDName
Laboratory data from cell lines are not to be used in clinical decision-making or any context involv...discouraged-001Clinical Decision-Making Without Validation
This is an interim/beta release with data not yet in completed final form. Some datasets are under t...discouraged-002Use During Incomplete Data Release
Datasets require domain expertise for meaningful analysis and interpretation. Not suitable for use w...discouraged-003Analysis Without Domain Expertise
ID
license-001
Name
Creative Commons Attribution Non-Commercial Share-Alike
Description
Data licensed for reuse under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International license (https://creativecommons.org/licenses/by-nc-sa/4.0/). Attribution is required to the copyright holders and the Cell Maps for Artificial Intelligence project. Any publications referencing this data or derived products should cite the bioRxiv article (Clark T, et al. Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines. BioRXiv, May 2024. doi:10.1101/2024.05.21.589311) and directly cite the data collection. Commercial use requires separate license negotiation with copyright holder (UCSD, Stanford, and/or UCSF depending upon specific data package). A Data Access Committee will supervise ethical matters related to dataset distribution and potential dual licensing for commercial use. Copyright (c) 2025 The Regents of the University of California except where otherwise noted. Spatial proteomics raw image data is copyright (c) 2025 The Board of Trustees of the Leland Stanford Junior University.
License Terms
  • Attribution required to copyright holders and authors
  • Non-commercial use only (commercial requires separate license)
  • Share-alike - derivative works must use same license
  • Must cite bioRxiv publication and data collection DOI
  • Data Access Committee oversight for ethical distribution
📤

Distribution

How will the dataset be distributed?

CC BY-NC-SA 4.0
🔄

Maintenance

How will the dataset be maintained?

ID
updates-001
Name
Quarterly Data Releases and Maintenance Plan
Description
Dataset regularly updated and augmented through end of project in November 2026. Beta releases on quarterly basis with periodic data augmentation. Initial alpha release (v0.5) provided as supplemental data. March 2025 Beta (V1.4) includes perturb-seq in KOLF2.1J iPSCs, SEC-MS in iPSCs and derivatives, and IF images in MDA-MB-468 under three conditions. June 2025 Beta (V2.1) revision adds RGB IF images, ro-crate metadata corrections, and naming convention changes. Future releases will include computed cell maps and complete integration of all data streams. Long-term preservation in University of Virginia Dataverse with committed institutional support.
Frequency
Quarterly updates through November 2026; long-term preservation thereafter
Update Details
  • Alpha release v0.5 (supplemental data)
  • March 2025 Beta release V1.4 (doi:10.18130/V3/B35XWX)
  • June 2025 Beta release V2.1 (doi:10.18130/V3/F3TD5R)
  • Quarterly augmentation through November 2026
  • Future releases to include computed cell maps
  • Final release expected November 2026
  • Long-term preservation in UVA Dataverse
👥

Human Subjects

Does the dataset relate to people?

ID
hsr-001
Name
CM4AI Non-Human Subjects Research
Description
CM4AI data are distinctive within Bridge2AI in that they are non-clinical data from tissue cultures and are considered to be de-identified as they cannot be matched, with current knowledge, to a human subject. Both cell lines (MDA-MB-468 and KOLF2.1J) are commercially available, ethically sourced, de-identified cell lines. MDA-MB-468 available from ATCC. KOLF2.1J available from HipSci resource for non-profit organizations via simple MTA. Ethics team developed comprehensive plan for ethical preparation, licensing, dissemination, and data access supervision balancing openness with IP protection and commercialization monitoring.
Involves Human Subjects
False
IRB Approval
  • Not applicable - de-identified cell lines from commercial sources
Ethics Review Board
  • CM4AI Ethics Module (Vardit Ravitsky, Jean-Christophe Bélisle-Pipon)
  • Data Access Committee (Jillian Parker)
  • Bridge2AI Ethics Working Group participation
Special Populations
  • MDA-MB-468 derived from 51-year-old black female (de-identified)
  • KOLF2.1J derived from healthy male Northern European donor (de-identified)
Generated on 2025-12-20 19:23:28 using Bridge2AI Data Sheets Schema