CM4AI Dataset Documentation

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

DescriptionID
Deliver machine-readable hierarchical maps of cell architecture as AI-Ready data from multimodal interrogation of disease-relevant cell lines to enable transformative biomedical AI research. CM4AI produces integrated cell maps from spatial proteomics, protein-protein interactions, and genetic perturbations using state-of-the-art mass spectrometry, cell imaging, and CRISPR technologies.
purpose-001
Address the grand challenge of interpretable genotype-phenotype learning in genomics and precision medicine. Machine learning models are often black boxes predicting phenotypes from genotypes without understanding the mechanisms. CM4AI enables visible machine learning systems informed by multi-scale cell and tissue architecture, allowing AI tools to interrogate how protein assemblies in the cell affect cell-level phenotype predictions.
purpose-002
Establish standards, best practices, and guidelines for ethical AI-readiness in biomedical data. This includes implementing FAIR principles, computing machine-readable provenance graphs, characterizing and validating all datasets with JSON-Schema mini-data-dictionaries, and mapping data elements to public ontology vocabularies where appropriate.
purpose-003
Enable development of visible neural networks (VNNs) and visible machine learning tools that use hierarchical cell maps as interpretable structures for AI model architectures, allowing interrogation of how protein assemblies affect cell-level phenotypes and interpretation of genetic variants and mutations.
purpose-004
  • ID
    funder-001
    Description
    NIH Common Fund Bridge2AI Program funded through National Institutes of Health grant 1OT2OD032742-01 (Bridge2AI Functional Genomics) and 5U54HG012513-02 (Bridge2AI Bridge Center), administered by NIH Office of the Director. Opportunity Number OTA-21-008. Project dates: September 1, 2022 to August 31, 2026. FY 2025 funding: $5,289,382 (Direct: $4,632,095, Indirect: $657,287). Additional funding from the Frederick Thomas Fund of the University of Virginia.
📊

Composition

What do the instances represent?

DescriptionID
MDA-MB-468 Breast Cancer Cell Line (RRID:CVCL_0419): Triple negative breast cancer cell line established from a metastatic site pleural effusion of a 51-year-old black female with a metastatic mammary adenocarcinoma, available from ATCC. This cell line has been extensively used to study triple-negative breast cancer and is well characterized with transcriptomic, mutational profile, and whole-genome sequencing data available. Cells are analyzed under three conditions: untreated, paclitaxel-treated, and vorinostat-treated. Cultured cell line from de-identified human tissue, ethically sourced from ATCC.
instance-001
KOLF2.1J Induced Pluripotent Stem Cells (RRID:CVCL_B5P3): Human induced pluripotent stem cell (iPSC) line derived from a healthy male Northern European donor, available from the Human Induced Pluripotent Stem Cells Initiative (HipSci) resource. Available for access by non-for-profit organizations via a simple MTA. Analyzed in undifferentiated state and after differentiation into neurons, neural progenitor cells (NPCs), and cardiomyocytes. Cultured cell line from de-identified human tissue, ethically sourced from HipSci.
instance-002
100 Chromatin Regulators: Near-comprehensive set of chromatin regulators encoded by the human genome analyzed via AP-MS, SEC-MS, IF imaging, and CRISPR perturbation screens across different cell states and treatment conditions. 17 genes endogenously tagged in MDA-MB-468 with AP-MS data under three conditions, with 34 additional genes in process. SEC-MS identified 72/100 chromatin modifiers, with 52 being integral components of protein complexes. Protein targets for multi-modal analysis.
instance-003
100 Metabolic Enzymes: Set of metabolic enzymes involved in cancer, neuropsychiatric, and cardiac disorders analyzed via multimodal interrogation including mass spectrometry, imaging, and perturbation screens. Protein targets for multi-modal analysis.
instance-004
DescriptionID
MDA-MB-468 Untreated - MDA-MB-468 breast cancer cells in control/untreated condition
subpop-001
MDA-MB-468 Paclitaxel-Treated - MDA-MB-468 breast cancer cells treated with paclitaxel chemotherapy
subpop-002
MDA-MB-468 Vorinostat-Treated - MDA-MB-468 breast cancer cells treated with vorinostat chemotherapy
subpop-003
KOLF2.1J Undifferentiated iPSCs - KOLF2.1J induced pluripotent stem cells in undifferentiated/naive state
subpop-004
KOLF2.1J iPSC-Derived Neurons - KOLF2.1J iPSCs differentiated into neurons
subpop-005
KOLF2.1J iPSC-Derived Neural Progenitor Cells (NPCs) - KOLF2.1J iPSCs differentiated into neural progenitor cells
subpop-006
KOLF2.1J iPSC-Derived Cardiomyocytes - KOLF2.1J iPSCs differentiated into cardiomyocytes
subpop-007
Access UrlsDescriptionIDName
https://doi.org/10.18130/V3/DXWOS5, https://doi.org/10.18130/V3/F3TD5R, https://dataverse.lib.virginia.edu/
Primary distribution through LibraData, University of Virginia's Dataverse instance (an NIH-approved generalist repository). All output data packaged as Research Object Crates (RO-Crate version 1.2) using FAIRSCAPE framework. RO-Crates contain datasets, metadata, provenance graphs, and software with resolvable references. Metadata expressed in JSON-LD graph language using vocabularies from Schema.org, EVI Evidence Graph Ontology, and other well-defined public ontologies. Data releases available with DOI 10.18130/V3/DXWOS5 (March 2025) and DOI 10.18130/V3/F3TD5R (June 2025). Quarterly updates planned through November 2026.
format-001University of Virginia Dataverse - RO-Crate Packages
https://dataverse.lib.virginia.edu/
Immunofluorescence images in standard confocal microscopy formats distributed as ZIP archives (2.6-4.6 GB per condition). RGB immunofluorescent images included in June 2025 release. Four-channel imaging data with DAPI (nuclei), calreticulin (ER), tubulin (microtubules), and protein-of-interest antibodies. Spatial localization of 464-563 proteins in MDA-MB-468 cells under three treatment conditions.
format-002Immunofluorescence Imaging Data - ZIP Archives
https://massive.ucsd.edu/
AP-MS and SEC-MS data in standard mass spectrometry data formats suitable for deposition to MassIVE Repository. Includes raw data and processed interaction networks for both human iPSCs and cancer cells. Data available with persistent identifiers for reproducibility.
format-003Mass Spectrometry Data - MassIVE Repository
https://www.ncbi.nlm.nih.gov/bioproject/, https://www.ncbi.nlm.nih.gov/sra/
Raw sequence data from CRISPR perturbation screens in FASTQ format deposited to NCBI BioProject and Sequence Read Archive (SRA) for long-term archival and public access. Processed cell atlas data in RO-Crate format with JSON metadata.
format-004CRISPR Perturbation Screens - NCBI SRA
https://www.ndexbio.org/
Hierarchical cell maps as directed acyclic graphs (DAG) compatible with Cytoscape, HiView, and NDEx. Cell maps shared via Network Data Exchange (NDEx) and can be visualized in a web browser or accessed via tools such as Cytoscape, HiView, and the Python ndex2 library. Includes JSON-Schema data dictionaries for all datasets.
format-005Hierarchical Cell Maps - NDEx
http://www.cm4ai.org
Open dissemination of CM4AI-generated data, maps, and tools via the CM4AI web portal using the U-BRITE platform to support open and trustworthy data management and sharing. CodeFest event materials and tutorials available.
format-006CM4AI Web Portal
🔍

Collection Process

How was the data acquired?

CM4AI
Cell Maps for Artificial Intelligence (CM4AI)
CM4AI is the Functional Genomics Data Generation Project in the U.S. National Institutes of Health's (NIH) Bridge to Artificial Intelligence (Bridge2AI) program. Its overarching mission is to produce ethical, AI-ready datasets of cell architecture, inferred from multimodal data collected for human cell lines, to enable transformative biomedical AI research. The project delivers machine-readable hierarchical maps of cell architecture as AI-Ready data produced from multimodal interrogation of 100 chromatin modifiers and 100 metabolic enzymes involved in cancer, neuropsychiatric, and cardiac disorders in disease-relevant cell lines under perturbed and unperturbed conditions. Data streams include immunofluorescence (IF) subcellular microscopy for spatial proteomics, affinity purification mass spectroscopy (AP-MS) and size exclusion mass spectroscopy (SEC-MS) for protein-protein interaction (PPI) data, and single-cell CRISPR-Cas perturbation screens by cell type. Input data streams are integrated via the Multi-Scale Integrated Cell (MuSIC) software pipeline employing deep learning models and community detection algorithms, and output cell maps are packaged with provenance graphs and rich metadata as AI-Ready datasets in RO-Crate format using the FAIRSCAPE framework.
en
  • Cell Maps
  • Artificial Intelligence
  • AI-Ready Data
  • Bridge2AI
  • Functional Genomics
  • Protein-Protein Interactions
  • Spatial Proteomics
  • CRISPR Perturbation
  • Hierarchical Cell Maps
  • MDA-MB-468
  • KOLF2.1J
  • iPSC
  • Breast Cancer
  • Chromatin Modifiers
  • Metabolic Enzymes
  • Immunofluorescence
  • Mass Spectrometry
  • AP-MS
  • SEC-MS
  • Perturb-Seq
  • FAIR Principles
  • RO-Crate
  • FAIRSCAPE
  • Visible Neural Networks
  • Deep Learning
  • MuSIC Pipeline
DescriptionID
Address the limitation that machine learning models in genomics and precision medicine are typically difficult-to-interpret black boxes by providing hierarchical cell maps that enable visible machine learning systems built directly on knowledge maps of cell and tissue architecture.
gap-001
Provide fully provenanced, ethically validated, and FAIR-compliant AI-ready datasets with machine-readable provenance graphs, complete schemas, validation procedures, and data sheets that can be reliably processed by AI applications with full explainability.
gap-002
Create integrated datasets combining protein localization (spatial proteomics), protein-protein interactions (AP-MS and SEC-MS), and transcriptional states (CRISPR perturbation screens) at multiple scales, enabling complex multi-modal AI analyses not feasible with single data types.
gap-003
RoleNameORCIDAffiliation
Contributorcreator-001-
Contributorcreator-002-
Contributorcreator-003-
Contributorcreator-004-
Contributorcreator-005-
Contributorcreator-006-
Contributorcreator-007-
Contributorcreator-008-
Contributorcreator-009-
Contributorcreator-010-
Contributorcreator-011-
Contributorcreator-012-
Contributorcreator-013-
Contributorcreator-014-
Contributorcreator-015-
DescriptionID
Spatial Proteomics IF Images - MDA-MB-468: Immunofluorescence-based staining (ICC-IF) and confocal microscopy images displaying spatial localization of 464-563 proteins of interest in MDA-MB-468 breast cancer cells under three conditions: untreated, paclitaxel-treated, and vorinostat-treated. Nuclei stained with DAPI (blue channel), endoplasmic reticulum with calreticulin antibody (yellow channel), microtubules with tubulin antibody (red channel), and antibody against protein of interest (green channel). Generated by Lundberg Lab at Stanford University using automated fixation and permeabilization protocols. Data released as ZIP archives (2.6-4.6 GB per condition) via University of Virginia Dataverse.
subset-001
Protein-Protein Interaction AP-MS Data: Affinity purification mass spectrometry (AP-MS) data on endogenously tagged cell lines mapping protein-protein interactions of chromatin regulators. 17 genes endogenously tagged in MDA-MB-468 with data acquired under three conditions (untreated, paclitaxel, vorinostat). 34 additional genes currently in tagging process. Orthogonal approach to SEC-MS for PPI mapping. Data deposited to MassIVE Repository for human cancer cells.
subset-002
Protein-Protein Interaction SEC-MS Data: Size exclusion chromatography coupled to mass spectrometry (SEC-MS) for proteome-wide complex/interaction mapping. Performed on MDA-MB-468 cells under three conditions (untreated, paclitaxel, vorinostat) and on KOLF2.1J iPSCs and derivatives (undifferentiated, neurons, NPCs, cardiomyocytes). Detected PPI profiles of over 1,000 complexes in MDA-MB-468 cells and over 700 protein complexes in iPSCs and differentiated neurons. Identified 72/100 chromatin modifiers with 52 being integral components of protein complexes. Data deposited to MassIVE Repository for human iPSCs and cancer cells.
subset-003
CRISPR Perturbation Cell Atlas: Genome-scale CRISPRi perturbation cell atlas in undifferentiated KOLF2.1J human induced pluripotent stem cells (hiPSCs) mapping transcriptional and fitness phenotypes associated with 11,739 targeted genes. Single-cell CRISPR screens performed using 10x Genomics 3'HT kit with CRISPR lentiviral library targeting 100 chromatin factors with 6 guide RNAs per gene. Screens conducted in MDA-MB-468 cells under 3 conditions (no treatment, paclitaxel, vorinostat) and KOLF2.1J iPSC in undifferentiated state. Includes raw sequence data deposited to NCBI BioProject/ Sequence Read Archive (SRA) and processed cell atlas data in RO-Crate format.
subset-004
Hierarchical Cell Maps via MuSIC: Integrated hierarchical cell maps produced by Multi-Scale Integrated Cell (MuSIC) pipeline from fusion of protein localization (IF images) and protein-protein interaction data (AP-MS and SEC-MS). Cell maps are hierarchical directed acyclic graphs (DAG) where each node represents an assembly of proteins in proximity at a given scale, spanning from large assemblies representing cell compartments to small assemblies of protein complexes. Maps contain 10 layers of depth representing communities at multiple resolutions, with communities at smaller distance nested inside larger communities similar to physical compartments within a cell. Not included in current interim releases but planned for future data releases.
subset-005
  • ID
    sampling-001
    Description
    Disease-Relevant Cell Line Selection: Purposive selection of two disease-relevant cell lines: MDA-MB-468 triple negative breast cancer cell line for cancer research, and KOLF2.1J iPSCs for neuropsychiatric and cardiac disorder research. Both cell lines ethically sourced and well-characterized in the literature. MDA-MB-468 chosen for triple-negative breast cancer research applications. KOLF2.1J chosen as reference iPSC line for large-scale collaborative studies. Both cell lines have extensive existing characterization data. Not a representative sample of all biological variants.
DescriptionID
Immunofluorescence Spatial Proteomics Imaging: Automated fixation and permeabilization protocols using pipetting robot for MDA-MB-468 and KOLF2.1J cell lines. Immunofluorescence-based staining (ICC-IF) with confocal microscopy to capture spatial subcellular organization. Completed spatial proteomics mapping of 100 chromatin regulators in MDA-MB-468 cells under three conditions (untreated, paclitaxel, vorinostat), with 500 additional proteins pending from genetic perturbations and PPI results. Antibodies from Human Protein Atlas resource. Generated by Lundberg Lab at Stanford University.
collection-001
Affinity Purification Mass Spectrometry: Endogenous tagging of genes in cell lines followed by affinity purification mass spectrometry (AP-MS) to map protein-protein interactions. 17 genes endogenously tagged in MDA-MB-468 with AP-MS data acquired under three conditions (untreated, paclitaxel, vorinostat). 34 additional genes currently in tagging process. Orthogonal approach to SEC-MS for comprehensive PPI mapping. Generated by Krogan Laboratory at UCSF.
collection-002
Size Exclusion Chromatography Mass Spectrometry: Size exclusion chromatography coupled to mass spectrometry (SEC-MS) for proteome-wide complex/interaction mapping. Performed in Krogan Laboratory at UCSF. Conducted on MDA-MB-468 cells under three conditions and on KOLF2.1J iPSCs and derivatives (undifferentiated, NPCs, neurons, cardiomyocytes). Enabled detection of over 1,000 protein complexes in MDA-MB-468 cells and over 700 complexes in iPSCs, with thousands of proteins exhibiting differential elution profiles between control and treated cells.
collection-003
CRISPR Perturbation Screens: Single-cell CRISPR screens using CRISPR lentiviral library targeting 100 chromatin factors with 6 guide RNAs per gene. Generated and characterized MDA-MB-468 and KOLF2.1J CRISPR lines expressing inducible dCas9. Screens performed in MDA-MB-468 cells under 3 conditions (no treatment, paclitaxel, vorinostat) and in undifferentiated KOLF2.1J iPSCs using 10x Genomics 3'HT kit. Genome-scale screens mapping transcriptional and fitness phenotypes for 11,739 targeted genes. Generated by Mali Laboratory at UC San Diego.
collection-004
DescriptionID
Confocal Microscopy for Subcellular Imaging: High-resolution confocal microscopy of immunofluorescence-stained cells capturing four channels: DAPI (nuclei, blue), calreticulin antibody (ER, yellow), tubulin antibody (microtubules, red), and antibody against protein of interest (green). Images processed using Human Protein Atlas deep learning model to reduce dimensionality, producing image embeddings containing information about protein localization.
acquisition-001
Mass Spectrometry for Protein Interactions: State-of-the-art mass spectrometry-based proteomics including AP-MS on endogenously tagged cell lines and SEC-MS for proteome-wide complex mapping. PPI networks processed using node2vec deep learning model to reduce dimensionality, producing PPI embeddings containing information about protein interactions.
acquisition-002
Single-Cell RNA Sequencing for Perturbation Mapping: Single-cell RNA sequencing using 10x Genomics 3'HT kit to capture transcriptional states following CRISPR perturbations. Generates genome-scale perturbation cell atlas mapping transcriptional and fitness phenotypes. Raw sequence data deposited to NCBI BioProject/Sequence Read Archive (SRA).
acquisition-003
MuSIC Pipeline Integration: Multi-Scale Integrated Cell (MuSIC) pipeline integrates PPI embeddings and image embeddings using contrastive deep learning to obtain co-embeddings for each protein. Community detection performed based on all-by-all similarities of protein pairs in co-embedding space, producing hierarchical cell maps as final output. Maps annotated using Gene Ontology, Reactome pathways, and large language model approaches.
acquisition-004
DescriptionID
Deep Learning Embedding Generation: PPI networks processed using node2vec deep learning model to reduce dimensionality and produce PPI embeddings. IF images processed using Human Protein Atlas deep learning model to reduce dimensionality and produce image embeddings. PPI and image embeddings integrated to obtain co-embeddings using contrastive deep learning, learning co-embeddings such that original embeddings can be reconstructed with minimal information loss.
preproc-001
Hierarchical Community Detection: Community detection performed on co-embedding space using multiscale community detection algorithms implemented in Cytoscape. Produces hierarchical directed acyclic graphs (DAG) of protein assemblies at multiple resolutions, with 10 layers of depth representing communities from large cell compartments to small protein complexes.
preproc-002
Cell Map Annotation: Two-pronged annotation approach: (1) Alignment to known protein function and pathway resources including Gene Ontology (GO) and Reactome to determine protein assemblies with high overlap with known cell biology, and (2) Large language model (LLM) approach to name sets of proteins and assign name confidence scores.
preproc-003
Quality Control for Imaging Data: Standardized automated fixation and permeabilization protocols using pipetting robot. Consistent staining protocols across conditions using Human Protein Atlas antibodies. Quality control of imaging data before release and processing through MuSIC pipeline.
preproc-004
Quality Control for Mass Spectrometry Data: Quality control and validation of AP-MS and SEC-MS data. Data currently undergoing quality checks before public release. Mass spectrometry data for human iPSCs deposited to MassIVE Repository, and data for human cancer cells also deposited to MassIVE Repository.
preproc-005
DescriptionID
FAIRSCAPE AI-Readiness Packaging: All datasets packaged using FAIRSCAPE framework which creates RO-Crate packages with datasets, metadata, provenance graphs, and software. FAIRSCAPE-CLI validates inputs and creates output RO-Crate packages. FAIRSCAPE server assigns persistent resolvable globally unique identifiers (ARK scheme with DOIs as supplementary PIDs for final-state publishable work), decomposes RO-Crates into components, and computes end-to-end provenance entailments using EVI Evidence Graph Ontology.
cleaning-001
Data Standard Mapping: Data mapped to applicable standards and ontologies including Gene Ontology (GO), Reactome, Protein Data Bank (PDB), AlphaFold Protein Structure Database, Schema.org, and EVI Evidence Graph Ontology. Keywords mapped to controlled vocabularies from NCI Thesaurus, BioAssay Ontology, Cell Ontology, CHEBI, Experimental Factor Ontology, and other ontologies.
cleaning-002
Integrative Structure Modeling Annotation: Bioinformatics pipeline developed for annotating MuSIC communities with available structural information from PDB, AlphaFoldDB, crosslinking mass spectrometry, and prediction of disordered sequence segments. Communities ranked by structural information amount as proxy for integrative modeling feasibility. Integrative structure modeling performed using Python Modeling Interface and Integrative Modeling Platform (IMP) version 2.18.
cleaning-003
DescriptionIDMaintainer DetailsName
Multi-institutional consortium managing dataset maintenance including University of California San Diego (lead institution), Stanford University (spatial proteomics), University of California San Francisco (protein-protein interactions), University of Virginia (standards and infrastructure), and Yale University (skills development).
maintainer-001UCSD - Project leadership and data integration, Stanford - Spatial proteomics imaging, UCSF - Mass spectrometry PPI data, UVA - Standards module and Dataverse hosting, Yale - Workforce developmentCM4AI Consortium
Supervised by Data Access Committee for ethical matters related to dataset distribution and potential dual licensing for commercial use. Contact: Jillian Parker (jillianparker@health.ucsd.edu). Ethical review by Vardit Ravitsky (ravitskyv@thehastingscenter.org) and Jean-Christophe Belisle-Pipon (jean-christophe_belisle-pipon@sfu.ca).
maintainer-002Data Access Committee oversight, Contact - jillianparker@health.ucsd.edu, Ethics review - ravitskyv@thehastingscenter.org, Ethics review - jean-christophe_belisle-pipon@sfu.caData Governance Committee
  1. ID
    sensitive-001
    Name
    De-identified Cell Line Data
    Description
    Data derived from commercially available de-identified human cell lines. While cell lines are considered de-identified and cannot be matched to human subjects with current knowledge, ethical considerations are embedded throughout data generation and AI system development. Ethics team develops tools for enhancing awareness of ethical, legal, and social ramifications. Mixed empirical research combines qualitative and quantitative methods to capture community insights for guidelines and best practices.
    Sensitive Elements Present
    False
    Sensitivity Details
    • De-identified cell lines (MDA-MB-468, KOLF2.1J)
    • No patient-identifiable information
    • Commercially available with MTAs
    • Ethically sourced from ATCC and HipSci
    • Value-Sensitive Design framework applied
    • Axiological repository of values for design standards
ID
retention-001
Name
Long-term Preservation in UVA Dataverse
Description
Long-term preservation in the University of Virginia Dataverse, supported by committed institutional funds. All dataset versions preserved with persistent DOIs for reproducibility. No planned deletion or retention limits. Quarterly updates through November 2026, with final release preserved indefinitely.
Retention Details
  • Indefinite preservation in UVA Dataverse
  • Committed institutional funding for long-term hosting
  • All versions preserved with persistent DOIs
  • No planned deletion or data expiration
  • Quarterly updates through project end (November 2026)
🚀

Uses

What (other) tasks could the dataset be used for?

DescriptionID
Multi-Scale Cell Mapping via MuSIC Pipeline: Integrate multimodal data streams (spatial proteomics via IF imaging, protein-protein interactions via AP-MS and SEC-MS, and genetic perturbations via CRISPR screens) using the Multi-Scale Integrated Cell (MuSIC) software pipeline employing deep learning models and community detection algorithms to produce hierarchical cell maps.
task-001
Disease-Relevant Cell Line Characterization: Characterize cell architecture and protein interactions in disease-relevant cell lines including treated and untreated MDA-MB-468 breast cancer cells (with paclitaxel and vorinostat) and differentiated and naive KOLF2.1J induced pluripotent stem cells (iPSCs) differentiated into neurons and cardiomyocytes.
task-002
Ethical AI Framework Development: Develop ethical AI frameworks and governance structures for biomedical data, including Value-Sensitive Design methodologies, axiological repositories, CM4AI Life Cycle framework, and guidelines for responsible design of datasets and AI technologies.
task-003
Skills and Workforce Development: Recruit and train a diverse biomedical AI/ML workforce through asynchronous virtual training, hosted virtual events (CodeFest), and in-person internships at Yale University and UC San Diego, with emphasis on underrepresented communities through partnerships with historically black colleges and universities.
task-004
DescriptionID
AI Model Training for Functional Genomics: Primary intended use is training and development of artificial intelligence and machine learning models for functional genomics research. AI-ready datasets with full provenance, metadata, and validation enable immediate use in AI/ML pipelines without reformatting.
use-001
Visible Neural Network Development: Development of visible neural networks (VNNs) that use hierarchical cell maps as interpretable model architectures. Unlike black box models, VNNs built on cell maps allow interrogation of how protein assemblies affect cell-level phenotypes, enabling interpretation of genetic variants and mutations in the context of cellular mechanisms.
use-002
Genotype-Phenotype Mapping Research: Research into interpretable genotype-phenotype learning using multi-scale cell maps. Enables understanding of mechanisms by which genotypes translate to phenotypes, supporting precision medicine applications.
use-003
Cellular Process Analysis: Analysis of cell architectural changes and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. Suitable for bioinformatics analysis requiring domain expertise.
use-004
  • ID
    prohibited-001
    Description
    Clinical Decision-Making: Laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval. This is research data from cell lines, not clinical diagnostic data.
ID
license-001
Name
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International
Description
Dataset licensed for reuse under CC BY-NC-SA 4.0 license (https://creativecommons.org/licenses/by-nc-sa/4.0/). Attribution is required to the copyright holders and the authors. Data are Copyright (c) 2025 The Regents of the University of California except where otherwise noted. Spatial proteomics raw image data is copyright (c) 2025 The Board of Trustees of the Leland Stanford Junior University. Commercial use of data requires a separate license negotiation with the copyright holder (UCSD, Stanford, and/or UCSF depending upon the specific data package in question). A Data Access Committee will supervise ethical matters related to dataset distribution and potential dual licensing for commercial use. All CM4AI integration and packaging software is freely available and licensed under nonrestrictive open-source licenses (FAIRSCAPE: MIT License, MuSIC: BSD-3 License, IMP: open source). No regulatory restrictions apply. Data derived from commercially available de-identified human cell lines. No human subjects research. No FDA regulation applicable. No export control restrictions.
License Terms
  • CC BY-NC-SA 4.0 license for dataset
  • Attribution required to copyright holders and authors
  • Copyright 2025 Regents of University of California (UCSD)
  • Spatial proteomics copyright 2025 Stanford University
  • Commercial use requires separate license negotiation
  • Contact UCSD, Stanford, or UCSF for commercial licensing
  • Data Access Committee supervises dual licensing
  • Citation required - Clark T, Parker J, et al. BioRxiv 2024 DOI 10.1101/2024.05.21.589311
  • Software under open source licenses (MIT, BSD-3)
  • No regulatory restrictions or export controls
  • No FDA regulation applicable
  • Ethically sourced de-identified cell lines
📤

Distribution

How will the dataset be distributed?

CC BY-NC-SA 4.0
🔄

Maintenance

How will the dataset be maintained?

ID
updates-001
Name
Quarterly Data Releases and Augmentation
Description
Dataset regularly updated and augmented through the end of the project in November 2026. Updates on a quarterly basis. Current releases are interim beta releases with some datasets under temporary pre-publication embargo. Computed cell maps will be added in future releases. Dataset versioned with major and minor releases tracked in Dataverse. March 2025 release (V1) and June 2025 release (V2) with corrections to ro-crate metadata and naming conventions. All versions preserved with DOIs for reproducibility. Long-term preservation in the University of Virginia Dataverse, supported by committed institutional funds.
Frequency
Quarterly updates through November 2026
Update Details
  • Quarterly data releases with version tracking
  • March 2025 release (V1) - Initial beta release
  • June 2025 release (V2) - Corrections and RGB images
  • Computed cell maps planned for future releases
  • All versions preserved with persistent DOIs
👥

Human Subjects

Does the dataset relate to people?

ID
hsr-001
Name
CM4AI Non-Clinical Cell Line Data
Description
CM4AI data are distinctive within Bridge2AI in that they are non-clinical (from tissue cultures) and are considered to be de-identified as they cannot be matched, with current knowledge, to a human subject. Data derived from commercially available de-identified human cell lines: MDA-MB-468 breast cancer cells (ATCC, RRID:CVCL_0419) and KOLF2.1J induced pluripotent stem cells (HipSci, RRID:CVCL_B5P3). No direct human subjects research. FDA Regulated: No. Ethically sourced cell lines with appropriate material transfer agreements. Ethics team employs Value-Sensitive Design (VSD) methodology, CM4AI Life Cycle framework, and mixed empirical research to ensure responsible design of datasets and AI technologies.
Involves Human Subjects
False
Ethics Review Board
  • Data Access Committee for ethical oversight
  • Ethics team (Vardit Ravitsky, Jean-Christophe Belisle-Pipon)
  • Value-Sensitive Design methodology framework
  • CM4AI Life Cycle governance framework
Generated on 2026-04-21 15:34:03 using Bridge2AI Data Sheets Schema