id: https://doi.org/10.18130/V3/DXWOS5
name: CM4AI
title: Cell Maps for Artificial Intelligence (CM4AI)
description: 'CM4AI is the Functional Genomics Data Generation Project in the U.S. National Institutes of Health''s (NIH) Bridge to Artificial Intelligence (Bridge2AI) program. Its overarching mission is to produce ethical, AI-ready datasets of cell architecture, inferred from multimodal data collected for human cell lines, to enable transformative biomedical AI research. The project delivers machine-readable hierarchical maps of cell architecture as AI-Ready data produced from multimodal interrogation of 100 chromatin modifiers and 100 metabolic enzymes involved in cancer, neuropsychiatric, and cardiac disorders in disease-relevant cell lines under perturbed and unperturbed conditions. Data streams include immunofluorescence (IF) subcellular microscopy for spatial proteomics, affinity purification mass spectroscopy (AP-MS) and size exclusion mass spectroscopy (SEC-MS) for protein-protein interaction (PPI) data, and single-cell CRISPR-Cas perturbation screens by cell type. Input data streams are integrated via the Multi-Scale Integrated Cell (MuSIC) software pipeline employing deep learning models and community detection algorithms, and output cell maps are packaged with provenance graphs and rich metadata as AI-Ready datasets in RO-Crate format using the FAIRSCAPE framework. A Nature publication (Schaffer, Hu et al., April 2025) demonstrates multimodal cell maps as a foundation for structural and functional genomics, integrating IF imaging, AP-MS, and SEC-MS with integrative structure modeling to produce multimodal cell maps of MDA-MB-468 breast cancer cells and KOLF2.1J iPSCs. '
doi: 10.18130/V3/DXWOS5
page: https://www.cm4ai.org
language: en
license: CC BY-NC-SA 4.0
version: '2.1'
download_url: https://doi.org/10.18130/V3/DXWOS5
publisher: University of California San Diego
is_tabular: false
conforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema
conforms_to_schema: src/data_sheets_schema/schema/data_sheets_schema_core.yaml
conforms_to_class: CoreDataset
keywords: - Cell Maps - Artificial Intelligence - AI-Ready Data - Bridge2AI - Functional Genomics - Protein-Protein Interactions - Spatial Proteomics - CRISPR Perturbation - Hierarchical Cell Maps - MDA-MB-468 - KOLF2.1J - iPSC - Breast Cancer - Chromatin Modifiers - Metabolic Enzymes - Immunofluorescence - Mass Spectrometry - AP-MS - SEC-MS - Perturb-Seq - FAIR Principles - RO-Crate - FAIRSCAPE - Visible Neural Networks - Deep Learning - Integrative Structure Modeling
purposes:
- id: cm4ai:purpose:1
description: 'Deliver machine-readable hierarchical maps of cell architecture as AI-Ready data from
multimodal interrogation of disease-relevant cell lines to enable transformative biomedical AI research.
CM4AI produces integrated cell maps from spatial proteomics, protein-protein interactions, and genetic
perturbations using state-of-the-art mass spectrometry, cell imaging, and CRISPR technologies.
'
- id: cm4ai:purpose:2
description: 'Address the grand challenge of interpretable genotype-phenotype learning in genomics and
precision medicine. Machine learning models are often "black boxes" predicting phenotypes from genotypes
without understanding the mechanisms. CM4AI enables "visible" machine learning systems informed by
multi-scale cell and tissue architecture, allowing AI tools to interrogate how protein assemblies
in the cell affect cell-level phenotype predictions.
'
- id: cm4ai:purpose:3
description: 'Establish standards, best practices, and guidelines for ethical AI-readiness in biomedical
data. This includes implementing FAIR principles, computing machine-readable provenance graphs, characterizing
and validating all datasets with JSON-Schema mini-data-dictionaries, and mapping data elements to
public ontology vocabularies where appropriate.
'
- id: cm4ai:purpose:4
description: 'Provide multimodal cell maps as a foundation for structural and functional genomics, enabling
integrative structure modeling of protein assemblies identified via the MuSIC pipeline. Demonstrated
in peer-reviewed publication in Nature (Schaffer, Hu et al., April 2025, doi:10.1038/s41586-025-08878-3)
integrating IF imaging, AP-MS, SEC-MS, and structure modeling.
'
tasks:
- id: cm4ai:task:1
description: 'Integrate multimodal data streams (spatial proteomics via IF imaging, protein-protein
interactions via AP-MS and SEC-MS, and genetic perturbations via CRISPR screens) using the Multi-Scale
Integrated Cell (MuSIC) software pipeline employing deep learning models and community detection algorithms
to produce hierarchical cell maps.
'
- id: cm4ai:task:2
description: 'Enable development of visible neural networks (VNNs) and visible machine learning tools
that use hierarchical cell maps as interpretable structures for AI model architectures, allowing interrogation
of how protein assemblies affect cell-level phenotypes and interpretation of genetic variants and
mutations.
'
- id: cm4ai:task:3
description: 'Characterize cell architecture and protein interactions in disease-relevant cell lines
including treated and untreated MDA-MB-468 breast cancer cells (with paclitaxel and vorinostat) and
differentiated and naive KOLF2.1J induced pluripotent stem cells (iPSCs) differentiated into neurons
and cardiomyocytes.
'
- id: cm4ai:task:4
description: 'Develop ethical AI frameworks and governance structures for biomedical data, including
Value-Sensitive Design methodologies, axiological repositories, CM4AI Life Cycle framework, and guidelines
for responsible design of datasets and AI technologies.
'
- id: cm4ai:task:5
description: 'Perform integrative structure modeling of MuSIC protein communities to determine structural
models using data from PDB, AlphaFoldDB, crosslinking mass spectrometry, and prediction of disordered
sequence segments, enabling structural and functional genomics applications.
'
addressing_gaps:
- id: cm4ai:gap:1
description: 'Address the limitation that machine learning models in genomics and precision medicine
are typically difficult-to-interpret "black boxes" by providing hierarchical cell maps that enable
visible machine learning systems built directly on knowledge maps of cell and tissue architecture.
'
- id: cm4ai:gap:2
description: 'Provide fully provenanced, ethically validated, and FAIR-compliant AI-ready datasets with
machine-readable provenance graphs, complete schemas, validation procedures, and data sheets that
can be reliably processed by AI applications with full explainability.
'
- id: cm4ai:gap:3
description: 'Create integrated datasets combining protein localization (spatial proteomics), protein-protein
interactions (AP-MS and SEC-MS), and transcriptional states (CRISPR perturbation screens) at multiple
scales, enabling complex multi-modal AI analyses not feasible with single data types.
'
- id: cm4ai:gap:4
description: 'Bridge the gap between protein interaction networks and structural biology by enabling
integrative structure modeling of protein communities identified from multimodal cell maps, combining
PDB, AlphaFoldDB, crosslinking MS, and sequence disorder predictions.
'
creators:
- id: cm4ai:creator:1
description: 'Trey Ideker, Contact PI/Project Leader, University of California San Diego, Department
of Internal Medicine/Medicine, ORCID: 0000-0002-1708-8454'
- id: cm4ai:creator:2
description: 'Jean-Christophe Bélisle-Pipon, Co-Investigator, Simon Fraser University, Ethics Module
Leader, ORCID: 0000-0002-8965-8153'
- id: cm4ai:creator:3
description: 'Timothy Clark, Co-Investigator, University of Virginia, Standards Module, ORCID: 0000-0003-4060-7360'
- id: cm4ai:creator:4
description: 'Jake Yue Chen, Co-Investigator, University of Alabama at Birmingham, Teaming Module, ORCID:
0000-0002-6112-415X'
- id: cm4ai:creator:5
description: 'Nevan J Krogan, Co-Investigator, University of California San Francisco, Data Acquisition
Module (Protein-Protein Interactions), ORCID: 0000-0003-4902-337X'
- id: cm4ai:creator:6
description: 'Emma Lundberg, Co-Investigator, Stanford University, Data Acquisition Module (Spatial
Proteomics), ORCID: 0000-0001-7034-0850'
- id: cm4ai:creator:7
description: 'Prashant Mali, Co-Investigator, University of California San Diego, Data Acquisition Module
(Genetic Perturbations), ORCID: 0000-0002-3383-1287'
- id: cm4ai:creator:8
description: 'Sarah J Ratcliffe, Co-Investigator, University of Virginia, ORCID: 0000-0002-6644-8284'
- id: cm4ai:creator:9
description: 'Vardit Ravitsky, Co-Investigator, University of Montreal, Ethics Module, ORCID: 0000-0002-7080-8801'
- id: cm4ai:creator:10
description: 'Andrej Sali, Co-Investigator, University of California San Diego, Integrative Structure
Modeling, ORCID: 0000-0003-0435-6197'
- id: cm4ai:creator:11
description: 'Wade Loren Schulz, Co-Investigator, Yale University, Skills and Workforce Development
Module, ORCID: 0000-0002-2048-4028'
- id: cm4ai:creator:12
description: 'Ying Ding, Co-Investigator, University of Texas at Austin, ORCID: 0000-0003-2567-2009'
- id: cm4ai:creator:13
description: 'Samah Fodeh, Co-Investigator, Yale University, ORCID: 0000-0003-4664-3143'
- id: cm4ai:creator:14
description: 'Cynthia Brandt, Co-Investigator, Yale University, ORCID: 0000-0001-8179-1796'
- id: cm4ai:creator:15
description: 'Pamela Payne-Foster, Co-Investigator, University of Alabama, ORCID: 0000-0002-3508-3577'
- id: cm4ai:creator:16
description: 'Jillian Parker, Program Manager, University of California San Diego, Data Governance Committee
Lead, ORCID: 0000-0003-4535-3486'
- id: cm4ai:creator:17
description: 'Leah V. Schaffer, Researcher, University of California San Diego, ORCID: 0000-0001-6339-9141'
- id: cm4ai:creator:18
description: 'Mengzhou Hu, Researcher, University of California San Diego, ORCID: 0000-0002-1571-8029'
- id: cm4ai:creator:19
description: 'Christopher P Churas, Researcher, University of California San Diego, ORCID: 0000-0001-9998-705X'
- id: cm4ai:creator:20
description: 'Sadnan Al Manir, Researcher, University of Virginia, Standards Module, ORCID: 0000-0003-4647-3877'
- id: cm4ai:creator:21
description: 'Maxwell Adam Levinson, Researcher, University of Virginia, Standards Module, ORCID: 0000-0003-0384-8499'
- id: cm4ai:creator:22
description: 'Dexter Pratt, Researcher, University of California San Diego, ORCID: 0000-0002-1471-9513'
- id: cm4ai:creator:23
description: 'Sami Nourreddine, Researcher, University of California San Diego, ORCID: 0000-0003-3881-7588'
- id: cm4ai:creator:24
description: 'Amir Dailamy, Researcher, University of California San Diego, ORCID: 0000-0002-6711-8260'
funders:
- id: cm4ai:funder:1
description: 'National Institutes of Health Common Fund Bridge2AI Program. Funded through NIH grant
1OT2OD032742-01 (Bridge2AI Functional Genomics) and 5U54HG012513-02 (Bridge2AI Bridge Center), administered
by NIH Office of the Director. Opportunity Number: OTA-21-008. Project dates: September 1, 2022 to
August 31, 2026. FY 2025 funding: $5,289,382 (Direct: $4,632,095, Indirect: $657,287). Additional
funding from the Frederick Thomas Fund of the University of Virginia.
'
instances:
- id: cm4ai:instance:1
description: 'MDA-MB-468: Triple negative breast cancer cell line (RRID:CVCL_0419) established from
a metastatic site pleural effusion of a 51-year-old black female with a metastatic mammary adenocarcinoma,
available from ATCC. This cell line has been extensively used to study triple-negative breast cancer
and is well characterized with transcriptomic, mutational profile, and whole-genome sequencing data
available. Cells are analyzed under three conditions: untreated, paclitaxel-treated, and vorinostat-treated.
'
- id: cm4ai:instance:2
description: 'KOLF2.1J: Human induced pluripotent stem cell (iPSC) line (RRID:CVCL_B5P3) derived from
a healthy male Northern European donor, available from the Human Induced Pluripotent Stem Cells Initiative
(HipSci) resource. Available for access by non-for-profit organizations via a simple MTA. Analyzed
in undifferentiated state and after differentiation into neurons, neural progenitor cells (NPCs),
and cardiomyocytes.
'
- id: cm4ai:instance:3
description: '100 Chromatin Regulators: Near-comprehensive set of chromatin regulators encoded by the
human genome analyzed via AP-MS, SEC-MS, IF imaging, and CRISPR perturbation screens across different
cell states and treatment conditions. 17 genes endogenously tagged in MDA-MB-468 with AP-MS data under
three conditions, with 34 additional genes in process. SEC-MS identified 72/100 chromatin modifiers,
with 52 being integral components of protein complexes.
'
- id: cm4ai:instance:4
description: '100 Metabolic Enzymes: Set of metabolic enzymes involved in cancer, neuropsychiatric,
and cardiac disorders analyzed via multimodal interrogation including mass spectrometry, imaging,
and perturbation screens.
'
sampling_strategies:
- id: cm4ai:sampling:1
description: 'Purposive selection of two disease-relevant cell lines: MDA-MB-468 triple negative breast
cancer cell line for cancer research, and KOLF2.1J iPSCs for neuropsychiatric and cardiac disorder
research. Both cell lines ethically sourced and well-characterized in the literature. Selection criteria:
(1) MDA-MB-468 chosen for triple-negative breast cancer research applications; (2) KOLF2.1J chosen
as reference iPSC line for large-scale collaborative studies; (3) both have extensive existing characterization
data and are commercially available and ethically sourced.
'
is_sample: false
is_random: false
is_representative: false
subpopulations:
- id: cm4ai:subpop:1
description: 'MDA-MB-468 Untreated: MDA-MB-468 breast cancer cells in control/untreated condition'
- id: cm4ai:subpop:2
description: 'MDA-MB-468 Paclitaxel-Treated: MDA-MB-468 breast cancer cells treated with paclitaxel
chemotherapy'
- id: cm4ai:subpop:3
description: 'MDA-MB-468 Vorinostat-Treated: MDA-MB-468 breast cancer cells treated with vorinostat
chemotherapy'
- id: cm4ai:subpop:4
description: 'KOLF2.1J Undifferentiated iPSCs: KOLF2.1J induced pluripotent stem cells in undifferentiated/naive
state'
- id: cm4ai:subpop:5
description: 'KOLF2.1J iPSC-Derived Neurons: KOLF2.1J iPSCs differentiated into neurons'
- id: cm4ai:subpop:6
description: 'KOLF2.1J iPSC-Derived Neural Progenitor Cells (NPCs): KOLF2.1J iPSCs differentiated into
neural progenitor cells'
- id: cm4ai:subpop:7
description: 'KOLF2.1J iPSC-Derived Cardiomyocytes: KOLF2.1J iPSCs differentiated into cardiomyocytes'
collection_mechanisms:
- id: cm4ai:collection:1
description: 'Immunofluorescence Spatial Proteomics Imaging: Automated fixation and permeabilization
protocols using pipetting robot for MDA-MB-468 and KOLF2.1J cell lines. Immunofluorescence-based staining
(ICC-IF) with confocal microscopy to capture spatial subcellular organization. Completed spatial proteomics
mapping of 100 chromatin regulators in MDA-MB-468 cells under three conditions (untreated, paclitaxel,
vorinostat), with 500 additional proteins pending from genetic perturbations and PPI results. Antibodies
from Human Protein Atlas resource. Generated by Lundberg Lab at Stanford University.
'
- id: cm4ai:collection:2
description: 'Affinity Purification Mass Spectrometry: Endogenous tagging of genes in cell lines followed
by affinity purification mass spectrometry (AP-MS) to map protein-protein interactions. 17 genes endogenously
tagged in MDA-MB-468 with AP-MS data acquired under three conditions (untreated, paclitaxel, vorinostat).
34 additional genes currently in tagging process. Orthogonal approach to SEC-MS for comprehensive
PPI mapping.
'
- id: cm4ai:collection:3
description: 'Size Exclusion Chromatography Mass Spectrometry: Size exclusion chromatography coupled
to mass spectrometry (SEC-MS) for proteome-wide complex/interaction mapping. Performed in Krogan Laboratory
at UCSF. Conducted on MDA-MB-468 cells under three conditions and on KOLF2.1J iPSCs and derivatives
(undifferentiated, NPCs, neurons, cardiomyocytes). Enabled detection of over 1,000 protein complexes
in MDA-MB-468 cells and over 700 complexes in iPSCs, with thousands of proteins exhibiting differential
elution profiles between control and treated cells.
'
- id: cm4ai:collection:4
description: 'CRISPR Perturbation Screens: Single-cell CRISPR screens using CRISPR lentiviral library
targeting 100 chromatin factors with 6 guide RNAs per gene. Generated and characterized MDA-MB-468
and KOLF2.1J CRISPR lines expressing inducible dCas9. Screens performed in MDA-MB-468 cells under
3 conditions (no treatment, paclitaxel, vorinostat) and in undifferentiated KOLF2.1J iPSCs using 10x
Genomics 3''HT kit. Genome-scale screens mapping transcriptional and fitness phenotypes for 11,739
targeted genes.
'
acquisition_methods:
- id: cm4ai:acquisition:1
description: 'Confocal Microscopy for Subcellular Imaging: High-resolution confocal microscopy of immunofluorescence-stained
cells capturing four channels: DAPI (nuclei, blue), calreticulin antibody (ER, yellow), tubulin antibody
(microtubules, red), and antibody against protein of interest (green). Images processed using Human
Protein Atlas deep learning model to reduce dimensionality, producing image embeddings containing
information about protein localization.
'
- id: cm4ai:acquisition:2
description: 'Mass Spectrometry for Protein Interactions: State-of-the-art mass spectrometry-based proteomics
including AP-MS on endogenously tagged cell lines and SEC-MS for proteome-wide complex mapping. PPI
networks processed using node2vec deep learning model to reduce dimensionality, producing PPI embeddings
containing information about protein interactions.
'
- id: cm4ai:acquisition:3
description: 'Single-Cell RNA Sequencing for Perturbation Mapping: Single-cell RNA sequencing using
10x Genomics 3''HT kit to capture transcriptional states following CRISPR perturbations. Generates
genome-scale perturbation cell atlas mapping transcriptional and fitness phenotypes. Raw sequence
data deposited to NCBI BioProject/Sequence Read Archive (SRA).
'
- id: cm4ai:acquisition:4
description: 'MuSIC Pipeline Integration: Multi-Scale Integrated Cell (MuSIC) pipeline integrates PPI
embeddings and image embeddings using contrastive deep learning to obtain co-embeddings for each protein.
Community detection performed based on all-by-all similarities of protein pairs in co-embedding space,
producing hierarchical cell maps as final output. Maps annotated using Gene Ontology, Reactome pathways,
and large language model approaches.
'
preprocessing_strategies:
- id: cm4ai:preproc:1
description: 'Deep Learning Embedding Generation: PPI networks processed using node2vec deep learning
model to reduce dimensionality and produce PPI embeddings. IF images processed using Human Protein
Atlas deep learning model to reduce dimensionality and produce image embeddings. PPI and image embeddings
integrated to obtain co-embeddings using contrastive deep learning, learning co-embeddings such that
original embeddings can be reconstructed with minimal information loss.
'
- id: cm4ai:preproc:2
description: 'Hierarchical Community Detection: Community detection performed on co-embedding space
using multiscale community detection algorithms implemented in Cytoscape. Produces hierarchical directed
acyclic graphs (DAG) of protein assemblies at multiple resolutions, with 10 layers of depth representing
communities from large cell compartments to small protein complexes.
'
- id: cm4ai:preproc:3
description: 'Cell Map Annotation: Two-pronged annotation approach: (1) Alignment to known protein function
and pathway resources including Gene Ontology (GO) and Reactome to determine protein assemblies with
high overlap with known cell biology, and (2) Large language model (LLM) approach to name sets of
proteins and assign name confidence scores.
'
- id: cm4ai:preproc:4
description: 'Quality Control for Imaging Data: Standardized automated fixation and permeabilization
protocols using pipetting robot. Consistent staining protocols across conditions using Human Protein
Atlas antibodies. Quality control of imaging data before release and processing through MuSIC pipeline.
'
- id: cm4ai:preproc:5
description: 'Quality Control for Mass Spectrometry Data: Quality control and validation of AP-MS and
SEC-MS data before public release. Mass spectrometry data for human iPSCs deposited to MassIVE Repository,
and data for human cancer cells also deposited to MassIVE Repository.
'
- id: cm4ai:preproc:6
description: 'Integrative Structure Modeling: Bioinformatics pipeline for annotating MuSIC communities
by available structural information about community members and their interactions. Structural information
includes PDB, AlphaFold Protein Structure Database, crosslinking mass spectrometry, and prediction
of disordered sequence segments. Communities ranked by structural information amount as proxy for
integrative modeling feasibility. Modeling protocol scripted using Python Modeling Interface package
based on Integrative Modeling Platform (IMP) version 2.18.
'
cleaning_strategies:
- id: cm4ai:cleaning:1
description: 'FAIRSCAPE AI-Readiness Packaging: All datasets packaged using FAIRSCAPE framework which
creates RO-Crate packages with datasets, metadata, provenance graphs, and software. FAIRSCAPE-CLI
validates inputs and creates output RO-Crate packages. FAIRSCAPE server assigns persistent resolvable
globally unique identifiers (ARK scheme), decomposes RO-Crates into components, and computes end-to-end
provenance entailments using EVI Evidence Graph Ontology.
'
- id: cm4ai:cleaning:2
description: 'Data Standard Mapping: Data mapped to applicable standards and ontologies including Gene
Ontology (GO), Reactome, Protein Data Bank (PDB), AlphaFold Protein Structure Database, Schema.org,
and EVI Evidence Graph Ontology. Keywords mapped to controlled vocabularies from NCI Thesaurus, BioAssay
Ontology, Cell Ontology, CHEBI, EFO, and other ontologies.
'
intended_uses:
- id: cm4ai:use:1
description: 'AI Model Training for Functional Genomics: Primary intended use is training and development
of artificial intelligence and machine learning models for functional genomics research. AI-ready
datasets with full provenance, metadata, and validation enable immediate use in AI/ML pipelines without
reformatting.
'
- id: cm4ai:use:2
description: 'Visible Neural Network Development: Development of visible neural networks (VNNs) that
use hierarchical cell maps as interpretable model architectures. Unlike black box models, VNNs built
on cell maps allow interrogation of how protein assemblies affect cell-level phenotypes, enabling
interpretation of genetic variants and mutations in the context of cellular mechanisms.
'
- id: cm4ai:use:3
description: 'Genotype-Phenotype Mapping Research: Research into interpretable genotype-phenotype learning
using multi-scale cell maps. Enables understanding of mechanisms by which genotypes translate to phenotypes,
supporting precision medicine applications and genomic variant interpretation.
'
- id: cm4ai:use:4
description: 'Drug Response and Synergy Prediction: Analysis of cellular responses to drug treatments
(paclitaxel, vorinostat) to predict drug response and synergy. Cell maps under different treatment
conditions enable visible machine learning for drug discovery and personalized medicine applications.
'
- id: cm4ai:use:5
description: 'Disease Mechanism Research: Study of disease mechanisms in cancer, neuropsychiatric disorders,
and cardiac disorders through analysis of chromatin modifiers and metabolic enzymes in disease-relevant
cell contexts. Supports understanding of disease pathways and identification of therapeutic targets.
'
- id: cm4ai:use:6
description: 'Structural and Functional Genomics: Use as foundation for integrative structure modeling
of protein communities, combining multimodal cell maps with PDB, AlphaFoldDB, and crosslinking mass
spectrometry data to determine structural models of protein assemblies, as demonstrated in Nature
publication (Schaffer, Hu et al., April 2025).
'
- id: cm4ai:use:7
description: 'Model for AI-Ready Biomedical Dataset Development: Use as exemplar for future AI-ready
biomedical dataset development, demonstrating best practices in FAIR principles implementation, provenance
tracking, ethical data governance, and AI-readiness packaging using RO-Crate and FAIRSCAPE frameworks.
'
discouraged_uses:
- id: cm4ai:discouraged:1
description: 'Clinical Decision-Making Without Validation: Laboratory data from cell lines are not to
be used in clinical decision-making or any context involving patient care without appropriate regulatory
oversight and approval. Requires domain expertise and clinical validation before any clinical applications.
Explicitly prohibited per dataset terms.
'
- id: cm4ai:discouraged:2
description: 'Use During Incomplete Data Release: This is an interim/beta release with data not yet
in completed final form. Some datasets are under temporary pre-publication embargo, protein interrogation
sets incompletely overlap across data modalities, and computed cell maps not yet included in releases.
Full integration and final cell maps will be available in future releases through November 2026.
'
- id: cm4ai:discouraged:3
description: 'Analysis Without Domain Expertise: Datasets require domain expertise for meaningful analysis
and interpretation. Current release is most suitable for bioinformatics analysis of individual datasets.
Not suitable for use without understanding of functional genomics, proteomics, cell biology, and AI/ML
methodologies. Training resources available through CM4AI Skills and Workforce Development module.
'
license_and_use_terms:
id: cm4ai:license:1
description: 'Data licensed for reuse under Creative Commons Attribution-NonCommercial-ShareAlike 4.0
International license (https://creativecommons.org/licenses/by-nc-sa/4.0/). Attribution is required
to the copyright holders and the Cell Maps for Artificial Intelligence project. Any publications referencing
this data or derived products should cite the Nature article (Schaffer LV, Hu M, et al. Multimodal
cell maps as a foundation for structural and functional genomics. Nature. 2025. doi:10.1038/s41586-025-08878-3)
and the bioRxiv preprint (Clark T, et al. Cell Maps for Artificial Intelligence: AI-Ready Maps of
Human Cell Architecture from Disease-Relevant Cell Lines. BioRXiv, May 2024. doi:10.1101/2024.05.21.589311)
and directly cite the data collection. Commercial use requires separate license negotiation with copyright
holder (UCSD, Stanford, and/or UCSF depending upon specific data package). A Data Access Committee
(led by Jillian Parker) supervises ethical matters related to dataset distribution and potential dual
licensing for commercial use. Copyright (c) 2025 The Regents of the University of California except
where otherwise noted. Spatial proteomics raw image data is copyright (c) 2025 The Board of Trustees
of the Leland Stanford Junior University.
'
distribution_formats:
- id: cm4ai:format:1
description: 'RO-Crate Packages with Provenance: All CM4AI output data packaged as Research Object Crate
(RO-Crate) packages containing datasets, metadata, provenance graphs, and software (or resolvable
references). RO-Crates assigned persistent globally unique identifiers (ARK scheme, DOIs planned for
publishable work) that resolve to machine- and human-readable landing pages with metadata in JSON-LD
using Schema.org and EVI vocabularies. Available at https://doi.org/10.18130/V3/DXWOS5, https://doi.org/10.18130/V3/B35XWX,
https://doi.org/10.18130/V3/F3TD5R, https://doi.org/10.18130/V3/K7TGEM, and https://www.cm4ai.org.
'
- id: cm4ai:format:2
description: 'Mass Spectrometry Data in MassIVE: Mass spectrometry data deposited to MassIVE Repository
(Proteomics community-supported repository). Separate depositions for human iPSC data and human cancer
cell data (SEC-MS for KOLF2.1J iPSCs and MDA-MB-468 cancer cells). Data will be uploaded to PRIDE
when available.
'
- id: cm4ai:format:3
description: 'Sequence Data in NCBI SRA: Raw sequence data from CRISPR perturbation screens deposited
to NCBI BioProject/Sequence Read Archive (SRA). Genome-scale CRISPRi perturbation cell atlas raw sequences
and processed data available.
'
- id: cm4ai:format:4
description: 'Hierarchical Cell Maps in NDEx: Cell maps shared via Network Data Exchange (NDEx) at https://www.ndexbio.org
for visualization and access. Maps can be visualized in web browser or accessed via tools such as
Cytoscape, HiView, and Python ndex2 library.
'
- id: cm4ai:format:5
description: 'University of Virginia Dataverse: Archived RO-Crates available in University of Virginia''s
LibraData data archive (instance of Harvard''s Dataverse, an NIH-approved generalist repository).
Long-term preservation supported by committed institutional funds. Quarterly updates through November
2026. Available at https://doi.org/10.18130/V3/DXWOS5.
'
distributions:
- id: cm4ai:dist:1
description: 'Primary RO-Crate release package (Beta V2.1) containing spatial proteomics IF images,
SEC-MS protein interaction data, CRISPR perturbation screen results, and RO-Crate metadata with provenance
graphs. Packaged as ZIP archive with JSON-LD metadata (ro-crate-metadata.json).
'
format: ZIP
media_type: application/zip
path: https://doi.org/10.18130/V3/DXWOS5
- id: cm4ai:dist:2
description: 'March 2025 Beta release V1.4 (doi:10.18130/V3/B35XWX): includes perturb-seq in KOLF2.1J
iPSCs, SEC-MS in iPSCs and derivatives, and IF images in MDA-MB-468 under three conditions. Packaged
as RO-Crate (ZIP with JSON-LD metadata).
'
format: ZIP
media_type: application/zip
path: https://doi.org/10.18130/V3/B35XWX
- id: cm4ai:dist:3
description: 'June 2025 Beta release V2.1 (doi:10.18130/V3/F3TD5R): revision adds RGB IF images, RO-Crate
metadata corrections, naming convention changes, and SEC-MS for MDA-MB-468. Packaged as RO-Crate (ZIP
with JSON-LD metadata).
'
format: ZIP
media_type: application/zip
path: https://doi.org/10.18130/V3/F3TD5R
- id: cm4ai:dist:4
description: 'October 2025 Beta release (doi:10.18130/V3/K7TGEM): adds Perturb-seq for MDA-MB-468 breast
cancer cells and additional SEC-MS data. Packaged as RO-Crate (ZIP with JSON-LD metadata).
'
format: ZIP
media_type: application/zip
path: https://doi.org/10.18130/V3/K7TGEM
maintainers:
- id: cm4ai:maintainer:1
description: 'CM4AI Consortium: Multidisciplinary consortium managing dataset maintenance including
University of California San Diego (lead), University of California San Francisco, Stanford University,
University of Virginia, Yale University, University of Alabama at Birmingham, Simon Fraser University,
and The Hastings Center. Data Governance Committee led by Jillian Parker (jillianparker@health.ucsd.edu).
Ethical Review by Vardit Ravitsky (ravitskyv@thehastingscenter.org) and Jean-Christophe Belisle-Pipon
(jean-christophe_belisle-pipon@sfu.ca).
'
updates:
id: cm4ai:updates:1
description: 'Dataset regularly updated and augmented through end of project in November 2026. Beta
releases on quarterly basis with periodic data augmentation. Initial alpha release (v0.5) provided
as supplemental data. March 2025 Beta (V1.4) includes perturb-seq in KOLF2.1J iPSCs, SEC-MS in iPSCs
and derivatives, and IF images in MDA-MB-468 under three conditions. June 2025 Beta (V2.1) revision
adds RGB IF images, ro-crate metadata corrections, and naming convention changes, plus SEC-MS for
MDA-MB-468. October 2025 Beta adds Perturb-seq for MDA-MB-468 breast cancer cells and additional SEC-MS
data. Future releases will include computed cell maps and complete integration of all data streams.
Long-term preservation in University of Virginia Dataverse with committed institutional support.
'
retention_limit:
id: cm4ai:retention:1
description: 'Digital data maintained according to NIH data sharing policies with long-term preservation
in University of Virginia''s LibraData repository supported by committed institutional funds. No planned
sunset for data availability. Archived RO-Crates with persistent identifiers (ARK, future DOIs) ensure
long-term accessibility and citability.
'
human_subject_research:
id: cm4ai:hsr:1
description: 'CM4AI data are distinctive within Bridge2AI in that they are non-clinical data from tissue
cultures and are considered to be de-identified as they cannot be matched, with current knowledge,
to a human subject. Both cell lines (MDA-MB-468 and KOLF2.1J) are commercially available, ethically
sourced, de-identified cell lines. MDA-MB-468 available from ATCC. KOLF2.1J available from HipSci
resource for non-profit organizations via simple MTA. Human Subjects: No. De-identified Samples: Yes.
FDA Regulated: No.
'
sensitive_elements:
- id: cm4ai:sensitive:1
description: 'Cell Line Origin Metadata: While cell lines are de-identified and cannot be matched to
specific individuals, metadata about cell line origins (age, sex, race of original donor) is retained
for scientific context. This metadata does not constitute identifiable human subjects data under current
knowledge. Data derived from commercially available de-identified human cell lines and does not represent
all biological variants in the population at large.
'
is_deidentified:
id: cm4ai:deidentified:1
description: 'CM4AI datasets are derived from commercially available, de-identified human cell lines
(MDA-MB-468 from ATCC; KOLF2.1J from HipSci). Data cannot be matched to individual human subjects
with current knowledge. Cell line origin metadata (donor age, sex, race) is retained for scientific
context only and does not constitute identifiable human subjects data. De-identification confirmed
by CM4AI Ethics Module and Data Access Committee.
'
ip_restrictions:
id: cm4ai:ip:1
description: 'Data licensed under CC BY-NC-SA 4.0. Commercial use requires separate license negotiation
with copyright holders (UCSD, Stanford, and/or UCSF, depending on specific data package). Spatial
proteomics raw image data copyright (c) 2025 The Board of Trustees of the Leland Stanford Junior University.
Other data copyright (c) 2025 The Regents of the University of California. Data Access Committee (Jillian
Parker) supervises ethical distribution matters and potential dual licensing for commercial use.
'
regulatory_restrictions:
id: cm4ai:regulatory:1
description: 'Data are non-clinical research data from tissue cultures and are not subject to FDA regulatory
oversight. NIH data sharing policy compliance required. No HIPAA obligations (de-identified cell lines,
not from living individuals under active care). Data Access Committee oversight required for commercial
use licensing. NIH grant terms apply to funded research uses.
'
external_resources:
- id: cm4ai:resource:1
description: 'CM4AI Project Website: Official project website and data portal using U-BRITE platform.
https://www.cm4ai.org'
- id: cm4ai:resource:2
description: 'NIH RePORTER Project Details: Federal grant information and project details for Bridge2AI
Functional Genomics. https://reporter.nih.gov/project-details/11211616'
- id: cm4ai:resource:3
description: 'University of Virginia Dataverse (LibraData): Repository with archived RO-Crates and data
releases. https://doi.org/10.18130/V3/DXWOS5'
- id: cm4ai:resource:4
description: 'Nature Publication: Schaffer LV, Hu M, Qian G, et al. Multimodal cell maps as a foundation
for structural and functional genomics. Nature. Published April 9, 2025. https://doi.org/10.1038/s41586-025-08878-3
'
- id: cm4ai:resource:5
description: 'bioRxiv Preprint: Clark T, et al. Cell Maps for Artificial Intelligence: AI-Ready Maps
of Human Cell Architecture from Disease-Relevant Cell Lines. BioRXiv, May 2024. https://doi.org/10.1101/2024.05.21.589311
'
- id: cm4ai:resource:6
description: 'FAIRSCAPE Framework Documentation: AI-readiness framework documentation, tutorial, and
installation instructions. https://fairscape.github.io'
- id: cm4ai:resource:7
description: 'Integrative Modeling Platform (IMP): Open source package for integrative structure modeling.
http://integrativemodeling.org'
- id: cm4ai:resource:8
description: 'Network Data Exchange (NDEx): Repository and visualization platform for cell maps and
networks. https://www.ndexbio.org'
- id: cm4ai:resource:9
description: 'Bridge2AI Program: Parent NIH Common Fund program supporting AI-ready biomedical datasets.
https://commonfund.nih.gov/bridge2ai'
- id: cm4ai:resource:10
description: 'NIH Common Fund Data Ecosystem (CFDE): Collaboration partner for data curation and integration.
https://www.nih-cfde.org'
- id: cm4ai:resource:11
description: 'MassIVE Proteomics Repository: Mass spectrometry data repository for iPSC and cancer cell
SEC-MS data.'
- id: cm4ai:resource:12
description: 'NCBI Sequence Read Archive (SRA): Repository for CRISPR perturbation screen raw sequence
data.'
- id: cm4ai:resource:13
description: 'Perturbation Cell Atlas Publication: Nourreddine S, Doctor Y, Dailamy A, et al. A PERTURBATION
CELL ATLAS OF HUMAN INDUCED PLURIPOTENT STEM CELLS. bioRxiv. 2024 Nov 4. PMCID: PMC11580897. https://doi.org/10.1101/2024.11.03.621734
'