================================================================================
CRATE-AUGMENTED SOURCE BUNDLE (de novo fork)
================================================================================
Project: CM4AI
Document bundle: data/preprocessed/concatenated/CM4AI_preprocessed.txt
Crate package: data/ro-crate_packages/CM4AI
Crate manifest: data/ro-crate_packages/crate_manifest.yaml

This bundle is the document corpus plus RO-Crate evidence. Artifacts
that are already in D4D or datasheet form are deliberately withheld so
that this arm extracts rather than transcribes; see the exclusion list
below and notes/D4D_GENERATION_ARMS.md.

CRATE EVIDENCE INCLUDED
--------------------------------------------------------------------------------
  + CM4AI_crate_metadata_reduced.json — crate JSON-LD with file inventories collapsed; the substantive evidence (rai:* fields, ethics, access, provenance)
  + ai_ready_score.json — AI-readiness self-assessment

CRATE ARTIFACTS WITHHELD
--------------------------------------------------------------------------------
  - CM4AI_crate_d4d.yaml — already a schema-valid D4D record — this is the deterministic fork's output; including it would make the de novo arm copy-through
  - ro-crate-linkml.yaml — upstream's own D4D-shaped mapping; same copy-through risk
  - ro-crate-datasheet.html — upstream-authored datasheet rendering of the same content; a datasheet is the artifact being generated, so it is withheld as input

================================================================================
================================================================================
CONCATENATED DOCUMENT
================================================================================
Input Directory: data/preprocessed/individual/CM4AI
Total Files: 10
Extensions: ['.txt']
Recursive: False
Selection Manifest: data/preprocessed/source_manifest.yaml
================================================================================

TABLE OF CONTENTS
--------------------------------------------------------------------------------
  1. www_nature_com_articles-s41586-025-08878-3_row2.txt
  2. biorxiv_2024.05.21.589311v1_row4.txt
  3. reporter_nih_gov_project-details-11211616_row7.txt
  4. cm4ai_org_row10.txt
  5. cm4ai_org_data-releases_row11.txt
  6. creativecommons_org_licenses-by-nc-sa_row15.txt
  7. dataverse_10.18130_V3_B35XWX_row16.txt
  8. dataverse_10.18130_V3_F3TD5R_row19.txt
  9. dataverse_10.18130_V3_K7TGEM_row16.txt
 10. dataverse_10.18130_V3_HIGT4C_2026-07-24.txt
================================================================================

FILE: www_nature_com_articles-s41586-025-08878-3_row2.txt
PATH: data/preprocessed/individual/CM4AI/www_nature_com_articles-s41586-025-08878-3_row2.txt
SIZE: 131754 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: CM4AI
Source ID: nature_publication
Source type: publication
Source URL: https://www.nature.com/articles/s41586-025-08878-3
Raw file: data/raw/CM4AI/www_nature_com_articles-s41586-025-08878-3_row2.html
--------------------------------------------------------------------------------
Multimodal cell maps as a foundation for structural and functional genomics | Nature
Skip to main content
Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain
the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in
Internet Explorer). In the meantime, to ensure continued support, we are displaying the site without styles
and JavaScript.
Advertisement
View all journals
Saved research
Search
Log in
Content
Explore content
About
the journal
Publish
with us
Sign up for alerts
RSS feed
nature
articles
article
Multimodal cell maps as a foundation for structural and functional genomics
Download PDF
Download PDF
Article
Open access
Published:
09 April 2025
Multimodal cell maps as a foundation for structural and functional genomics
Leah V. Schaffer
ORCID:
orcid.org/0000-0001-6339-9141
1
na1
,
Mengzhou Hu
ORCID:
orcid.org/0000-0002-1571-8029
1
na1
,
Gege Qian
ORCID:
orcid.org/0000-0001-5068-0399
1
,
2
,
Kyung-Mee Moon
ORCID:
orcid.org/0000-0003-3796-720X
3
,
Abantika Pal
4
,
Neelesh Soni
ORCID:
orcid.org/0000-0002-9333-7491
4
,
Andrew P. Latham
ORCID:
orcid.org/0000-0002-9338-7253
4
,
Laura Pontano Vaites
5
,
Dorothy Tsai
ORCID:
orcid.org/0000-0002-1484-7961
1
,
Nicole M. Mattson
1
,
Katherine Licon
1
,
Robin Bachelder
ORCID:
orcid.org/0000-0001-7502-2907
1
,
Anthony Cesnik
ORCID:
orcid.org/0000-0002-5326-7134
6
,
Ishan Gaur
6
,
Trang Le
6
,
William Leineweber
ORCID:
orcid.org/0000-0003-3069-398X
6
,
Aji Palar
ORCID:
orcid.org/0000-0001-9357-6970
4
,
Ernst Pulido
ORCID:
orcid.org/0000-0002-3616-9743
6
,
Yue Qin
1
,
7
,
Xiaoyu Zhao
ORCID:
orcid.org/0000-0002-6081-2196
1
,
Christopher Churas
1
,
Joanna Lenkiewicz
1
,
Jing Chen
1
,
Keiichiro Ono
1
,
Dexter Pratt
ORCID:
orcid.org/0000-0002-1471-9513
1
,
Peter Zage
8
,
Ignacia Echeverria
9
,
10
,
Andrej Sali
ORCID:
orcid.org/0000-0003-0435-6197
4
,
10
,
11
,
J. Wade Harper
ORCID:
orcid.org/0000-0002-6944-7236
5
,
Steven P. Gygi
ORCID:
orcid.org/0000-0001-7626-0034
5
,
Leonard J. Foster
ORCID:
orcid.org/0000-0001-8551-4817
3
,
Edward L. Huttlin
ORCID:
orcid.org/0000-0002-1822-1173
5
,
Emma Lundberg
ORCID:
orcid.org/0000-0001-7034-0850
6
,
12
,
13
,
14
&
…
Trey Ideker
ORCID:
orcid.org/0000-0002-1708-8454
1
,
15
,
16
Show authors
Nature
volume
642
,
pages
222–231 (
2025
)
Cite this article
Save article
View saved research
63k
Accesses
41
Citations
242
Altmetric
Metrics
details
Subjects
Data integration
Machine learning
Network topology
Proteome informatics
A
Publisher Correction
to this article was published on 02 October 2025
This article has been
updated
Abstract
Human cells consist of a complex hierarchy of components, many of which remain unexplored
1
,
2
. Here we construct a global map of human subcellular architecture through joint measurement of biophysical interactions and immunofluorescence images for over 5,100 proteins in U2OS osteosarcoma cells. Self-supervised multimodal data integration resolves 275 molecular assemblies spanning the range of 10
−8
to 10
−5
m, which we validate systematically using whole-cell size-exclusion chromatography and annotate using large language models
3
. We explore key applications in structural biology, yielding structures for 111 heterodimeric complexes and an expanded Rag–Ragulator assembly. The map assigns unexpected functions to 975 proteins, including roles for C18orf21 in RNA processing and DPP9 in interferon signalling, and identifies assemblies with multiple localizations or cell type specificity. It decodes paediatric cancer genomes
4
, identifying 21 recurrently mutated assemblies and implicating 102 validated new cancer proteins. The associated Cell Visualization Portal and Mapping Toolkit provide a reference platform for structural and functional cell biology.
Similar content being viewed by others
High-plex imaging of RNA and proteins at subcellular resolution in fixed tissue by spatial molecular imaging
Article
06 October 2022
A multi-scale map of cell structure fusing protein images and interactions
Article
24 November 2021
Deep profiling of gene expression across 18 human cancers
Article
17 December 2024
Main
Human cells are organized across a spatial hierarchy of components, ranging from small protein complexes at the scale of nanometres to large condensates, compartments and organelles at the scale of micrometres
5
,
6
. One of the ultimate goals of the biological sciences is to understand this multiscale subcellular organization and its relationship to biological function and human disease. As much of cell structure still remains uncharted, there has been long-standing interest in strategies to map this architecture systematically
7
,
8
,
9
.
A variety of complementary technologies have been implemented for systematically determining subcellular organization across scales. In particular, methods such as whole-cell electron microscopy have led to maps of subcellular organelles and their placement within cells
10
,
11
. Protein immunofluorescence (IF) staining
12
and endogenous fluorescent tagging
13
, coupled to confocal microscopy imaging, have begun to reveal the subcellular locations of proteins. Biochemical proteomics approaches, such as affinity purification–mass spectrometry (AP–MS)
14
, cross-linking MS
15
, size-exclusion chromatography–MS (SEC–MS)
16
,
17
, proximity labelling
18
and isotope tagging
19
,
20
have revealed patterns of protein–protein interaction and subcellular localization that inform the makeup of protein complexes and organelles. Although these cell mapping technologies have typically been applied separately, integration of multiple complementary data modalities provides the opportunity to incorporate biological structure robustly across physical scales. Towards this aim, we recently demonstrated proof-of-concept for how two modalities—protein IF and AP–MS profiles—can be computationally fused to systematically map subcellular assemblies, with the initial version covering 661 human proteins
21
.
Here we substantially scale the cell mapping datasets and pipeline, yielding protein biophysical interactions and protein IF images for a matched set of more than 5,100 proteins in U2OS cells (Fig.
1
). Integrating these data produces a global cell biology reference map with extensive coverage of human subcellular components, including 275 distinct protein assemblies. We systematically annotate this map, assisted by recent advances in large language models (LLMs), then systematically validate its assemblies by generating a third distinct data modality—proteome-wide SEC–MS—in the same U2OS cellular context. Finally, we examine how such proteome-wide cell maps can be used to guide diverse biological studies including structural biology, protein functional annotation, analyses of cell-type specificity and multi-localization, and interpretation of the cancer genome.
Fig. 1: Study overview.
Full size image
Proteins are purified from whole-cell biochemical extracts and their biophysical interactions are determined using AP–MS. In parallel, proteins are illuminated by IF and their subcellular distributions are determined using high-resolution confocal imaging. These IF imaging and biophysical interaction data are integrated into a multimodal cell map, which is explored across five biological use cases and in an interactive visualization portal. MS and confocal microscopy illustrations are from the NIAID NIH BIOART Source (
https://bioart.niaid.nih.gov/bioart/286
;
https://bioart.niaid.nih.gov/bioart/86
).
Multimodal proteomics data acquisition
We systematically tagged proteins in U2OS osteosarcoma cells through lentiviral expression of C-terminal Flag–HA-tagged baits available in the human ORFeome library
14
. A total of 2,174 proteins were successfully tagged and isolated from U2OS whole-proteome extracts using affinity purification, and interacting partners were identified by tandem MS (AP–MS) to yield a total of 36,842 interactions among 7,543 proteins (
Methods
and Extended Data Fig.
1a,b
). Data were required to pass a panel of quality-control measures implemented as previously described
14
,
22
; these measures included sequence validation of lentiviral clones, detection of tagged bait proteins in each AP–MS run, and monitoring for sufficient numbers of protein and peptide identifications (
Methods
). Additional quality-control metrics included recovery of known complexes (Extended Data Fig.
1c,d
), for which the new interactions showed coverage comparable to previous AP–MS datasets.
To match these protein interactions with parallel information on protein subcellular locations, we amassed a large collection of confocal images of U2OS cells stained with IF antibodies against each of 10,348 proteins (20,660 images total;
Methods
). Each sample was simultaneously co-stained with reference markers for nucleus, endoplasmic reticulum and microtubules, providing a reference set of subcellular landmarks common to all images. Of these data, 17,368 images were collected in a previous publication
12
, and the remaining 3,292 images were more recently generated and validated according to the Human Protein Atlas (HPA) standard procedures for image and antibody quality control.
Combining across the interaction and imaging data, a total of 5,147 proteins was well represented in both modalities. These proteins captured approximately half of the detectable U2OS proteome
12
and provided representative coverage over the full catalogue of human protein functions, other than under-representation of transmembrane and immunoglobulin proteins (Extended Data Fig.
1e
). We found that the protein pairs measured as most similar by one modality were enriched for pairs similar in the other, showing that the biophysical interaction and imaging data share information (Extended Data Fig.
1f,g
).
Construction of a global cell map
We devised a self-supervised machine learning approach for fusing protein confocal imaging and biophysical interaction data to create a global map of protein subcellular organization (
Methods
). First, the two data streams were processed separately to generate protein features for each modality; this information was subsequently fused to create a unified multimodal embedding for each protein. Achieving a quality embedding—a low-dimensional representation extracted from complex high-dimensional data—has been a major focus of machine learning research in recent years
23
,
24
. Here we adopted a self-supervised embedding approach (Extended Data Fig.
2a
), in which proteins were positioned such that the original imaging and AP–MS features could each be reconstructed with minimal loss of information (reconstruction loss) while capturing the relative similarities and differences of each protein to others in both data modalities (contrastive loss). This multimodal embedding exhibited good performance in recovering known subcellular organization (Fig.
2a
and Extended Data Fig.
2b–d
), performing as well as, or better than, alternative supervised and unsupervised approaches (
Methods
and Extended Data Fig.
2e
).
Fig. 2: Multiscale integrated map of a U2OS cell.
Full size image
a
, Multimodal embedding of proteins based on integration of AP–MS and imaging data, reduced to two dimensions using the UMAP method
56
(left). The points are proteins that are coloured and annotated on the basis of the top-level protein communities that can be resolved. Right, enlargement of the embedding, centred on the endomembrane community and its substructure.
b
, A multiscale hierarchical view of subcellular assemblies resolved in the U2OS cell map. The nodes represent assemblies, and the edges represent containment of a smaller assembly (lower) by a larger one (upper). The node size is proportional to the estimated size in nanometres. The node colour is based on three categories of overlap with known subcellular components (defined in pie chart). The dashed boxes denote assemblies described in the text and figures.
c
, Calibrating the sizes of assemblies in the cell map (number of proteins) to the physical diameters of known structures (nanometres).
d
, GPT-4 self-confidence in generating informative names for assemblies in the cell map, shown for the categories of assemblies denoted in
b
and random assemblies (grey). The distributions of confidence scores are shown as violin plots, with the thick black lines representing the median confidence in each category. The significance of differences between distributions was calculated using one-sided Mann–Whitney
U
-tests; ****
P
< 0.0001.
Once the multimodal embedding had been learned, all pairwise protein–protein distances were computed and analysed using the multiscale community detection technique (
Methods
). Using this procedure, protein assemblies were resolved as modular communities of proteins in close proximity to one another, with such detection performed at multiple resolutions to identify protein assemblies at increasing diameters.
Application of this analytical pipeline to the data generated in U2OS osteosarcoma cells identified a hierarchy of 275 discrete protein assemblies (Fig.
2b
and Supplementary Table
1
). By calibrating the map using 13 well-known subcellular components with characterized physical sizes (for example, nucleus, mitochondria and proteasome; Supplementary Table
2
), we found that we could translate the size of an assembly (number of proteins) to an estimate of its physical diameter (in nanometres,
R
2
= 0.90) along with a prediction interval on this estimate (
Methods
). Estimated assembly diameters spanned the relevant scales of cell biology from 10
1
nm to 10
4
nm (Fig.
2c
), with assemblies robustly identified at each of these scales (
Methods
and Extended Data Fig.
3a
). By contrast, we found that maps constructed from only the imaging data tended to recover large assemblies but miss small ones, while maps constructed from only the AP–MS data recovered small assemblies but tended to miss large ones (Extended Data Fig.
3b,c
). Overall, the integrated map identified the largest number of assemblies, including 104 that were not resolved by either individual modality (Extended Data Fig.
3d
and Supplementary Table
1
).
Annotation of the U2OS cell map
To study and annotate the U2OS cell map, we held a series of in-person Annotation Jamborees, during which approximately a dozen individuals worked in pairs to assign names and putative functional roles to assemblies on the basis of expert knowledge and literature curation. First we examined the correspondence of assemblies to known subcellular components documented in the Comprehensive Resource of Mammalian protein complexes (CORUM)
25
, Gene Ontology (GO)
26
or HPA
12
(
Methods
; Jaccard index ≥ 10%). We found that 41 assemblies closely reconstructed a known component (Jaccard index ≥ 50%) while 90 had moderate agreement, with some unexpected differences (20% ≤ Jaccard index < 50%).
The remaining 144 assemblies were designated as not previously documented assemblies. In these cases, team members worked collaboratively to consider the current biological literature relevant to the assembly’s protein subunits and their potential functions. This process was greatly informed by suggestions from OpenAI’s pre-trained transformer (GPT-4)
27
, a generative LLM that we recently showed is capable of providing insightful names and functional interpretations for gene sets identified in omics data
3
. As in this previous study, we used an engineered prompt and pipeline (
Methods
and Extended Data Fig.
4a,b
) to guide the LLM to generate descriptive names for gene sets indicative of their biological roles, along with a fully referenced analysis essay providing its rationale (Extended Data Fig.
4c
) and a self-assessment of confidence in the suggested name. When applied to the U2OS cell map, we found that the LLM assigned names to known assemblies with very high confidence (median of 0.92 for both high overlap and substantial variation; Fig.
2d
) and to the previously undocumented assemblies with moderately high confidence (median, 0.85), contrasting starkly with its confidence for sets of proteins drawn randomly without any correspondence to biological structure (median, 0.0). For 104 out of the 144 not previously documented assemblies, the literature about the various proteins was sufficiently coherent for GPT-4 to propose a confident assembly name (confidence ≥ 0.85), each of which was subsequently passed to the human curation team for final naming determination (Supplementary Table
1
).
We noted that the highest level of organization in the cell map covers previously documented organelles and large subcellular compartments of >100 proteins, including the nucleus with 102 nuclear subassemblies, the mitochondrion with 16 mitochondrial subassemblies, 127 assemblies inside the cytosol and 3 assemblies related to microtubules (Fig.
2b
). Organized within the nucleus are subcomponents such as nucleoli and the nucleoplasm, which itself hierarchically resolves 67 components including the Mediator and RNA polymerase complexes and an array of other transcriptional machines. Notably, components of the plasma membrane and cytosolic periphery, such as G-protein and clathrin-coated-pit complexes, are tightly associated with numerous other cytosolic proteins under a single large compartment, which we simply labelled ‘cytosol’ (Fig.
2b
). Major expected components of the cytosolic compartment, such as the endoplasmic reticulum and Golgi apparatus, are also resolved. We found 48 assemblies that are potential biomolecular condensates
28
on the basis of their enrichment for proteins with intrinsically disordered regions, proteins predicted to phase separate or proteins recorded in the CD-Code condensate database (
Methods
and Supplementary Table
3
). Of these, 39 had a significant overlap with a recent complementary effort to predict protein condensates through integration of diverse biochemical protein features
29
(hypergeometric test FDR < 5%), while the remaining nine putative condensates had not been previously identified (Supplementary Table
3
).
Systematic validation by SEC–MS
We next sought to systematically validate the cell map components using whole-cell SEC–MS as an orthogonal approach. Using this technique, cellular extracts from a cell population of interest are separated by SEC, followed by identification of proteins in each size fraction by tandem MS (Fig.
3a
). Here we subjected triplicate cultures of U2OS cells to SEC–MS of 40 separate chromatography fractions, yielding quantitative fractionation profiles for 5,509 proteins in at least two replicates, of which 3,020 were present in the cell map. Quality assessment of the SEC–MS dataset showed that elution profiles were largely reproducible across replicate biological measurements (Extended Data Fig.
5a,b
), with protein peaks present across the full range of fractions (Extended Data Fig.
5c
).
Fig. 3: Global analysis of assemblies with SEC–MS.
Full size image
a
, Overview of the SEC–MS experiment.
b
, The elution fractions (columns) for proteins (rows) in representative cell map assemblies. The intensity (colour) represents the relative protein abundance in each fraction, scaled in the range [0–1] for each protein across all fractions. Previously undocumented assemblies are indicated in purple font.
c
, The distribution of SEC–MS Pearson correlations among pairs of proteins in assemblies, shown separately for small assemblies, medium assemblies and random pairs. Significant differences from random pairs are indicated, determined using one-sided Wilcoxon rank-sum tests; ****
P
< 0.0001.
d
, Cell map with assemblies coloured on the basis of the significance of validation by SEC. The assembly colour indicates the FDR as determined using one-sided Wilcoxon rank-sum tests comparing SEC co-elution profiles for protein pairs in that assembly versus protein pairs that do not co-occur in any assembly under the root.
Integration of these measurements with the multiscale cell map revealed significant agreement, with proteins in the same assembly (as identified earlier by AP–MS and imaging) having a strong tendency to co-elute in the same chromatography size fractions (Fig.
3b,c
). Overall, SEC data validated 89 assemblies (5% false-discovery rate (FDR)), corresponding to 43% of assemblies (76 out of 175) with more than 5 proteins and 61% of assemblies (59 out of 96) with more than 15 proteins (Fig.
3d
,
Methods
and Supplementary Table
4
). Among small-to-medium size assemblies of <50 proteins, we found 39 for which the SEC data had specifically corroborated the inclusion of unexpected members (
Methods
and Supplementary Table
4
), with functions related to heat shock, stress response and vesicle trafficking.
At this stage of the study, we had interrogated U2OS cells with multimodal proteomics data; integrated these data to resolve subcellular components at multiple scales; annotated these components; and lent support to many using an independent whole-cell profiling technique. We next turned our attention from map construction to use, exploring key impacts in structural and functional biology (use cases 1–5: three-dimensional (3D) structural modelling; revealing protein function; studying cell type specificity; protein multi-localization; and interpreting tumour mutations).
3D structural modelling
We first explored the cell map as a platform to guide 3D structural modelling projects, interfacing with the recent advances in structure prediction enabled by artificial intelligence (AI)
30
. We used AlphaFold-Multimer
31
to predict structural models for every pair of proteins arising in the same focal protein assembly (142 assemblies of <10 proteins, 1,666 protein pairs in total; Supplementary Table
5
). We noted that the estimated accuracies of these structures (AlphaFold pTM and ipTM scores;
Methods
) were significantly higher than expected at random, supporting that these protein pairs have direct biophysical interaction interfaces (one-sided Mann–Whitney
U
-test,
P
= 2.7 × 10
−12
). Particularly high structural accuracy was indicated for 161 pairs, which also received highly confident per-residue scores at the protein–protein interaction interface (Fig.
4a
and
Methods
).
Fig. 4: Use of the cell map in driving studies of subcellular structure and function.
Full size image
a
, The results from AlphaFold-Multimer folding of heterodimeric protein complexes in the U2OS cell map.
b
,
c
, SEC–MS plot (
b
) and the corresponding structure (
c
) for the DPYSL2–DYSL3 heterodimer.
d
,
e
, SEC–MS plot (
d
) and the structure (
e
) for TARS3 and EPRS1.
f
,
g
, SEC–MS plot (
f
) and the structure (
g
) for ERH and CCDC9B, excluding disordered regions.
h
, IF images for representative members of the Rag–Ragulator complex. Members are immunostained (green) with cytoskeleton counterstain (red). Scale bar, 2 µm.
i
, Biophysical interaction data for the Rag–Ragulator complex.
j
, Integrative structure model of the Rag–Ragulator complex. The structural ensembles of ITPA and BORCS6 are presented as 3D localization probability densities, with surfaces transparent for visual clarity.
k
, Biophysical interaction data for representative RNase MRP complex members.
l
, IF images for four RNase MRP proteins, immunostained (green) and with cytoskeleton counterstain (red). Scale bar, 5 µm.
m
, Differential expression (z score, colour bar) after CRISPR knockdown of genes encoding the RNase MRP complex (top rows, green) versus a random sampling of other proteins. The rows represent CRISPR knockdowns, and the columns represent genes with the 20 most variable differential expression patterns across the full dataset.
Of these high-confidence structures, 111 had not been previously documented in the Protein Data Bank (PDB). An example was a biophysical assembly identified among DPYSL2, DPYSL3 and DPYSL4, a family of phosphoproteins important for nervous system development
32
. Their initial association was validated by SEC–MS co-elution profiling (Fig.
4b
), after which AlphaFold-Multimer yielded high-confidence structures for all pairwise interactions of these proteins (Fig.
4c
). Additional complexes that were validated first by SEC–MS, then resolved structurally by AlphaFold-Multimer, included an interaction between TARS3, a threonyl-tRNA synthetase, and EPRS1, a member of the aminoacyl-tRNA synthetase multienzyme subsystem
33
(Fig.
4d,e
); another example was a structure involving ERH and CCDC9B (Fig.
4f,g
).
We also examined how AI predictions can be integrated with experimental structural data to create a 3D model of a large protein assembly. We selected the Rag–Ragulator complex, which is located on the lysosomal membrane where it regulates growth signalling through the activation of the mammalian target of rapamycin complex 1 (mTORC1)
34
. The assembly that we had resolved in the cell map (Fig.
4h,i
) included members of the recombination-activating genes (RAG) and Ragulator protein families (LAMTOR1–5, RRAGA, RRAGC, SLC38A9) as well as two unexpected proteins, BORCS6 and ITPA. We built an integrative structural model
35
of this Rag–Ragulator assembly (
Methods
), incorporating and expanding on the base cryo-EM structure
36
(PDB:
6WJ2
), AlphaFold single structure predictions of BORCS6 and ITPA, as well as pairwise AlphaFold-Multimer predictions of BORCS6 or ITPA interactions with each of the other members of the Rag–Ragulator complex. The integrated structure (Fig.
4j
) indicated that BORCS6 interacts with LAMTOR2 and is proximal to LAMTOR1, LAMTOR3 and LAMTOR5. Similarly, the model supported the interaction of ITPA with LAMTOR1, LAMTOR3 and LAMTOR4. These examples illustrate how a data-driven compendium of subcellular components can identify new target protein components for downstream 3D structural studies.
Revealing protein function
Notably, 138 proteins of previously unknown function
37
were present in the cell map, of which 24 fell in small-to-medium size assemblies of fewer than 25 proteins. Most of these assemblies had been assigned robust biological names during map curation (see above), enabling us to propose functions for their uncharacterized proteins through guilt by association (Supplementary Table
6
). One such functional assignment was for C18orf21, which our cell mapping data placed robustly in the RNase mitochondrial RNA processing (MRP) complex (Fig.
4k,l
). Corroborating this assignment, we observed that knockdown of
C18orf21
induces a distinct transcriptional cell state very similar to knockdowns of other MRP genes (Fig.
4m
).
Expanding to proteins with some previous functional annotation, we found 951 cases in which a protein was assigned to an unexpected assembly of fewer than 25 proteins, suggesting new functional roles (Supplementary Table
6
). For example, the interferon-stimulated gene factor 3 (ISGF3) complex
38
, previously defined as consisting of STAT1, STAT2 and IRF9, also included dipeptidyl peptidase 9 (DPP9), a serine protease previously associated with inflammation
39
. Our AP–MS data implicated DPP9 as a potential member of this complex based on the STAT2 pull-down (Extended Data Fig.
6a
) and this association was reinforced by the confocal images, which indicated similar cytosolic patterns of localization with ISGF3 proteins (Extended Data Fig.
6b
). We observed that inhibition of DPP9 by 1G244 (a selective DPP9 inhibitor
40
) upregulated the canonical ISG targets of STAT transcription factors, including IFNβ1, IFNγ1 and IFNγ2, while a non-ISG control was unaffected (
Methods
and Extended Data Fig.
6c
), suggesting that DPP9 acts to suppress the IFN response (Extended Data Fig.
6d
). These examples illustrate how a data-derived reference cell map provides a substantial aid in completing the functional annotation of the human proteome.
Studying cell type specificity
Defining a global map of a given cell type confers the potential to distinguish subcellular components that are specific to that type from those that are more widely conserved. As an initial proof of concept towards this aim, we examined each protein assembly in the U2OS cell map for evidence of shared versus distinct biophysical interaction patterns in comparison to HEK293 human embryonic kidney cells (previously characterized by AP–MS in the BioPlex 3.0 resource
22
;
Methods
and Extended Data Fig.
7a
). Of the 258 assemblies with AP–MS data coverage in both cell types, we identified 103 that were conserved across cell types (Extended Data Fig.
7b
and Supplementary Table
7
). These included large assemblies, including the nucleus and cytosol, as well as small assemblies such as the spliceosome, the 9–1–1 RAD–RFC complex (Extended Data Fig.
7c,d
) and components of the SNARE complex. The remaining 155 assemblies showed biophysical interaction patterns that were significantly different between HEK293 and U2OS cell types. For example, a cytosolic component named the energy metabolism regulation complex was robustly identified in the U2OS AP–MS data, but none of the corresponding interactions were detected in HEK293 cells (Extended Data Fig.
7e
). These examples illustrate how a data-driven cell map can elucidate protein assemblies that are specific or shared between cell types, providing a basis to explain different cell phenotypes and identify cell-type-specific drug targets.
Protein multi-localization
A substantial fraction of proteins have been postulated to multi-localize, that is, to have a role in multiple subcellular assemblies or compartments
12
,
41
. To this point, we noted that approximately 30% of proteins in the cell map (1,520 out of 5,147) are present in more than one distinct assembly (Extended Data Fig.
8a
and Supplementary Table
8
). For example, XAB2, a known factor of the spliceosome and transcription-coupled repair
42
, localized not only to nuclear assemblies as expected, but also to the endomembrane (Extended Data Fig.
8b
). Evidence for such localizations was present in the fluorescence images as well as in the AP–MS interaction network, in which XAB2 showed strong interactions with both nuclear spliceosomal and membrane-associated stress factors (Extended Data Fig.
8c
).
Moving beyond single proteins, we also investigated whether there was evidence of multiple localizations for entire protein assemblies, noting 23 that were indeed documented to multi-localize according to the U2OS cell map (Extended Data Fig.
8d,e
). For example, the amyloid precursor protein (APP) complex (APP, APBA2, APBA3, APLP2, TJAP1) was clearly resolved in both the cytosol and endomembrane compartments (Extended Data Fig.
8d
) on the basis of evidence from both the protein imaging and biophysical interaction modalities (Extended Data Fig.
8f,g
). This finding aligns with previous studies showing that APP and its homologue, APLP2, have a role in subcellular trafficking from the endoplasmic reticulum to the cell surface
43
(with vesicular and endoplasmic reticulum localizations captured in our U2OS imaging data; Extended Data Fig.
8g
). APBA2 and APBA3 are members of the X11 adaptor protein family, which is known to regulate the translocation of APP
44
. These examples illustrate how a multimodal cell map can reveal both single proteins and whole assemblies that localize to multiple subcellular compartments, suggesting pleiotropic functions.
Interpreting tumour mutations
Determining how diverse genetic alterations disrupt common molecular machines is critical to understanding the complexity of diseases such as cancer. Towards this aim, we obtained genome-wide somatic mutation profiles for a compendium of 772 paediatric primary tumours encompassing 18 tumour types
4
(Supplementary Table
9
). We then analysed these mutational profiles using the U2OS cell map, looking for mutational selection on the set of genes of an assembly as a whole (
Methods
). Each assembly was tested for mutation within each tumour type separately and across the entire pan-cancer cohort. While individual gene mutations are rare in paediatric cancer, with only 6 genes altered in >2% of tumours (Fig.
5a
), we identified a total of 11 recurrently mutated assemblies at this same 2% threshold (Fig.
5b
). For example, the
SMARCA4
SWI–SNF transcriptional activator is a well-known cancer driver that is genetically altered in 2.5% of paediatric tumours
45
(Fig.
5a
), but this frequency increases to 6.0% when including coding alterations across all 13 proteins in SWI–SNF complexes (Fig.
5b
). Some recurrently mutated assemblies were highly specific to certain cancer types, as was the case for an unexpected finding of frequent mutations of cell junctions in B cell lymphoblastic lymphoma (Fig.
5c
). Other assemblies appeared to be under mutational selection more generally across tumours, as in the case of the nuclear pore (Fig.
5c
). Cumulative across subtypes, this analysis identified a total of 21 assemblies that were recurrently mutated, suggesting positive selective pressure during tumour evolution (Fig.
5d,e
and Supplementary Table
10
). Mutated assemblies were identified at all size scales but had a clear preference for small complexes of fewer than 50 proteins (Fig.
5e
).
Fig. 5: Protein assemblies as convergence points for paediatric cancer mutations.
Full size image
a
, The mutation frequencies of the top 550 proteins (
x
axis), quantified in the pan-paediatric cancer cohort (
y
axis,
n
= 772 tumours). Non-silent point mutations or insertion/deletions are included. Proteins with magenta bars were previously reported as being under significant mutational pressure
4
.
b
, The mutation frequencies of 98 cancer protein assemblies (
x
axis), quantified in the same pan-paediatric cancer cohort (
y
axis). The magenta bars highlight assemblies under significant mutational pressure (FDR ≤ 0.4, Methods). Inset (top right): expansion of one of these assemblies (SWI–SNF complex) by the protein-level mutation frequencies of its members (grey bars). QC, quality control; reg., regulation.
c
, The mutation frequencies (colour gradient) of assemblies (rows) within paediatric tumour types (columns). Pink gradient is used for recurrently mutated assemblies detected in the pan-cancer analysis. Navy gradient is used for recurrently mutated assemblies detected in individual tumour cohorts. MBL, medulloblastoma; HGG, high-grade glioma; ATRT, atypical teratoid/rhabdoid tumour; NHL, non-Hodgkin lymphoma; AML, acute myeloid leukaemias; WT, Wilms’ tumours; RBL, retinoblastoma; OS, osteosarcoma; ES, Ewing’s sarcoma; BLL, B cell lymphoblastic leukaemia/lymphoma; RMS, rhabdomyosarcoma; NBL, neuroblastoma; PAST, pilocytic astrocytoma; EPM, ependymoma.
d
, Cell map indicating assemblies that are under mutational pressure across the pan-paediatric patient cohort (magenta,
n
= 14) or in individual tumour cohorts (navy,
n
= 7). The assembly indicated by a dashed rectangle is further discussed in Extended Data Fig.
9
.
e
, The distribution of sizes for the recurrently mutated assemblies.
Within these assemblies, we focused on 250 putative cancer proteins, defined as proteins that are not only present in recurrently mutated assemblies but are also themselves mutated in multiple tumour samples (
Methods
). To further investigate a role for these proteins in cancer, we performed a large meta-analysis of transposon-based mutagenesis screens in mouse tumour models
46
(
Methods
and Extended Data Fig.
9a
). The putative cancer proteins showed a very high degree of enrichment for genes in which transposon mutagenesis leads to tumour development (Extended Data Fig.
9b)
, with specific validation support for 102 proteins (FDR < 0.3). The majority of these proteins had not been implicated in previous gene-level mutational analysis of either adult or paediatric cancer (Extended Data Fig.
9c,d
and Supplementary Table
10
). For example, the significantly mutated NCOR-associated transcriptional regulation assembly (Extended Data Fig.
9e
) contained a total of 28 proteins, of which 16 were impacted by paediatric cancer mutations (Supplementary Table
10
). Two proteins in this complex, NCOR1 and TBL1XR1, had been previously reported as cancer driver genes and shown to regulate key signalling pathways in modulating tumour growth
47
,
48
. Of others in this complex, we found that three validate as cancer drivers through mouse transposon mutagenesis (GTFIRD1, NRIP1, NCOR2). We also noted that proteins in this complex show a high proclivity to phase separate (22 out of 28;
Methods
and Supplementary Table
3
) with distinct punctae in the IF images, suggestive of nuclear condensate formation (Extended Data Fig.
9f
). These findings demonstrate how knowledge of cancer protein assemblies can focus a genome analysis to increase the sensitivity of detecting cancer mutational events.
Cell map toolkit and portal
To enable interactive exploration of the human cell map, we developed the companion Multiscale Integrated Cell visualization portal (available at
http://musicmaps.ai/u2os-cellmap/
), which combines a high-performance graphical web interface with the general analysis functionality of the widely used Cytoscape application
49
. The map is browsable as a tree view (that is, the hierarchy in Fig.
2b
) or a cell view, in which hierarchical assembly relationships are represented as nested circles (Extended Data Fig.
10
). Tables provide key information such as the proteins comprising each assembly, estimated assembly sizes in nanometres and links to confocal images. Each assembly can be selected to display its supporting subnetwork of evidence, including biophysical interactions (denoting proteins with high subcellular proximity as revealed by AP–MS pull-downs) and imaging interactions (denoting proteins with high subcellular proximity as revealed by the confocal images). Built-in search functionality is used to select and highlight assemblies that contain proteins of interest, and the platform also integrates LLM functional interpretation (Extended Data Fig.
4
) to allow assemblies to be explored for insightful names and functional interpretations
3
. To facilitate continued map improvement, incorporation of new datasets, and construction of new cell maps across subtypes and disease states, we also developed the Cell Mapping Toolkit (
https://github.com/idekerlab/cellmaps_pipeline
), which implements the end-to-end pipeline described here as a series of Python packages complete with full user documentation. This toolkit provides a flexible and generalizable framework for cell map construction, enabling researchers to integrate and construct cell maps via multiple input modalities.
Discussion
Although the basic sequence of the human genome has been known for over two decades
50
, knowledge of how its proteins are organized within cells is still very much evolving. To advance this cause, we have developed a reference human cell map with extensive coverage of subcellular assemblies spanning four orders of magnitude (around 10
−8
to 10
−5
m). Achieving coverage across proteins and scales relied on at least two advances: interrogating the cell with matched proteome-wide datasets tuned to complementary types of information, and integrating these views systematically through a multimodal deep learning workflow. These advances provide a blueprint for mapping subcellular architecture that can be readily applied across human cell types and disease states. They also pave the way to expanded cell maps incorporating new modalities, such as proximity labelling, subcellular fractionation or cryo-electron tomography, as well as time-dependent measurements, such as monitoring of subcellular dynamics over a progression of cell cycle phases.
With such generality in mind, we surveyed a series of use cases representing common areas of investigation in which a global data-driven cell map can powerfully drive biological discovery. First, we examined how protein assemblies provide the starting material for 3D structural modelling, leading to the generation of high-confidence heterodimeric structures using AlphaFold (Fig.
4a
and Supplementary Table
5
) and a large integrative model of the Rag–Ragulator complex combining computational predictions with experimental 3D coordinates (Fig.
4h–j
). A second key impact was in the study of individual proteins, in which the cell map suggests unexpected roles for numerous proteins (Supplementary Table
6
). As a proof of concept, we further investigated a role for C18orf21 in the RNase MRP complex (Fig.
4k–m
) and for DPP9 in the ISGF3 complex (Extended Data Fig.
6
). Other key applications were in the study of cell type specificity (Supplementary Table
7
and Extended Data Fig.
7
), molecular condensates (Supplementary Table
3
) and multi-localizing proteins and protein assemblies (Supplementary Table
8
and Extended Data Fig.
8
). A final, critical demonstration was in decoding human genetics. By identifying patterns of genetic mutations that converge on protein assemblies (Supplementary Table
10
), numerous proteins were implicated that had not been previously reported as paediatric cancer drivers (Extended Data Fig.
9c,d
).
Through multimodal analysis, the human cell map presented here unifies and extends multiple ongoing efforts that have thus far progressed independently. In this respect, we found that the integration of multiple modes of data substantially broadens the sensitivity and robustness with which subcellular components can be resolved across scales (Extended Data Fig.
3
). These benefits translate to real impacts in biological discovery as exhibited in the use cases. Approximately half of AlphaFold structures (47 out of 111; Supplementary Table
5
) and 40% of new protein functional annotations (Supplementary Table
6
) were driven by assemblies that were robustly identified only by integrating both AP–MS and imaging datasets.
A separate distinct benefit of a multimodal analysis is that, by design, it provides multiple lines of evidence for new biological findings. In a typical omics study, a single modality of data is presented and analysed with many putative findings, only a few of which can be validated or pursued at any depth. By contrast, each new finding of the U2OS cell map is derived from two complementary experimental platforms by default (AP–MS biochemical pull-downs and spatial proteomics imaging), and the systematic lines of evidence deepen further in the use cases through support from SEC–MS, AlphaFold predictions, perturb-seq and/or transposon mutagenesis. For example, the assembly of multifunctional protein ERH with RNA-binding protein CCDC9B was supported by an AP–MS interaction, image subcellular annotations, SEC–MS elution profiles (Fig.
4f
) and a high-confidence AlphaFold 3D model (Fig.
4g
). Such confluence of data, also seen in other recent multi-omic studies
51
,
52
, increases the confidence in each result and provides substantial additional structural, functional and/or spatial information. This aspect pushes towards a new mode of end-to-end cell biology whereby multiple datasets are generated, integrated and simultaneously corroborated, informing a unified and foundational representation of the cell
9
,
53
,
54
,
55
.
Methods
AP–MS data collection
U2OS cell cultures were processed for protein–protein physical interaction mapping by AP–MS, according to a previously described protocol developed as part of the BioPlex project
14
. U2OS cells were obtained from American Type Culture Collection (ATCC) and tested for
Mycoplasma
contamination. C-terminal HA-Flag-tagged DNA constructs targeting each of 2,174 bait proteins were constructed using clones from the human ORFeome library
57
and introduced into U2OS cells by lentiviral transfection. Baits were selected based on success in previous pull-down experiments and to ensure broad sampling of the interactome as observed previously
22
. Immobilized and pre-washed mouse monoclonal anti-HA agarose resin was incubated with cell lysates to extract protein baits and their associated protein complexes. Subsequently, these were eluted with HA peptide then reduced and digested with trypsin. Approximately 1 µg of peptide was loaded for reversed-phase liquid chromatography with a C18 microcapillary column followed by tandem MS (Thermo Fisher Q-Exactive HFX) using data-dependent acquisition selecting the top 20 precursors for MS2 analysis. Proteins were identified from the MS2 spectra using Sequest
58
, filtered to 1% protein-level FDR with additional entropy-based filtering
14
. The CompPASS algorithm
59
,
60
was used to select high-confidence (top 2%) protein–protein interactions on the basis of the abundance of proteins in each immunoprecipitation compared with their average levels across all other immunoprecipitations. Interactions were further filtered with CompPASS-Plus at a 1% FDR
14
,
61
. Steps for quality control were as follows. Clones were sequence-validated as described previously
57
. AP–MS analyses required the bait protein to be detected in the Sequest results; moreover, bait proteins were required to have a higher abundance (based on spectral counting) in their own pull-down compared with the other pull-downs on the same 96-well plate. To remove under-loaded samples, we required LC–MS runs to contain a minimum of around 5,000 PSMs and about 700 proteins. Enrichment of interactions within CORUM complexes (Extended Data Fig.
1c,d
; CORUM v.4.1) was computed using a one-sided binomial test, assuming background probability of interaction equal to the network’s interaction density, with Benjamini–Hochberg (BH) FDR correction. CORUM complexes for each case were limited to those with at least three proteins and at least one AP–MS bait in the network. Randomized networks were constructed preserving the overall number of interactions per bait (node degrees).
Matched protein IF imaging data
U2OS cell cultures were analysed using IF confocal imaging as part of the Human Protein Atlas project (HPA) using a previously described protocol
12
. U2OS cells were obtained from ATCC and were authenticated according to the manufacturer using morphology, karyotyping and PCR-based approaches to confirm the identity and to exclude intraspecies and interspecies contaminations. U2OS cells were seeded in 96-well glass-bottom plates and grown to a confluence of 60 to 70% at 37 °C in McCoy 5A medium, supplemented with 10% fetal bovine serum (FBS) and 5% CO
2
for propagation. Cells were then fixed in 4% paraformaldehyde followed by permeabilization with Triton X-100 detergent and incubated with the HPA primary antibody for the target protein, overnight at 4 °C. HPA antibodies were diluted to 2–4 μg ml
−1
in blocking buffer with 1 μg ml
−1
mouse anti-tubulin and 1 μg ml
−1
chicken anti-calreticulin. The next day, cells were incubated at 90 min at room temperature with secondary antibodies (goat anti-rabbit AlexaFluor 488; goat anti-mouse and goat anti-chicken AlexaFluor 647; or goat anti-rat AlexaFluor 647) diluted to 1 μg ml
−1
and counterstained with 4′,6-diamidino-2-phenylindole (DAPI). IF images were acquired using a Leica SP5 confocal microscope equipped with a ×63 HCX PL APO 1.40 oil CS objective. Each IF image contains four colour channels, one for the protein of interest and the other three channels for reference markers corresponding to nucleus (DAPI), microtubule (anti-tubulin antibody) and endoplasmic reticulum (anti-calreticulin antibody). Antibody quality was scored according to a standard HPA protocol (
https://www.proteinatlas.org/about/antibody+validation
); the highest scoring antibody per protein was selected with up to two technical replicate images.
SEC–MS data collection
We collected a proteomic SEC–MS dataset in the U2OS cell line according to a previously described procedure
62
. U2OS cells were tested for
Mycoplasma
contamination. Three 15 cm dishes of confluent U2OS cells for each replicate (
n
= 3) were washed and collected in ice-cold SEC buffer (50 mM KCl, 50 mM NaCH
3
COO, 50 mM Tris, pH 7.2, containing 1× EDTA-free HALT protease and Thermo Fisher Scientific phosphatase inhibitor cocktail). These samples were subjected to a fractionation protocol described previously
63
, with modifications. In brief, cells were lysed using a Dounce homogenizer with a tight pestle for 3.5 min on ice. Lysates were ultracentrifuged at 100,000 rcf for 15 min at 4 °C, and the supernatants were concentrated over 100 kDa molecular mass cut-off spin columns (Sartorius). A standard Bradford assay was performed to inject 600 µg of protein for each replicate into a single 300 × 7.8 mm BioSep-4000 column (Phenomenex) using SEC buffer without protease inhibitors. The samples were then separated into 40 fractions at 15 s per fraction using the 1290 Series semi-preparative HPLC (Agilent Technologies) system at a flow rate of 0.6 ml min
−1
at 6 °C. The collection end point was predetermined by measuring the end of the BSA standard peak, discarding anything smaller than a single BSA protein size. The resulting fraction volumes of protein were denatured by adding to a final concentration 20% (v/v) 2,2,2-trifluoroethanol (Sigma-Aldrich), reduced and alkylated
64
. Subsequently, we added an equal volume of 50 mM ammonium bicarbonate for overnight digestion with trypsin (New England Biolabs) at 37 °C. The resulting peptides were cleaned with C-18 STop And Go Extraction (STAGE) tips
65
using 40% (v/v) acetonitrile and 0.1% (v/v) formic acid in water as the elution buffer. Peptide concentrations were measured on a NanoDrop One instrument (Thermo Fisher Scientific, 205 nm, Scopes method), after which we loaded approximately 50 ng of peptides onto the TimsTOF Pro2 (Bruker Daltonics) system with CaptiveSpray source coupled to a nanoElute UHPLC (Bruker Daltonics) device using an Aurora Series Gen2 analytical column (25 cm × 75 μm, 1.6 μm FSC C18; Ion Opticks). The instrument was set to acquire in DIA-PASEF mode as previously outlined
66
. The sample batch was randomized before injection. Acquired SEC–MS data were searched on DIA-NN (v.1.8.1.0)
67
against the UniProt human sequences (
UP000005640
, downloaded 2 June 2023) and common contaminant sequences (229 sequences). Library-free search was enabled, using trypsin/P protease specificity and 1 missed cleavages. Other search parameters included 1 maximum number of variable modifications, N-terminal M excision, carbamidomethylation of C and oxidation of M. Peptide length ranged from 7 to 30, precursor charge ranged from 1–4, precursor
m
/
z
ranged from 300 to 1,800, and fragment ion
m
/
z
ranged from 200 to 1,800. Precursor FDR was set to 1%, with 0 for settings ‘mass accuracy’, ‘MS1 accuracy’ and ‘scan window’. The settings ‘heuristic protein inference’, ‘use isotopologues’, ‘match between run (MBR)’ and ‘no shared spectra’ were all enabled. ‘Protein name from FASTA’ was chosen for the protein inference parameter along with ‘double-pass mode’ for neural network classifier. Robust LC (high precision) was used for the quantification strategy, RT-dependent mode for cross-run normalization, and smart profiling mode for library generation. Analyses of SEC–MS data used the protein elution profiles, defined as the protein-level quantification values reported by DIA-NN across all fractions. The similarity was calculated between the elution profiles for every pair of proteins, taking the mean Pearson correlation across the three replicates. For assessment of reproducibility across biological measurements (Extended Data Fig.
5b
), we first selected the set of proteins present in all three replicates (
n
= 5,018). For each replicate, we determined each protein’s elution pattern, defined as the set of Pearson correlations between that protein and every other of the 5,018 proteins. We then calculated the Pearson correlation of protein elution patterns across replicates for the same protein or, alternatively, between random pairs of proteins.
AP–MS and IF data preprocessing
Proteins were first pre-processed within the AP–MS and IF modalities separately. For the AP–MS data, the node2vec
68
Python3 implementation (
https://github.com/eliorc/node2vec
) was used to represent each protein
i
as a 1,024-dimension feature vector (
x
i
) based on its protein–protein interaction neighbourhood (
p
= 2,
q
= 1, walk length = 80, number of walks = 10). For the IF data, we applied DenseNet-121, a convolutional neural network pre-trained for object recognition in protein IF confocal images
69
. DenseNet-121 was used to represent each protein as a 1,024-dimension feature vector (
y
i
) from the four channels of the colour image.
Multimodal embedding overview
We developed a self-supervised multimodal machine learning model to integrate (co-embed) the AP–MS and IF protein representations into a single low-dimensional (128-dimension) embedding space (Extended Data Fig.
2a
). Our model is based on the autoencoder architecture known as multimodal structured embedding
70
with modifications. Parameters of the autoencoder are trained using a two-component loss function that combines reconstruction loss and triplet (contrastive) loss. Details are provided in the ‘Encoder/decoder architecture’, ‘Loss functions’ and ‘Model training’ sections below.
Encoder/decoder architecture
The separate AP–MS and IF vector inputs (
x
i
and
y
i
for each protein
i
, see above) are compressed by modality-specific encoders (
f
x
and
f
y
) yielding 128-dimension vectors
a
and
b
:
$$\begin{array}{l}{{\bf{a}}}_{i}\,=\,{f}_{x}\,({{\bf{x}}}_{i})\\ \,=\,{\rm{Tanh}}({\rm{BatchNorm}}({\rm{Linear}}({\rm{Dropout}}({\rm{ELU}}({\rm{BatchNorm}}\\ \,({\rm{Linear}}({\rm{Dropout}}({{\bf{x}}}_{i}))))))))\end{array}$$
$$\begin{array}{l}{{\bf{b}}}_{i}\,=\,{f}_{y}\,({{\bf{y}}}_{i})\\ \,=\,{\rm{Tanh}}({\rm{BatchNorm}}({\rm{Linear}}({\rm{Dropout}}({\rm{ELU}}({\rm{BatchNorm}}\\ \,({\rm{Linear}}({\rm{Dropout}}({{\bf{y}}}_{i}))))))))\end{array}$$
where Dropout indicates dropout layers
71
; Linear indicates linear transformation layers; BatchNorm indicates batch normalization
72
; Tanh indicates a hyperbolic tangent function; and ELU indicates an exponential linear unit function. The
a
and
b
vectors are then input to a joint encoder
f
z
that learns the L2-normalized 128-dimension latent representation
z
i
:
$${{\bf{z}}}_{i}={f}_{z}\,[{\rm{concat}}({{\bf{a}}}_{i},{{\bf{b}}}_{i})]={\rm{L}}2{\rm{Norm}}({\rm{BatchNorm}}({\rm{Linear}}({\rm{Dropout}}({\rm{concat}}({{\bf{a}}}_{i},{{\bf{b}}}_{i})))))$$
Values of
z
i
constitute the self-supervised multimodal embedding used for subsequent cell map evaluation (see the ‘Evaluation of embedding approaches’ section below) and construction (see the ‘Pan-resolution community detection’ section below). For the decoder step,
z
is reverse-transformed to extract 128-dimension modality-specific features through weight matrices
w
x
and
w
y
:
$${{\bf{c}}}_{i}={w}_{x}{{\bf{z}}}_{i}$$
$${{\bf{d}}}_{i}={w}_{y}{{\bf{z}}}_{i}$$
Finally, these features are passed to modality-specific decoders (
g
x
and
g
y
), yielding the 1,024-dimension reconstructed inputs (
\(\hat{{\bf{x}}}\)
i
,
ŷ
i
):
$${\widehat{{\bf{x}}}}_{i}={g}_{x}({{\bf{c}}}_{i})={\rm{Linear}}({\rm{Tanh}}({\rm{Linear}}({\rm{ELU}}({\rm{Linear}}({{\bf{c}}}_{i})))))$$
$${\widehat{{\bf{y}}}}_{i}={g}_{y}({{\bf{d}}}_{i})={\rm{Linear}}({\rm{Tanh}}({\rm{Linear}}({\rm{ELU}}({\rm{Linear}}({{\bf{d}}}_{i})))))$$
Loss functions
To compute the reconstruction loss
R
, the (
\(\hat{{\bf{x}}}\)
i
,
ŷ
i
) outputs of the autoencoder are compared to the original input values (
x
i
,
y
i
) for each modality:
$${R}_{x}=\frac{1}{n}\mathop{\sum }\limits_{i=1}^{n}{||{{\bf{x}}}_{i}-\hat{{{\bf{x}}}_{i}}||}_{2}$$
$${R}_{y}=\frac{1}{n}\mathop{\sum }\limits_{i=1}^{n}{||{{\bf{y}}}_{i}-\hat{{{\bf{y}}}_{i}}||}_{2}$$
where
n
is the total number of proteins. The overall reconstruction loss is the sum of modality-specific reconstruction losses and a regularization term, where
λ
regularization
is the regularization weight and ||w||
F
is the
F
-norm of the matrix:
$$R={R}_{x}+{R}_{y}+{\lambda }_{{\rm{regularization}}}({\parallel {w}_{x}\parallel }_{F}+{\parallel {w}_{y}\parallel }_{F})$$
To compute triplet loss
T
, clustering using the Louvain algorithm
73
is performed on the (
a
,
b
) vectors of each modality (during early training clusters are defined using input (
x
,
y
) values instead; see the ‘Model training’ section below). This clustering defines selection functions
S
x
and
S
y
for each modality, with
S
(
i
,
j
) = 1 for proteins
i
,
j
in the same cluster, else 0. This information is used to compute
T
for each modality:
$${T}_{x}=\frac{1}{m}\sum _{i\varepsilon N}\sum _{j\varepsilon N\,,j\ne i}\sum _{k\varepsilon N\,,k\ne i\,,j}{S}_{x}(i\,,j)(1\,-\,{S}_{x}(i,k))\times {\rm{\text{max}}}(D({{\bf{z}}}_{i},{{\bf{z}}}_{j})-D({{\bf{z}}}_{i},{{\bf{z}}}_{k})+\varepsilon ,0)$$
$${T}_{y}=\frac{1}{m}\sum _{i\varepsilon N}\sum _{j\varepsilon N\,,j\ne i}\sum _{k\varepsilon N\,,k\ne i,j}{S}_{y}(i\,,j)(1\,-\,{S}_{y}(i,k))\times {\rm{\text{max}}}(D({{\bf{z}}}_{i},{{\bf{z}}}_{j})-D({{\bf{z}}}_{i},{{\bf{z}}}_{k})+\varepsilon ,0)$$
where
N
is the set of all proteins,
D
denotes the cosine distance (1 – cosine similarity), and
m
is the total number of terms inside the summation that are greater than 0. The full loss function
L
is a weighted sum of the reconstruction and triplet losses:
$$L=R+{\lambda }_{{\rm{triplet}}}({T}_{x}+{T}_{y})$$
Model training
Model parameters were trained with standard neural network learning procedures provided by Pytorch
74
v.2.0.1, based on backpropagation using the Adam stochastic gradient descent method
75
. Training occurred in three phases: (1) Over the first 200 epochs, only the reconstruction loss
R
was used for backpropagation. (2) Over an additional 200 epochs, the full loss function
L
was used for backpropagation, with
S
x
and
S
y
defined using input
x
,
y
vectors. (3) Over a final 500 epochs of training, the full loss function
L
was used for backpropagation, with
S
x
and
S
y
defined using
a
,
b
vectors (updated every 200 epochs). Values of hyperparameters were set based on previous work
70
without fine-tuning: batch size = 64,
λ
regularization
= 5,
λ
triplet
= 5, Adam optimization learning rate = 0.0001. Triplet loss margin and dropout percentages (
ε
= 0.10, dropout = 0.25) were set based on commonly recommended values
76
,
77
.
Evaluation of embedding approaches
The above self-supervised embedding model was evaluated in comparison to two alternative multimodal embedding approaches: (1) simple unsupervised concatenation of the separate AP–MS and IF inputs (
x
,
y
); and (2) a random forest regression model supervised to use (
x
,
y
) to predict protein–protein semantic similarities from the Gene Ontology (June 2023 release), trained as previously described
21
(Python Scikit-learn package, fivefold cross-validation, n_estimators=1000, max_depth=30). These embedding models were each scored for their recovery of interacting protein pairs documented in three complementary reference databases: (1) high-confidence protein–protein interactions in STRING
78
,
79
(v.12, NDEx uuid 0b04e9eb-8e60-11ee-8a13-005056ae23aa; Extended Data Fig.
2b,e
); (2) protein pairs assigned to the same CORUM
25
complex (v.4.1, NDEx uuid 764f7471-9b79-11ed-9a1f-005056ae23aa; Extended Data Fig.
2c,e
); or (3) protein pairs with high functional similarity in a genome-wide CRISPR-perturbation/mRNA sequencing screen (perturb-seq
80
; Extended Data Fig.
2d,e
). Here, high functional similarity was defined as the top 1% of protein pairs by Pearson correlation between the profiles of mRNA transcriptional changes induced by CRISPR disruptions of the two proteins (see the ‘Analysis of perturb-seq data’ section below).
Pan-resolution community detection
The cosine similaritiy between the multimodal embeddings for each pair of proteins was used to generate a series of protein–protein proximity networks in which edges were defined from the most similar 0.2, 0.3, 0.4, 0.5, 1.0, 2.0, 3.0, 4.0, 5.0 or 10.0% pairs, respectively, yielding 10 networks in total. Pan-resolution community detection was performed in each of these networks using the Hierarchical community Decoding Framework (HiDeF;
https://github.com/fanzheng10/HiDeF
)
81
, with a persistence threshold (
k
) of 10 and a maximum resolution (maxres) of 80, with other parameters kept at the default settings. HiDeF identifies protein communities at different resolutions and represents their hierarchical relationships as a directed acyclic graph (DAG). In this DAG, the nodes represent communities and the directed edges (
a
→
b
) represent that community
a
contains community
b
. The DAG was refined by assigning parent–child containment relationships between assemblies with containment index ≥ 75% and removing redundant systems with Jaccard index ≥ 90% with parent systems. This final DAG defines the cell map referenced in Fig.
2b
.
Estimation of assembly diameter
A subset of 13 protein assemblies was selected from the cell map corresponding to assemblies with a known physical diameter documented in the literature (Supplementary Table
2
). Linear regression was used to fit the log
10
-transformed diameter (nm,
y
) against the log
10
-transformed size of the assembly (number of proteins,
x
):
y
= 1.27
x
− 0.31. This linear equation was then used to estimate a diameter
ŷ
for each assembly in the map. A 95% prediction interval (PI) was estimated on the basis of the standard error as follows:
$${\log }_{10}{\rm{PI}}=\widehat{y}\pm ({t}_{(1-\alpha /2,n-2)}\times {\rm{s.e.}}(\widehat{y}))$$
with
t
determined by the Student’s
t
-distribution (
t
= 2.2 with d.f. =
n
− 2,
n
= 13 components). The s.e. is the standard error between predicted and measured sizes, calculated as follows:
$${\rm{s.e.}}(\widehat{y})={s}_{e}\sqrt{1+\frac{1}{n}+\frac{{(x-\bar{x})}^{2}}{\mathop{\sum }\limits_{i=1}^{n}{({x}_{i}-\bar{x})}^{2}}}$$
where,
\({s}_{e}=\sqrt{\frac{{\sum }_{i=1}^{n}{({y}_{i}-\hat{y})}^{2}}{n-2}}\)
. Relevant to Fig.
2c
.
Evaluation of assembly robustness
The robustness of protein assemblies was evaluated using a statistical jackknifing approach, as described previously
21
. A random set of 10% of proteins was removed before multimodal embedding (see the ‘Multimodal embedding overview’ section above); integration and community detection were then performed using the same parameters described in the ‘Model training’ and ‘Pan-resolution community detection’ sections. This randomization procedure was repeated 300 times to create a set of jackknifed hierarchies. The robustness of each assembly from the original hierarchy was then calculated as the fraction of all jackknifed hierarchies that contained at least one matching assembly, defined as substantial and significant overlap between the protein sets representing the target and the match (Jaccard index ≥ 40% and hypergeometric statistic FDR < 0.001). To assess the dependence of each assembly on the protein imaging data, we created a dataset with AP–MS features randomized (1,024-dimension random vectors sampled from a normal distribution) before the statistical jackknifing procedure, and the robustness of each assembly was computed as described above. For assessing the dependence of each assembly in the map on the AP–MS data, a reciprocal procedure was performed in which image embeddings were randomized. Relevant to Extended Data Fig.
3
.
Annotation of cell map assemblies
The cell map was annotated by first aligning assemblies with the GO cellular component branch (June 2023 release), CORUM (4.1 human complexes) or HPA (v.23) resources. Each of these cell biology resources defines a list of protein sets (GO terms, CORUM complex, HPA subcellular localizations), referred to here as components. Hypergeometric tests were performed for each assembly versus each component in the resource, and the FDR was determined using BH correction. The results were tabulated for all assembly–component pairs with Jaccard index ≥ 10% and hypergeometric statistic FDR < 0.01 (Supplementary Table
1
). Assemblies in the map were labelled as high overlap with known assembly (Jaccard index ≥ 50% for at least one of the three resources); substantial variation on known assembly (Jaccard index < 50% for all three resources and 20% ≤ Jaccard index < 50% for at least one of the resources); or not previously documented assembly (Jaccard index < 20% for all three resources) based on this enrichment analysis. We also used our recently developed Gene Set AI (GSAI) pipeline
3
to guide the GPT-4 model
27
(v.gpt-4-1106-preview) to annotate assemblies with <1,000 proteins (Extended Data Fig.
4a
). This approach uses a well-engineered prompt that follows the chain-of-thought
82
and one-shot
83
strategies to query GPT-4 for a descriptive name, a confidence score and a detailed reasoning assay of the protein members from each assembly. One example is shown in Extended Data Fig.
4c
, and the full result for each assembly is available in Supplementary Table
1
. Literature references are provided by a separate GPT-4 based citation module developed in the previous study
3
(Extended Data Fig.
4b
) to aid in interpretability. The citation model extracts gene symbols and functional keywords from each paragraph of the LLM-generated analysis text; these are used to construct and execute PubMed queries that search titles and abstracts. The returned publications are prioritized based on relevance and the number of matching genes in their abstracts. Finally, a separate GPT-4 instance is asked to evaluate whether the top three publication titles and abstracts provide supporting evidence for factual statements in the original analysis paragraph, selecting those that satisfy this requirement as references. To evaluate the reproducibility of GPT-4 naming (Extended Data Fig.
4d
), we performed the GSAI pipeline for five additional replicate runs of GPT-4 and calculated the semantic similarity between the assembly names generated in each of these runs versus the original run. Similarity was computed using the SapBERT model
84
from huggingface (cambridgeltl/SapBERT-from-PubMedBERT-fulltext) using the transformers package
85
(v.4.29.2). Assemblies that were not named by the original run were eliminated from the reproducibility test.
Biological condensate analysis
To analyse the cell map for biological condensates, we used three resources: IUPred3.0
86
, a sequence-based predictor of protein disorder; FuzDrop
87
, a sequence-based predictor for the ability of a protein to drive condensate formation; and CD-Code
88
, a database containing proteins known to participate in biological condensates. IUPred3.0 predicts the probability of each amino acid in a sequence as being disordered. Proteins containing a contiguous sequence of amino acids >30 residues, where each amino acid has a >50% chance of being disordered, were annotated as likely disordered. FuzDrop assigns a probability of a sequence driving phase separation, which we thresholded at >60% to annotate a protein as ‘likely phase-separated’. Finally, we searched for each gene’s UniProtID in CD-Code (accessed 31 May 2023) under ‘
Homo sapiens
’, enabling us to annotate a protein as ‘associated with known condensates’. We used a hypergeometric test to assign statistical significance (
P
< 0.01) to each protein assembly that was enriched in proteins that were likely disordered, likely phase-separated, or associated with known condensates. Assemblies that were significant in one of these three analyses were considered possible biological condensates (Supplementary Table
3
).
Validation of protein assemblies and subunits by SEC–MS data
For the set of proteins in each assembly, we determined the Pearson correlation in SEC–MS elution profiles for all pairs of these proteins (see the ‘SEC–MS data collection’ section). This similarity distribution was then compared to a null distribution (all pairs of proteins not in any common U2OS assembly, that is, assigned to root node only) using a one-sided Wilcoxon rank-sum test with BH correction (Fig.
3d
and Supplementary Table
4
). Assemblies with FDR < 5% were considered validated. A similar analysis was performed using PrinCE
89
(
https://github.com/fosterlab/PrInCE
) scores to rank protein pairs rather than Pearson correlations, with PrinCE run using the default parameters. We found that 90 assemblies were validated at 5% FDR in the complementary analysis using PrInCE, including 70 assemblies validated by both Pearson correlation and PrinCE similarity measures (Supplementary Table
4
). For validation of unexpected protein subunits within assemblies, for each assembly <50 proteins, ‘unexpected proteins’ were defined as those not included in the best matching cellular component from any of three cell biology resources (GO, CORUM, HPA; see the ‘Annotation of cell map assemblies’ section above). For each unexpected member, its SEC–MS elution profile was compared against all other proteins in the assembly using Pearson correlation; this similarity distribution was compared to the null distribution as described above to compute an FDR. Unexpected proteins with FDR < 5% were considered validated (Supplementary Table
4
).
AlphaFold-Multimer analysis
All pairs of proteins in small assemblies (<10 proteins) were selected for AlphaFold-Multimer analysis. AlphaFold-Multimer was run on each pair using localcolabfold (
https://github.com/YoshitakaMo/localcolabfold
) with the default settings
90
. Sequences were acquired from the complete human protein UniProt FASTA file (
UP000005640
, reviewed sequences, downloaded 11 September 2023). For each predicted heterodimeric structure, we calculated a weighted average between the predicted template modelling score (PTM, an estimate of the similarity between the predicted and ground truth structures) and the ipTM score (the pTM score modified to score the interfaces across different proteins)
31
:
$${\rm{model\; score}}=0.8\times {\rm{ipTM}}+0.2\times {\rm{pTM}}$$
We calculated the median score out of five independent models generated per protein pair. A null score distribution was generated by repeating this score computation for pairs of proteins drawn randomly from those pairs that were not part of the same small assembly (<10 proteins as above). This null distribution was used to calculate an FDR for actual protein pair scores, selecting a cut-off of 30% corresponding to a weighted PTM score of 0.39. Pairs were further evaluated for the presence of a confident interface residue (within 10 Å of the other protein and plDDT score > 80). Relevant to Fig.
4a
.
Integrative structure modelling of the Rag–Ragulator complex
A structural model of the Rag–Ragulator community was computed by using an integrative modelling approach
35
,
91
,
92
,
93
, proceeding through the standard four stages
35
,
91
,
94
as follows. (1) Gathering input information: the Rag–Ragulator model in the cell map included LAMTOR1 through LAMTOR5, RRAGA, RRAGC, SLC38A9, BORCS6, NUDT3 and ITPA. An integrative model was computed based on the SLC38A9–RagA–RagC–Ragulator comparative model (PDB:
6WJ2
template)
36
, AlphaFold
30
predictions for BORCS6 and ITPA, and pairwise AlphaFold-Multimer predictions
31
for BORCS6 or ITPA versus all other members of the Rag–Ragulator complex. One-hundred AlphaFold-Multimer models were generated for each pair and evaluated using FoldDock
95
. The model excluded NUDT3 because AlphaFold-Multimer did not produce high-confidence models of NUDT3 and other Rag–Ragulator components according to FoldDock. (2) Representing subunits and translating data into spatial restraints: the components of the Rag–Ragulator community were represented as rigid bodies. Alternative models were ranked through a scoring function corresponding to a sum of terms, each one of which restrains some aspect of the model based on a subset of input information. The spatial restraints included a binary binding mode restraint on the position and orientation of pairs of proteins as derived from ensembles of AlphaFold-Multimer predictions, connectivity restraints between consecutive pairs of beads in a subunit and excluded volume restraints between non-bonded pairs of beads. (3) Configurational sampling to produce an ensemble of structures that satisfies the restraints: the initial positions and orientations of rigid bodies and flexible beads were randomized. The generation of structural models was performed using replica exchange Gibbs sampling, based on the Metropolis Monte Carlo algorithm
96
. Each Monte Carlo step consisted of a series of random translations of flexible beads and random translations and rotations of rigid bodies. (4) Analysing and validating the data and ensemble structures: model validation
93
,
97
included selection of the models for validation; estimation of sampling precision; estimation of model precision; and quantification of the degree to which a model satisfies the information used to compute it. The above four-step modelling protocol was scripted using the Python Modelling Interface (PMI) package, a library for modelling macromolecular complexes based on the open-source Integrative Modelling Platform (IMP) package v.2.18 (
https://integrativemodeling.org
)
91
. The configuration of the rigid Rag–Ragulator complex, ITPA protein and the two BORCS6 domains was computed by minimizing the violations of the spatial restraints implied by the input information, using IMP
91
. Relevant to Fig.
4j
.
Analysis of perturb-seq data
The K562 day-8 perturb-seq dataset
80
was acquired at
https://gwps.wi.mit.edu
(BioProject:
PRJNA831566
). This dataset provides single-cell transcriptional profiles for 9,867 distinct gene knockouts, which underwent filtering based on the following criteria: (1) gene knockout corresponds to a protein in our U2OS cell map; (2) gene knockout has efficient on-target mRNA reduction of >30%; (3) gene knockout induces a strong transcriptional phenotype defined by ≥20 differentially expressed genes at a significance of
P
< 0.05 on the basis of the Anderson–Darling test followed by BH correction. This filtering process resulted in a list of 1,289 gene knockouts. The functional cell states due to each of these perturbations were represented using the mean-normalized differential expression profile. Relevant to Fig.
4m
and Extended Data Fig.
2d,e
.
Analysis of DPP9 inhibition
U2OS cells were seeded in triplicate at 300,000 cells per well in a six-well plate (two biological replicates). The next day, cells were treated with 1G244, a DPP9 inhibitor (HY-116304, MedChem Express) at the indicated concentrations for a total of 6 h. After treatment, The medium was aspirated and washed once with ice-cold PBS. Cells were collected in 500 µl of cold TRIzol reagent (15596026, Invitrogen) using a cell scraper. 100 µl of chloroform was added to the TRIzol lysate and vortexed for 20 s followed by a 3 min incubation at room temperature. The homogenate was centrifuged at 10,000
g
for 18 min at 4 °C. A total of 200 µl of aqueous phase was removed with a pipette and transferred to a new Eppendorf tube. An equal volume of 100% ethanol was slowly added to the aqueous phase and mixed by gentle pipetting. The entire sample was transferred to an RNeasy Mini spin column placed in a 2 ml collection tube (74104, Qiagen). The rest of the extraction was carried out according to the Qiagen RNeasy protocol. 2 µg of RNA per sample was reverse-transcribed according to the iScript cDNA Synthesis Kit protocol (1708890, Bio-Rad, interferon beta 1: Hs01077958_s1; interferon gamma 1, Hs00194264_m1; interferon gamma 2, Hs00988304_m1; non-ISG—18S, 4333760T; and
GAPDH
, Hs0275889q_g1). qPCR was carried out in triplicates in a 96-well plate according to the TaqMan Fast Advanced Master Mix protocol (4444557, Thermo Fisher Scientific) on a CFX96 Touch Real-Time PCR Detection System from Bio-Rad. The expression levels were compared against a housekeeping gene (
GAPDH
), and the relative expression levels were compared against the DMSO control. Relevant to Extended Data Fig.
6
.
Analysis of conservation of U2OS assemblies in a second cell type
We downloaded the AP–MS BioPlex v3 network from NDEx (uuid 6b995fc9-2379-11ea-bb65-0ac135e8bacf), which provides high coverage of human protein interactions in a second cell type, HEK293 cells (14,033 proteins, 127,732 protein–protein interactions). Node2vec was used to represent the interaction pattern of each protein in this HEK293 network (see the ‘AP–MS and IF data preprocessing’ section). The cosine similarity in interaction patterns was then computed for all protein pairs (separately for HEK293 and U2OS). For the set of proteins included in each U2OS assembly, the distribution of pairwise protein similarities in HEK293 were compared to those in U2OS cells using the two-sided Mann–Whitney
U
-test. This test was translated to an effect size using Cliff’s delta
98
; assemblies with Cliff’s delta ≥ 0.5 were considered to be increasingly U2OS-specific whereas those with Cliff’s delta < 0.5 were considered to be increasingly conserved. Relevant to Extended Data Fig.
7
; in Extended Data Fig.
7b
, Cliff’s delta scores of <0 are set to 0.
Multi-localization analysis
For each protein, we identified its terminal locations in the cell map hierarchy, defined as assemblies (hierarchy nodes) where the protein appeared but was absent in all subassemblies (child nodes). We then counted the number of unique paths from these terminal locations to the root of the hierarchy (root node). Proteins with multiple distinct paths to the root were classified as multi-localized, indicating their presence in different branches of the cell map. Multi-localized assemblies were identified as assemblies with more than one parent node in the hierarchy. Relevant to Extended Data Fig.
8
.
Pre-processing of paediatric cancer mutational profiles
Data were obtained from a pan-paediatric cancer study
4
of 914 individual patients with cancer aged under 25 years (study ID: pediatric_dkfz_2017, downloaded from cBioPortal
99
,
100
). We selected the following types of non-silent somatic mutation events: ‘Frame_Shift_Del’, ‘Frame_Shift_Ins’, ‘In_Frame_Del’, ‘In_Frame_Ins’, ‘Missense_Mutation’, ‘Nonsense_Mutation’, ‘Nonstop_Mutation’, ‘RNA’, ‘Splice_Region’, ‘Splice_Site’ and ‘Translation_Start_Site’. A total of 772 primary tumour samples, spanning 18 cancer types, were in the resulting list (Supplementary Table
9
). We recorded the number of tumours in the pan-paediatric cohort, as well as each individual tumour cohort, in which each gene was observed to have at least one somatic mutation event (
N
(
g
,obs)
). Moreover, we calculated the expected number of mutations for each gene in the pan-paediatric cohort (
N
(
g
,exp)
) using the default setting of MutSigCV v.1.4, as described in a previous study
101
. For expected mutation counts for individual cancer cohorts, we down-scaled the pan-paediatric cancer
N
(
g
,exp)
based on the proportion of patients (for example, 44 patients with Wilms’ tumours (WT) account for 5.7% of the pan-paediatric cohort, so
N
g
,exp,WT
= 0.057 ×
N
g
,exp,pan-paediatric
). Finally, the corrected log mutation count of each gene (
M
g
) for each cohort was calculated as:
$${M}_{g}={\log }_{2}(\max ({N}_{(g,{\rm{obs}})-}{N}_{(g,\exp )},0)+1)$$
Statistical identification of recurrently mutated assemblies
We applied a previously described statistical model, HiSig
101
(
https://github.com/fanzheng10/HiSig
), to calculate the mutation selection pressure on assemblies with the default parameter settings. HiSig implements linear regression (with L1 lasso regularization) of the mutation count against the organization of proteins in assemblies. We calculated an empirical
P
value by comparing the mutational selection on assemblies against 10,000 randomly permuted assignments of proteins to assemblies. The FDR was calculated by BH correction. Recurrently mutated assemblies were selected on the basis of FDR ≤ 0.4. Assembly-level mutation frequencies were calculated from the number of distinct patients who carried at least one mutated protein in the assembly. Tumour types with fewer than 15 patients were excluded from analysis, as were mutated assemblies with >50 mutated proteins.
Validation of cancer driver genes
Genes mutated in more than one patient with cancer and located in the significantly recurrent mutated assemblies (see above) were defined as putative cancer proteins. We obtained a large collection of transposon-based mutagenesis screens in mice from the Candidate Cancer Gene Database (CCGD)
46
(
http://ccgd-starrlab.oit.umn.edu/index.html
, downloaded on 26 March 2024). This database consists of a total of 72 studies with mouse transposon insertion mutagenesis screens across 13 tumour categories (Extended Data Fig.
9a
). We determined the number of studies in which a gene was disrupted by transposon insertion in mice tumours. Mutated genes in cancer assemblies were designated positives (genes expected to have high study counts because they are mutated), and all other genes were designated negatives (genes not expected to have high study counts). We calculated the kernel density estimation (KDE) for the mutated genes in cancer assemblies and other genes in the cell map using the stat.gaussian_kde function from the Python package scipy (v1.7.3). The area under the KDE curves was integrated using the trapz function from Python package numpy (v.1.21.6). The FDR was then computed as the ratio of the area under the curve for false positives (Area
FP
) to the total area under the KDE curve representing both false positives and true positives (Area
TP
), mathematically shown as:
\({\rm{FDR}}=\frac{{{\rm{Area}}}_{{\rm{FP}}}}{{{\rm{Area}}}_{{\rm{FP}}}+{{\rm{Area}}}_{{\rm{TP}}}}\)
. We specified the minimum number of screens reporting a gene at 4 (
x
≥ 4), corresponding to FDR = 0.28, as the threshold cut-off for validated cancer drivers (Extended Data Fig.
9b
). Adult cancer driver genes were collected from the TCGA Pan-Cancer Atlas
102
; significantly mutated genes in the pan-paediatric cancer cohort were collected from refs.
4
,
103
. These genes were defined as known cancer genes in Extended Data Fig.
9c,d
.
Running the cell mapping toolkit
The Cell Mapping Toolkit (
https://github.com/idekerlab/cellmaps_pipeline
) implements a series of Python packages to execute the end-to-end pipeline described herein. Specific packages include steps for processing the protein imaging and biophysical interaction datasets (cellmaps_imagedownloader, cellmaps_ppidownloader), embedding the input modalities (cellmaps_image_embedding, cellmaps_ppi_embedding), integrating the modalities (cellmaps_coembedding), constructing the hierarchical cell map (cellmaps_generate_hierarchy) and annotating the cell map with known resources such as GO (cellmaps_hierarchyeval). Each package is pip-installable and is linked to complete user documentation hosted at ReadTheDocs (
https://cellmaps-pipeline.readthedocs.io/
). A step-by-step guide is provided at the GitHub repository.
Statistics and reproducibility
Statistical tests were performed using SciPy
104
with BH multiple-testing correction where appropriate. Statistics involving comparison between two data distributions were calculated using Mann–Whitney U-tests or Wilcoxon rank-sum tests (Figs.
2d
,
3c,d
and Extended Data Figs.
2b–d
,
7b
,
9b
). Statistics for assessing the enrichment of proteins or protein pairs were calculated using hypergeometric tests (Fig.
2b
and Extended Data Fig.
3
) unless stated otherwise. The SEC–MS data were reproduced in three biological replicates. The IF stainings were reproduced in at least two different cell lines in HPA (Fig.
4h,l
and Extended Data Figs.
6b
,
7d
,
8c,g
,
9f
). The qPCR experiment for DPP9 was repeated for two biological replicates and three technical replicates each (Extended Data Fig.
6c
).
Reporting summary
Further information on research design is available in the
Nature Portfolio Reporting Summary
linked to this article.
Data availability
The Multiscale Integrated Cell web portal (
musicmaps.ai/u2os-cellmap
) provides links to all major data and derived resources associated with this study, including AP–MS protein interactions, protein IF images, SEC data and the online interactive U2OS cell map. The U2OS cell map is available at
https://ndexbio.org
under uuid f693137a-d2d7-11ef-8e41-005056ae3c32. Protein assemblies in the cell map are also available at the European Bioinformatics Institute (EBI) Protein Complex Portal (
https://www.ebi.ac.uk/complexportal
) with the query CLO:0009454. The AP–MS protein interaction data are available at
https://ndexbio.org
under uuid 95bc75d5-d1d1-11ee-8a40-005056ae23aa. In addition to its release here, the U2OS protein interaction network will be included as part of the upcoming BioPlex
105
v.4.0 database release (E.L.H. et al., manuscript in preparation). AP–MS raw MS files are available at MassIVE under the identifier
MSV000097168
. The entire image dataset is included in the Human Protein Atlas v23 release. SEC–MS raw MS files and search results are available at the Proteome Xchange under the identifier
PXD052362
. All structural models are available at the ModelArchive Database (
https://modelarchive.org
) with the identifiers ma-idk-u2osmap and ma-m5og4. Other public databases and resources used in this study include Gene Ontology (June 2023 release;
https://geneontology.org
), CORUM (v.4.1 release,
https://mips.helmholtz-muenchen.de/corum/
), UniProt
Homo sapiens
proteome (accessed 2 June and 11 September 2023;
https://uniprot.org
), STRING interactome (v.12; NDEx uuid: 0b04e9eb-8e60-11ee-8a13-005056ae23aa), OpenCell interactions (
https://opencell.czbiohub.org/download
), CD-CODE condensate database (accessed 31 May 2023;
https://cd-code.org
), FuzDrop (dataset S7 in ref.
87
), Protein Condensate Atlas (supplementary dataset 8 in ref.
29
), K562 day-8 perturb-seq dataset (
https://gwps.wi.mit.edu
), HEK-293 BioPlex v.3.0 (NDEx uuid: 6b995fc9-2379-11ea-bb65-0ac135e8bacf), paediatric cancer mutation data (
https://www.cbioportal.org/study/summary?id=pediatric_dkfz_2017
) and transposon-based mutagenesis screens from the Candidate Cancer Gene Database (
http://ccgd-starrlab.oit.umn.edu/index.html
; downloaded 26 March 2024).
Code availability
Open-source software for cell map construction (Cell Mapping Toolkit) has been released as a series of Python PyPI packages at GitHub (
https://github.com/idekerlab/cellmaps_pipeline
) and is also linked through the Multiscale Integrated Cell web portal (
https://musicmaps.ai/u2os-cellmap
).
Change history
29 April 2025
In the version of the article initially published, the text “Knut and Alice Wallenberg Foundation (2021.0346 E.L.), Göran Gustafsson Foundation (E.L), Stanford Institute for Human-Centered AI and Param Hansa Philanthropies” was missing from the Acknowledgments section and has now been added to the HTML and PDF versions of the article.
02 October 2025
A Correction to this paper has been published:
https://doi.org/10.1038/s41586-025-09648-x
References
Greenblatt, J. F., Alberts, B. M. & Krogan, N. J. Discovery and significance of protein-protein interactions in health and disease.
Cell
187
, 6501–6517 (2024).
Article
CAS
PubMed
Google Scholar
Kustatscher, G. et al. Understudied proteins: opportunities and challenges for functional proteomics.
Nat. Methods
19
, 774–779 (2022).
Article
CAS
PubMed
Google Scholar
Hu, M. et al. Evaluation of large language models for discovery of gene set function.
Nat. Methods
https://doi.org/10.1038/s41592-024-02525-x
(2024).
Gröbner, S. N. et al. The landscape of genomic alterations across childhood cancers.
Nature
555
, 321–327 (2018).
Article
ADS
PubMed
Google Scholar
Schaffer, L. V. & Ideker, T. Mapping the multiscale structure of biological systems.
Cell Syst.
12
, 622–635 (2021).
Article
CAS
PubMed
PubMed Central
Google Scholar
Harold, F. M. Molecules into cells: specifying spatial architecture.
Microbiol. Mol. Biol. Rev.
69
, 544–564 (2005).
Article
CAS
PubMed
PubMed Central
Google Scholar
Tomita, M. Whole-cell simulation: a grand challenge of the 21st century.
Trends Biotechnol.
19
, 205–210 (2001).
Article
CAS
PubMed
Google Scholar
Robinson, C. V., Sali, A. & Baumeister, W. The molecular sociology of the cell.
Nature
450
, 973–982 (2007).
Article
ADS
CAS
PubMed
Google Scholar
Johnson, G. T. et al. Building the next generation of virtual cells to understand cellular biology.
Biophys. J.
122
, 3560–3569 (2023).
Article
ADS
CAS
PubMed
PubMed Central
Google Scholar
Heinrich, L. et al. Whole-cell organelle segmentation in volume electron microscopy.
Nature
599
, 141–146 (2021).
Article
ADS
CAS
PubMed
Google Scholar
Nogales, E. & Mahamid, J. Bridging structural and cell biology with cryo-electron microscopy.
Nature
628
, 47–56 (2024).
Article
ADS
CAS
PubMed
PubMed Central
Google Scholar
Thul, P. J. et al. A subcellular map of the human proteome.
Science
356
, eaal3321 (2017).
Article
PubMed
Google Scholar
Mikuni, T., Nishiyama, J., Sun, Y., Kamasawa, N. & Yasuda, R. High-throughput, high-resolution mapping of protein localization in mammalian brain by in vivo genome editing.
Cell
165
, 1803–1817 (2016).
Article
CAS
PubMed
PubMed Central
Google Scholar
Huttlin, E. L. et al. The BioPlex Network: a systematic exploration of the human interactome.
Cell
162
, 425–440 (2015).
Article
CAS
PubMed
PubMed Central
Google Scholar
Wheat, A. et al. Protein interaction landscapes revealed by advanced in vivo cross-linking–mass spectrometry.
Proc. Natl Acad. Sci. USA
118
, e2023360118 (2021).
Article
CAS
PubMed
PubMed Central
Google Scholar
Havugimana, P. C. et al. A census of human soluble protein complexes.
Cell
150
, 1068–1081 (2012).
Article
CAS
PubMed
PubMed Central
Google Scholar
Bludau, I. et al. Complex-centric proteome profiling by SEC-SWATH-MS for the parallel detection of hundreds of protein complexes.
Nat. Protoc.
15
, 2341–2386 (2020).
Article
CAS
PubMed
Google Scholar
Go, C. D. et al. A proximity-dependent biotinylation map of a human cell.
Nature
595
, 120–124 (2021).
Article
ADS
CAS
PubMed
Google Scholar
Dunkley, T. P. J., Watson, R., Griffin, J. L., Dupree, P. & Lilley, K. S. Localization of organelle proteins by isotope tagging (LOPIT).
Mol. Cell. Proteom.
3
, 1128–1134 (2004).
Article
CAS
Google Scholar
Mulvey, C. M. et al. Using hyperLOPIT to perform high-resolution mapping of the spatial proteome.
Nat. Protoc.
12
, 1110–1135 (2017).
Article
CAS
PubMed
Google Scholar
Qin, Y. et al. A multi-scale map of cell structure fusing protein images and interactions.
Nature
600
, 536–542 (2021).
Article
ADS
CAS
PubMed
PubMed Central
Google Scholar
Huttlin, E. L. et al. Dual proteome-scale networks reveal cell-specific remodeling of the human interactome.
Cell
184
, 3022–3040 (2021).
Article
CAS
PubMed
PubMed Central
Google Scholar
Girdhar, R. et al. ImageBind: one embedding space to bind them all. In
Proc. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
15180–15190 (IEEE, 2023).
Schroff, F., Kalenichenko, D. & Philbin, J. FaceNet: a unified embedding for face recognition and clustering. In
Proc. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
815–823 (IEEE, 2015).
Tsitsiridis, G. et al. CORUM: the comprehensive resource of mammalian protein complexes–2022.
Nucleic Acids Res.
51
, D539–D545 (2022).
Article
PubMed Central
Google Scholar
Ashburner, M. et al. Gene Ontology: tool for the unification of biology.
Nat. Genet.
25
, 25–29 (2000).
Article
CAS
PubMed
PubMed Central
Google Scholar
OpenAI. GPT-4 Technical Report. Preprint at
https://doi.org/10.48550/arXiv.2303.08774
(2023).
Pappu, R. V., Cohen, S. R., Dar, F., Farag, M. & Kar, M. Phase transitions of associative biomacromolecules.
Chem. Rev.
123
, 8945–8987 (2023).
Article
CAS
PubMed
PubMed Central
Google Scholar
Saar, K. L. et al. Protein Condensate Atlas from predictive models of heteromolecular condensate composition.
Nat. Commun.
15
, 5418 (2024).
Article
ADS
CAS
PubMed
PubMed Central
Google Scholar
Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold.
Nature
596
, 583–589 (2021).
Article
ADS
CAS
PubMed
PubMed Central
Google Scholar
Evans, R. et al. Protein complex prediction with AlphaFold-Multimer. Preprint at
bioRxiv
https://doi.org/10.1101/2021.10.04.463034
(2021).
Desprez, F., Ung, D. C., Vourc’h, P., Jeanne, M. & Laumonnier, F. Contribution of the dihydropyrimidinase-like proteins family in synaptic physiology and in neurodevelopmental disorders.
Front. Neurosci.
17
, 1154446 (2023).
Article
PubMed
PubMed Central
Google Scholar
Kim, M. H. & Kang, B. S. Structure and dynamics of the human multi-tRNA synthetase complex.
Subcell. Biochem.
99
, 199–233 (2022).
Article
CAS
PubMed
Google Scholar
Colaço, A. & Jäättelä, M. Ragulator—a multifaceted regulator of lysosomal signaling and trafficking.
J. Cell Biol.
216
, 3895–3898 (2017).
Article
PubMed
PubMed Central
Google Scholar
Rout, M. P. & Sali, A. Principles for integrative structural biology studies.
Cell
177
, 1384–1403 (2019).
Article
CAS
PubMed
PubMed Central
Google Scholar
Fromm, S. A., Lawrence, R. E. & Hurley, J. H. Structural mechanism for amino acid-dependent Rag GTPase nucleotide state switching by SLC38A9.
Nat. Struct. Mol. Biol.
27
, 1017–1023 (2020).
Article
CAS
PubMed
PubMed Central
Google Scholar
Duek, P., Gateau, A., Bairoch, A. & Lane, L. Exploring the uncharacterized human proteome using neXtProt.
J. Proteome Res.
17
, 4211–4226 (2018).
Article
CAS
PubMed
Google Scholar
Platanias, L. C. Mechanisms of type-I- and type-II-interferon-mediated signalling.
Nat. Rev. Immunol.
5
, 375–386 (2005).
Article
CAS
PubMed
Google Scholar
Zhang, H., Chen, Y., Keane, F. M. & Gorrell, M. D. Advances in understanding the expression and function of dipeptidyl peptidase 8 and 9.
Mol. Cancer Res.
11
, 1487–1496 (2013).
Article
CAS
PubMed
Google Scholar
Jiaang, W.-T. et al. Novel isoindoline compounds for potent and selective inhibition of prolyl dipeptidase DPP8.
Bioorg. Med. Chem. Lett.
15
, 687–691 (2005).
Article
CAS
PubMed
Google Scholar
Breckels, L. M. et al. Advances in spatial proteomics: mapping proteome architecture from protein complexes to subcellular localizations.
Cell Chem. Biol.
31
, 1665–1687 (2024).
Article
CAS
PubMed
Google Scholar
Donnio, L.-M. et al. XAB2 dynamics during DNA damage-dependent transcription inhibition.
eLife
11
, e77094 (2022).
Article
CAS
PubMed
PubMed Central
Google Scholar
Kaden, D., Munter, L. M., Reif, B. & Multhaup, G. The amyloid precursor protein and its homologues: structural and functional aspects of native and pathogenic oligomerization.
Eur. J. Cell Biol.
91
, 234–239 (2012).
Article
CAS
PubMed
Google Scholar
Rogelj, B., Mitchell, J. C., Miller, C. C. J. & McLoughlin, D. M. The X11/Mint family of adaptor proteins.
Brain Res. Rev.
52
, 305–315 (2006).
Article
CAS
PubMed
Google Scholar
Mittal, P. & Roberts, C. W. M. The SWI/SNF complex in cancer—biology, biomarkers and therapy.
Nat. Rev. Clin. Oncol.
17
, 435–448 (2020).
Article
CAS
PubMed
PubMed Central
Google Scholar
Abbott, K. L. et al. The Candidate Cancer Gene Database: a database of cancer driver genes from forward genetic screens in mice.
Nucleic Acids Res.
43
, D844–D848 (2015).
Article
CAS
PubMed
Google Scholar
Zhao, Y., Lin, H., Jiang, J., Ge, M. & Liang, X. TBL1XR1 as a potential therapeutic target that promotes epithelial-mesenchymal transition in lung squamous cell carcinoma.
Exp. Ther. Med.
17
, 91–98 (2019).
CAS
PubMed
Google Scholar
St-Jean, S. et al. NCOR1 sustains colorectal cancer cell growth and protects against cellular senescence.
Cancers
13
, 4414 (2021).
Shannon, P. et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks.
Genome Res.
13
, 2498–2504 (2003).
Article
CAS
PubMed
PubMed Central
Google Scholar
Lander, E. S. et al. Initial sequencing and analysis of the human genome.
Nature
409
, 860–921 (2001).
Article
ADS
CAS
PubMed
Google Scholar
Cho, N. H. et al. OpenCell: Endogenous tagging for the cartography of human cellular organization.
Science
375
, eabi6983 (2022).
Article
CAS
PubMed
PubMed Central
Google Scholar
Vickovic, S. et al. SM-Omics is an automated platform for high-throughput spatial multi-omics.
Nat. Commun.
13
, 795 (2022).
Article
ADS
CAS
PubMed
PubMed Central
Google Scholar
Schaff, J., Fink, C. C., Slepchenko, B., Carson, J. H. & Loew, L. M. A general computational framework for modeling cellular structure and function.
Biophys. J.
73
, 1135–1146 (1997).
Article
CAS
PubMed
PubMed Central
Google Scholar
Karr, J. R. et al. A whole-cell computational model predicts phenotype from genotype.
Cell
150
, 389–401 (2012).
Article
CAS
PubMed
PubMed Central
Google Scholar
Bunne, C. et al. How to build the virtual cell with artificial intelligence: priorities and opportunities.
Cell
187
, 7045–7063 (2024).
Article
CAS
PubMed
PubMed Central
Google Scholar
McInnes, L., Healy, J., Saul, N. & Großberger, L. UMAP: uniform manifold approximation and projection.
J. Open Source Softw.
3
, 861 (2018).
Article
Google Scholar
Yang, X. et al. A public genome-scale lentiviral expression library of human ORFs.
Nat. Methods
8
, 659–661 (2011).
Article
CAS
PubMed
PubMed Central
Google Scholar
Eng, J. K., McCormack, A. L. & Yates, J. R. An approach to correlate tandem mass spectral data of peptides with amino acid sequences in a protein database.
J. Am. Soc. Mass. Spectrom.
5
, 976–989 (1994).
Article
ADS
CAS
PubMed
Google Scholar
Sowa, M. E., Bennett, E. J., Gygi, S. P. & Harper, J. W. Defining the human deubiquitinating enzyme interaction landscape.
Cell
138
, 389–403 (2009).
Article
CAS
PubMed
PubMed Central
Google Scholar
Behrends, C., Sowa, M. E., Gygi, S. P. & Harper, J. W. Network organization of the human autophagy system.
Nature
466
, 68–76 (2010).
Article
ADS
CAS
PubMed
PubMed Central
Google Scholar
Huttlin, E. L. et al. Architecture of the human interactome defines protein communities and disease networks.
Nature
545
, 505–509 (2017).
Article
ADS
CAS
PubMed
PubMed Central
Google Scholar
Skinnider, M. A. et al. An atlas of protein-protein interactions across mouse tissues.
Cell
184
, 4073–4089 (2021).
Article
CAS
PubMed
Google Scholar
Kristensen, A. R., Gsponer, J. & Foster, L. J. A high-throughput approach for measuring temporal changes in the interactome.
Nat. Methods
9
, 907–909 (2012).
Article
CAS
PubMed
PubMed Central
Google Scholar
Goodman, J. K., Zampronio, C. G., Jones, A. M. E. & Hernandez-Fernaud, J. R. Updates of the in-gel digestion method for protein analysis by mass spectrometry.
Proteomics
18
, e1800236 (2018).
Article
PubMed
Google Scholar
Ishihama, Y., Rappsilber, J., Andersen, J. S. & Mann, M. Microcolumns with self-assembled particle frits for proteomics.
J. Chromatogr. A
979
, 233–239 (2002).
Article
CAS
PubMed
Google Scholar
McAfee, A., Chapman, A., Bao, G., Tarpy, D. R. & Foster, L. J. Investigating trade-offs between ovary activation and immune protein expression in bumble bee (
Bombus impatiens
) workers and queens.
Proc. Biol. Sci.
291
, 20232463 (2024).
CAS
PubMed
PubMed Central
Google Scholar
Demichev, V., Messner, C. B., Vernardis, S. I., Lilley, K. S. & Ralser, M. DIA-NN: neural networks and interference correction enable deep proteome coverage in high throughput.
Nat. Methods
17
, 41–44 (2020).
Article
CAS
PubMed
Google Scholar
Grover, A. & Leskovec, J. node2vec: scalable feature learning for networks. In
Proc. 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
855–864 (ACM, 2016).
Ouyang, W. et al. Analysis of the Human Protein Atlas Image Classification competition.
Nat. Methods
16
, 1254–1261 (2019).
Article
CAS
PubMed
PubMed Central
Google Scholar
Bao, F. et al. Integrative spatial analysis of cell morphologies and transcriptional states with MUSE.
Nat. Biotechnol.
40
, 1200–1209 (2022).
Article
CAS
PubMed
Google Scholar
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I. & Salakhutdinov, R. Dropout: a simple way to prevent neural networks from overfitting.
J. Mach. Learn. Res.
15
, 1929–1958 (2014).
MathSciNet
MATH
Google Scholar
Ioffe, S. & Szegedy, C. Batch normalization: accelerating deep network training by reducing internal covariate shift. In
Proc. 32nd International Conference on International Conference on Machine Learning—Volume 37
(eds Bach, F. & Blei, D.) 448–456 (JMLR, 2015).
Blondel, V. D., Guillaume, J.-L., Lambiotte, R. & Lefebvre, E. Fast unfolding of communities in large networks.
J. Stat. Mech.
2008
, P10008 (2008).
Article
Google Scholar
Paszke, A. et al. PyTorch: an imperative style, high-performance deep learning library. In
Proc. 33rd International Conference on Neural Information Processing Systems
(eds Wallach, H. M. et al.) 8026–8037 (Curran, 2019).
Kingma, D. P. & Ba, J. Adam: a method for stochastic optimization. In
Proc.
3rd International Conference on Learning Representations
(ICLR) (eds Bengio, Y. & LeCun, Y.) (DBLP, 2014).
Choubineh, A., Chen, J., Coenen, F. & Ma, F. Applying Monte Carlo dropout to quantify the uncertainty of skip connection-based convolutional neural networks optimized by big data.
Electronics
12
, 1453 (2023).
Article
Google Scholar
Zhao, Y., Jin, Z., Qi, G.-J., Lu, H. & Hua, X.-S. An adversarial approach to hard triplet generation. In
Proc. European Conference on Computer Vision (ECCV)
(eds Ferrari, V. & Hebert, M.) 508–524 (Springer, 2018).
Snel, B., Lehmann, G., Bork, P. & Huynen, M. A. STRING: a web-server to retrieve and display the repeatedly occurring neighbourhood of a gene.
Nucleic Acids Res.
28
, 3442–3444 (2000).
Article
CAS
PubMed
PubMed Central
Google Scholar
Szklarczyk, D. et al. STRING v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets.
Nucleic Acids Res.
47
, D607–D613 (2018).
Article
PubMed Central
Google Scholar
Replogle, J. M. et al. Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq.
Cell
185
, 2559–2575 (2022).
Article
CAS
PubMed
PubMed Central
Google Scholar
Zheng, F. et al. HiDeF: identifying persistent structures in multiscale’omics data.
Genome Biol.
22
, 21 (2021).
Article
PubMed
PubMed Central
Google Scholar
Wei, J. et al. Chain of thought prompting elicits reasoning in large language models. In
Proc. Advances in Neural Information Processing Systems
(eds Koyejo, S. et al.) 24824–24837 (Curran, 2022).
Brown, T. B. et al. Language models are few-shot learners. In
Proc. 34th International Conference on Neural Information Processing Systems
(eds Larochelle, H. et al.) 1877–1901 (Curran, 2020).
Liu, F., Shareghi, E., Meng, Z., Basaldella, M. & Collier, N. Self-alignment pretraining for biomedical entity representations. In
Proc. 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
(eds Toutanova, K. et al.) 4228–4238 (Association for Computational Linguistics, 2021).
Wolf, T. et al. Transformers: state-of-the-art natural language processing. In
Proc. 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations
(eds Liu, Q. & Schlangen, D.) 38–45 (Association for Computational Linguistics, 2020).
Erdős, G., Pajkos, M. & Dosztányi, Z. IUPred3: prediction of protein disorder enhanced with unambiguous experimental annotation and visualization of evolutionary conservation.
Nucleic Acids Res.
49
, W297–W303 (2021).
Article
PubMed
PubMed Central
Google Scholar
Hardenberg, M., Horvath, A., Ambrus, V., Fuxreiter, M. & Vendruscolo, M. Widespread occurrence of the droplet state of proteins in the human proteome.
Proc. Natl Acad. Sci. USA
117
, 33254–33262 (2020).
Article
ADS
CAS
PubMed
PubMed Central
Google Scholar
Rostam, N. et al. CD-CODE: crowdsourcing condensate database and encyclopedia.
Nat. Methods
20
, 673–676 (2023).
Article
CAS
PubMed
PubMed Central
Google Scholar
Stacey, R. G., Skinnider, M. A., Scott, N. E. & Foster, L. J. A rapid and accurate approach for prediction of interactomes from co-elution data (PrInCE).
BMC Bioinform.
18
, 457 (2017).
Article
Google Scholar
Mirdita, M. et al. ColabFold: making protein folding accessible to all.
Nat. Methods
19
, 679–682 (2022).
Article
CAS
PubMed
PubMed Central
Google Scholar
Russel, D. et al. Putting the pieces together: integrative modeling platform software for structure determination of macromolecular assemblies.
PLoS Biol.
10
, e1001244 (2012).
Article
CAS
PubMed
PubMed Central
Google Scholar
Saltzberg, D. et al. Modeling biological complexes using Integrative Modeling Platform.
Methods Mol. Biol.
2022
, 353–377 (2019).
Article
CAS
PubMed
Google Scholar
Saltzberg, D. J. et al. Using integrative modeling platform to compute, validate, and archive a model of a protein complex structure.
Protein Sci.
30
, 250–261 (2021).
Article
CAS
PubMed
Google Scholar
Sali, A. From integrative structural biology to cell biology.
J. Biol. Chem.
296
, 100743 (2021).
Article
CAS
PubMed
PubMed Central
Google Scholar
Zhu, W., Shenoy, A., Kundrotas, P. & Elofsson, A. Evaluation of AlphaFold-Multimer prediction on multi-chain protein complexes.
Bioinformatics
39
, btad424 (2023).
Article
CAS
PubMed
PubMed Central
Google Scholar
Metropolis, N. & Ulam, S. The Monte Carlo method.
J. Am. Stat. Assoc.
44
, 335–341 (1949).
Article
CAS
PubMed
MATH
Google Scholar
Viswanath, S., Chemmama, I. E., Cimermancic, P. & Sali, A. Assessing exhaustiveness of stochastic sampling for integrative modeling of macromolecular structures.
Biophys. J.
113
, 2344–2353 (2017).
Article
ADS
CAS
PubMed
PubMed Central
Google Scholar
Cliff, N. Dominance statistics: ordinal analyses to answer ordinal questions.
Psychol. Bull.
114
, 494–509 (1993).
Article
Google Scholar
Cerami, E. et al. The cBio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data.
Cancer Discov.
2
, 401–404 (2012).
Article
PubMed
Google Scholar
Gao, J. et al. Integrative analysis of complex cancer genomics and clinical profiles using the cBioPortal.
Sci. Signal.
6
, l1 (2013).
Article
Google Scholar
Zheng, F. et al. Interpretation of cancer mutations using a multiscale map of protein systems.
Science
374
, eabf3067 (2021).
Article
CAS
PubMed
PubMed Central
Google Scholar
Bailey, M. H. et al. Comprehensive characterization of cancer driver genes and mutations.
Cell
173
, 371–385 (2018).
Article
CAS
PubMed
PubMed Central
Google Scholar
Ma, X. et al. Pan-cancer genome and transcriptome analyses of 1,699 paediatric leukaemias and solid tumours.
Nature
555
, 371–376 (2018).
Article
ADS
CAS
PubMed
PubMed Central
Google Scholar
Virtanen, P. et al. SciPy 1.0: fundamental algorithms for scientific computing in Python.
Nat. Methods
17
, 261–272 (2020).
Article
CAS
PubMed
PubMed Central
Google Scholar
Huttlin, E. L. et al.
BioPlex Interactome
https://bioplex.hms.harvard.edu/
(2015).
Thomas, P. D. et al. PANTHER: making genome-scale phylogenetics accessible to all.
Protein Sci.
31
, 8–22 (2022).
Article
CAS
PubMed
Google Scholar
Download references
Acknowledgements
We acknowledge funding from Schmidt Futures (T.I., E.L.), the Bridge2AI Program (NIH Common Fund; OT2 OD032742; T.I., E.L., A.S.), the Cancer Cell Map Initiative (NCI Center for Cancer Systems Biology; U54 CA274502; T.I., E.L., A.S.), the Cytoscape (5U24HG012107; T.I.) and the Network Data Exchange programs (5U24CA269436; T.I., D.P.), the Canada Foundation for Innovation and Genome BC (374PRO, L.J.F.), Knut and Alice Wallenberg Foundation (2021.0346 E.L.), Göran Gustafsson Foundation (E.L), Stanford Institute for Human-Centered AI and Param Hansa Philanthropies, and NIH grants R01GM083960 and P41GM109824 (A.S.). We acknowledge funding for the BioPlex project from NHGRI U24 HG006673 (S.P.G., J.W.H., E.L.H.), Third Rock Ventures (S.P.G., J.W.H., E.L.H.), Google Ventures (S.P.G., J.W.H., E.L.H.), Interline Therapeutics (S.P.G., E.L.H.) and Xaira (S.P.G., E.L.H.).
Author information
Author notes
These authors contributed equally: Leah V. Schaffer, Mengzhou Hu
Authors and Affiliations
Department of Medicine, University of California San Diego, La Jolla, CA, USA
Leah V. Schaffer, Mengzhou Hu, Gege Qian, Dorothy Tsai, Nicole M. Mattson, Katherine Licon, Robin Bachelder, Yue Qin, Xiaoyu Zhao, Christopher Churas, Joanna Lenkiewicz, Jing Chen, Keiichiro Ono, Dexter Pratt & Trey Ideker
Bioinformatics and Systems Biology Program, University of California San Diego, La Jolla, CA, USA
Gege Qian
Department of Biochemistry & Molecular Biology, Michael Smith Laboratories, University of British Columbia, Vancouver, British Columbia, Canada
Kyung-Mee Moon & Leonard J. Foster
Department of Bioengineering and Therapeutic Sciences, University of California San Francisco, San Francisco, CA, USA
Abantika Pal, Neelesh Soni, Andrew P. Latham, Aji Palar & Andrej Sali
Department of Cell Biology, Harvard Medical School, Boston, MA, USA
Laura Pontano Vaites, J. Wade Harper, Steven P. Gygi & Edward L. Huttlin
Department of Bioengineering, Stanford University, Palo Alto, CA, USA
Anthony Cesnik, Ishan Gaur, Trang Le, William Leineweber, Ernst Pulido & Emma Lundberg
Broad Institute of MIT and Harvard, Boston, MA, USA
Yue Qin
Department of Pediatrics, Division of Hematology-Oncology, University of California San Diego, La Jolla, CA, USA
Peter Zage
Department of Cellular and Molecular Pharmacology, University of California San Francisco, San Francisco, CA, USA
Ignacia Echeverria
Quantitative Biosciences Institute, University of California San Francisco, San Francisco, CA, USA
Ignacia Echeverria & Andrej Sali
Department of Pharmaceutical Chemistry, University of California San Francisco, San Francisco, CA, USA
Andrej Sali
Department of Pathology, Stanford University, Palo Alto, CA, USA
Emma Lundberg
Science for Life Laboratory, School of Engineering Sciences in Chemistry, Biotechnology and Health, KTH Royal Institute of Technology, Stockholm, Sweden
Emma Lundberg
Chan Zuckerberg Biohub, San Francisco, CA, USA
Emma Lundberg
Department of Computer Science and Engineering, University of California San Diego, La Jolla, CA, USA
Trey Ideker
Department of Bioengineering, University of California San Diego, La Jolla, CA, USA
Trey Ideker
Authors
Leah V. Schaffer
View author publications
Search author on:
PubMed
Google Scholar
Mengzhou Hu
View author publications
Search author on:
PubMed
Google Scholar
Gege Qian
View author publications
Search author on:
PubMed
Google Scholar
Kyung-Mee Moon
View author publications
Search author on:
PubMed
Google Scholar
Abantika Pal
View author publications
Search author on:
PubMed
Google Scholar
Neelesh Soni
View author publications
Search author on:
PubMed
Google Scholar
Andrew P. Latham
View author publications
Search author on:
PubMed
Google Scholar
Laura Pontano Vaites
View author publications
Search author on:
PubMed
Google Scholar
Dorothy Tsai
View author publications
Search author on:
PubMed
Google Scholar
Nicole M. Mattson
View author publications
Search author on:
PubMed
Google Scholar
Katherine Licon
View author publications
Search author on:
PubMed
Google Scholar
Robin Bachelder
View author publications
Search author on:
PubMed
Google Scholar
Anthony Cesnik
View author publications
Search author on:
PubMed
Google Scholar
Ishan Gaur
View author publications
Search author on:
PubMed
Google Scholar
Trang Le
View author publications
Search author on:
PubMed
Google Scholar
William Leineweber
View author publications
Search author on:
PubMed
Google Scholar
Aji Palar
View author publications
Search author on:
PubMed
Google Scholar
Ernst Pulido
View author publications
Search author on:
PubMed
Google Scholar
Yue Qin
View author publications
Search author on:
PubMed
Google Scholar
Xiaoyu Zhao
View author publications
Search author on:
PubMed
Google Scholar
Christopher Churas
View author publications
Search author on:
PubMed
Google Scholar
Joanna Lenkiewicz
View author publications
Search author on:
PubMed
Google Scholar
Jing Chen
View author publications
Search author on:
PubMed
Google Scholar
Keiichiro Ono
View author publications
Search author on:
PubMed
Google Scholar
Dexter Pratt
View author publications
Search author on:
PubMed
Google Scholar
Peter Zage
View author publications
Search author on:
PubMed
Google Scholar
Ignacia Echeverria
View author publications
Search author on:
PubMed
Google Scholar
Andrej Sali
View author publications
Search author on:
PubMed
Google Scholar
J. Wade Harper
View author publications
Search author on:
PubMed
Google Scholar
Steven P. Gygi
View author publications
Search author on:
PubMed
Google Scholar
Leonard J. Foster
View author publications
Search author on:
PubMed
Google Scholar
Edward L. Huttlin
View author publications
Search author on:
PubMed
Google Scholar
Emma Lundberg
View author publications
Search author on:
PubMed
Google Scholar
Trey Ideker
View author publications
Search author on:
PubMed
Google Scholar
Contributions
L.V.S., M.H., E.L. and T.I. designed the study. L.V.S., M.H., G.Q., X.Z., T.L., A. Pal, A.P.L., Y.Q., P.Z., E.L., and T.I. developed ideas for data analyses. L.V.S. and M.H. implemented computational methods and analyses. R.B., A.C., I.G., T.L., W.L., A. Palar, E.P., L.V.S., M.H., D.P., I.E., A.S., E.L. and T.I. annotated the hierarchy. A. Pal, N.S., I.E. and A.S. designed and performed structural modelling. E.L.H., L.P.V., J.W.H. and S.P.G. generated and analysed the AP–MS data. K.-M.M. and L.J.F. generated and analysed the SEC–MS data. D.T., K.L. and N.M.M. generated qPCR data. L.V.S., C.C., J.L. and D.P. designed and implemented the cell map toolkit. K.O., D.P. and J.C. designed and implemented Cytoscape Web. L.V.S., M.H., E.L. and T.I. wrote the manuscript with input from all of the authors.
Corresponding authors
Correspondence to
Edward L. Huttlin
,
Emma Lundberg
or
Trey Ideker
.
Ethics declarations
Competing interests
T.I. is a co-founder, advisor and holder of equity for Data4Cure and Serinus Biosciences, and he is an advisor and shareholder for Ideaya BioSciences. The terms of these arrangements have been reviewed and approved by UC San Diego in accordance with its conflict of interest policies. E.L. is an advisor for and has equity interest in Cartography Biosciences, Element Biosciences, Santa Ana Bio, Pixelgen Technologies and Moleculent. The terms of these arrangements have been reviewed and approved by Stanford University in accordance with its conflict of interest policies. The BioPlex project has been supported by Interline Therapeutics and Xaira Therapeutics (S.P.G. and E.L.H.). J.W.H. is a co-founder for Caraway Therapeutics (a subsidiary of Merck) and is a scientific advisory board member for Lyterian Therapeutics. S.P.G. is on the scientific advisory board for Thermo Fisher Scientific, Cell Signaling Technology and Frontier Medicine. E.L.H. is a consultant for Matchpoint Therapeutics, Flagship Ventures and Calico Life Sciences. The other authors declare no competing interests.
Peer review
Peer review information
Nature
thanks Georg Kustatscher and the other, anonymous, reviewer(s) for their contribution to the peer review of this work.
Peer reviewer reports
are available.
Additional information
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended data figures and tables
Extended Data Fig. 1 Data quality assessment.
a)
Complete network of U2OS protein-protein interactions as measured by AP-MS.
b)
Histogram of number of interactions per protein.
c)
Fraction of CORUM complexes significantly enriched (1% FDR,
Methods
) for AP-MS interactions measured in USOS (left) alongside interactions ascertained in two previously published AP-MS networks for other cell lines
22
,
51
(middle and right). Error bars denote 95% confidence intervals.
d)
Interaction networks for select CORUM complexes, with FDR q-values as per panel (
c
). Blue nodes denote bait proteins and grey nodes denote prey proteins.
e)
PANTHER
106
classifications of protein function (top 30 largest classes by number of proteins), shown for proteins covered by U2OS cell map in comparison to the entire human proteome (UniProt, downloaded September 11, 2023).
f)
Protein pairs ranked by cosine similarity in AP-MS features enrich for the most similar protein pairs (top 1% in the immunofluorescent protein image features and
g)
vice versa.
Extended Data Fig. 2 Self-supervised embedding of multiple data modalities.
a)
Architecture of self-supervised multimodal embedding model. Columns of squares represent feature vectors, with the dimensionality written just below each column. Regions enclosed by dotted lines represent neural networks with layers described. Protein coordinates in the joint multimodal embedding (z) are used for computing pairwise protein-protein similarities in subsequent panels (cosine similarity function).
b)
Distribution of similarities shown for protein pairs with a ‘high-confidence interaction’ denoted in the STRING database (green) in comparison to all other protein pairs (grey).
c)
Similar to (b) but for protein pairs in the same CORUM complex.
d)
Similar to (b) but for protein pairs that yield highly similar transcriptional profiles (top 1% pairs) when genetically disrupted by CRISPR, drawn from a recent perturb-seq functional genomics study
80
. **** denotes significant difference,
p
< 0.0001 by one-sided Wilcoxon rank-sum test.
e)
Different protein embedding approaches (coloured points,
Methods
) are evaluated by their degree of enrichment (x-axis) across orthogonal functional and physical interaction resources (y-axis, resources from panels b-d above). Supervised Random Forest trained using the Gene Ontology (
Methods
). Enrichment computed using Cliff’s Delta (1,000 samplings of 1,000 protein pairs with replacement,
Methods
) yielding values in range [–1,1], with positive values indicating enrichment above random expectation. Error bars denote standard deviations across 1000 bootstrap resamplings with the center at the mean. * denotes significant difference in comparison with self-supervised multimodal embedding results (two-tailed
p
< 0.05 across bootstrap resamplings).
Extended Data Fig. 3 Robustness of cell map assemblies.
(a-c)
Cell maps are coloured by assembly robustness, measured as the fraction of 300 jackknife resamplings where an assembly was recovered (
Methods
). Three panels show maps built using
a)
both imaging and AP-MS data,
b)
AP-MS data only, or
c)
imaging data only.
d)
Number of robust assemblies (recovered in >50% jackknife resamplings) versus size of assembly in number of proteins. Grey bars denote total number of assemblies in each size category.
Extended Data Fig. 4 Annotation of protein assemblies with GPT-4.
a)
Left: Annotation workflow extended from Hu et al.
3
, in which the cell map is used to query GPT-4 for a descriptive name, a confidence score and a supporting rationale. Right: Composition of prompt used for GPT-4 query.
b)
Schematic of GPT-4 assisted citation module. GPT-4 is asked to provide gene symbol keywords and functional keywords separately. Multiple gene keywords and functions are combined and used to search PubMed for relevant paper titles and abstracts in the scientific literature.
c)
Example assembly with GPT-4 name and supporting analysis paragraphs with citations generated from citation module (see panel b).
d)
Semantic similarity of the original GPT-4 name given an assembly vs. the name assigned in each of five replicate GPT-4 runs. Results for two example assemblies are shown (yellow and green points), one of which is named identically across replicates (yellow) and one of which shows variation (green). The average performance over all U2OS assemblies (n = 271, excluding assemblies with more than 1000 proteins) is shown in dark blue.
Extended Data Fig. 5 Quality assessment of SEC-MS dataset.
a)
Overlap of protein identifications across three biological replicates.
b)
Violin plots showing the distribution of the Pearson correlation between replicate measurements of each protein’s elution pattern (purple, n = 5018) vs. random pairings of different proteins across replicates (white, n = 5018), with thick black lines representing the median.
c)
Histogram of number of proteins with maximum intensity in each elution fraction for replicate 1 (top), replicate 2 (middle), and replicate 3 (bottom). Select CORUM complexes are highlighted at the median maximum intensity of proteins in the complex.
Extended Data Fig. 6 DPP9 association with STAT interferon signalling.
a)
Interaction data for the ISGF3 complex.
b)
Immunofluorescence images for the ISGF3 complex. Members immunostained (green) with cytoskeleton counterstain (red). Scale bar, 3 µm.
c)
Relative mRNA expression level of IFNβ1, IFNγ1, IFNγ2, and negative control (Non-ISG 18S) upon DPP9 inhibition, separate samples. Expression levels normalized to DMSO control. Points (n = 6, 2 biological replicates and 3 technical replicates each) denote replicate measurements. Light grey whiskers represent mean ± SE. Significance (p-values) are determined by a two-sided student’s t-test.
d)
Canonical function of ISGF3 complex with putative upstream activity of DPP9.
Extended Data Fig. 7 Comparison of protein assemblies across U2OS and HEK293 cell lines.
a)
Schematic of the approach for determining conserved vs. U2OS-specific protein assemblies in the cell map. For each assembly, the cosine similarities are determined between protein AP-MS features for U2OS and HEK293, separately. The U2OS similarities are then compared to the HEK293 similarities with a two-sided Mann-Whitney U-test.
b)
U2OS cell map (see Fig.
2b
), where assembly colour indicates effect size (Cliff’s Delta). 18 assemblies that did not have sufficient data in both cell lines were removed. Dashed boxes denote examples of strongly conserved and strongly U2OS-specific assemblies.
c)
Biophysical interaction data for 9-1-1 RAD-RFC complex; edge colour signifies presence in HEK293 (orange) or U2OS (blue) interaction networks.
d)
Immunofluorescence images for RFC1 (orange) in HEK293 cells (top) or U2OS cells (bottom), with cytoskeleton counterstain (red). Scale bar, 5 µm.
e)
Biophysical interaction data for the Energy metabolism regulation complex; edge colour signifies presence in HEK293 (orange) or U2OS (blue) interaction networks.
Extended Data Fig. 8 Analysis of multi-localized proteins and assemblies.
a)
Number of distinct assemblies per protein, defined as the number of distinct paths to the root of the U2OS cell map hierarchy (shown in Fig.
2b
).
b)
Tripartite localization of XAB2 in the cell map visualized in circle-packing mode. XAB2 is highlighted by red circle and the cell map assemblies in which it participates are highlighted with orange borders. Cytosol and nucleus are filled as blue and pink respectively.
c)
Top: Immunofluorescence images of XAB2 (middle) and representative interacting partners in cytosol (WDR83, left) or nucleus (DHX8, right). These proteins are immunostained (green) with cytoskeleton counterstain (red). Scale bar, 2.5 µm. Bottom: Biophysical interaction partners of XAB2 with cell map localizations in cytosol and membrane (turquoise box) or nucleus (pink box).
d)
Cell map coloured to indicate multi-localized assemblies (red nodes). The multiple containing assemblies are indicated in each case (red edges). Dashed box denotes the assembly detailed in text and in panel f–g.
e)
Number of assemblies with single (left) versus multiple (right) localizations.
f)
Biophysical interaction data for Amyloid Precursor Protein (APP) complex.
g)
Immunofluorescence images for proteins in APP complex, immunostained (green) with cytoskeleton counterstain (red). Scale bar, 2.5 µm.
Extended Data Fig. 9 Association of protein assemblies with tumour growth via transposon mutagenesis.
a)
LEFT: Schematic overview of transposon-based genetic screens in mouse tumour models
46
. RIGHT: Number of screens by cancer type.
b)
Distribution of the number of transposon screens identifying each human gene as a cancer driver. Separate curves show cancer proteins in recurrently mutated assemblies (magenta curve) versus all other proteins in the cell map (grey curve).
P
value between the two distributions are determined by one-sided Mann Whitney U test. Black dashed line represents the threshold number of screens used to call cancer drivers (threshold = 4, FDR < 0.3).
c)
Number of proteins validated as cancer drivers in recurrently mutated assemblies (magenta), versus cancer drivers previously identified in pediatric (gold) or adult cancer studies (lavender).
d)
Number of transposon screening studies identifying a protein as a cancer driver (total n = 72), shown for select cancer assemblies (rows). Light grey whiskers represent mean ± SE.
e)
As for panel d, focusing on proteins in NATR assembly.
f)
Immunofluorescence images for six representative proteins in the NATR assembly found to be mutated in certain pediatric tumours. Members immunostained (green) with cytoskeleton counterstain (red). Scale bar, 2.5 µm.
Extended Data Fig. 10 Navigating the multiscale cell map.
Portal for visualization of the U2OS cell map and associated data, available at
http://musicmaps.ai/u2os-cellmap/
. In the main display, subcellular components (protein assemblies) are displayed either as a hierarchy (as in Fig.
2b
) or as a circle-packing diagram (shown here, left). In circle-packing mode, nesting of one circle inside another indicates containment of an assembly inside another; proteins are represented by the innermost circles which do not contain others; the outermost circle represents the whole cell. An endonuclease assembly has been selected in the main display (circle with yellow border), with its supporting multimodal data shown in detail in a supplemental window display (upper right). Edge colours represent evidence types: dark blue: AP-MS edges; pink: similar images; purple: similar multimodal embeddings. Further selection of nodes in the subnetwork displays protein-level information with direct links to raw data, including immunofluorescence images (lower right).
Supplementary information
Supplementary Tables (download XLSX
)
Supplementary Tables 1–10.
Reporting Summary (download PDF
)
Peer Review File (download PDF
)
Rights and permissions
Open Access
This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit
http://creativecommons.org/licenses/by/4.0/
.
Reprints and permissions
About this article
Cite this article
Schaffer, L.V., Hu, M., Qian, G.
et al.
Multimodal cell maps as a foundation for structural and functional genomics.
Nature
642
, 222–231 (2025). https://doi.org/10.1038/s41586-025-08878-3
Download citation
Received
:
03 June 2024
Accepted
:
10 March 2025
Published
:
09 April 2025
Version of record
:
09 April 2025
Issue date
:
05 June 2025
DOI
:
https://doi.org/10.1038/s41586-025-08878-3
Share this article
Anyone you share the following link with will be able to read this content:
Get shareable link
Sorry, a shareable link is not currently available for this article.
Copy shareable link to clipboard
Provided by the Springer Nature SharedIt content-sharing initiative
Download PDF
Advertisement
Explore content
Research articles
News
Opinion
Research Analysis
Careers
Books & Culture
Podcasts
Videos
Current issue
Browse issues
Collections
Subjects
Follow us on Facebook
Follow us on Bluesky
Follow us on X
Sign up for alerts
RSS feed
About the journal
Journal Staff
About the Editors
Research Cross-Journal Editorial Team
Journal Information
Journal Metrics
Our publishing models
Editorial Values Statement
Editorial policies
Journalistic Principles
History of Nature
Awards
Contact
Send a news tip
Publish with us
For Authors
For Referees
Language editing services
Open access funding
Submit manuscript
Search
Search articles by subject, keyword or author
Show results from
All journals
This journal
Search
Advanced search
Quick links
Explore articles by subject
Find a job
Guide to authors
Editorial policies
Nature
(
Nature
)
ISSN
1476-4687
(online)
ISSN
0028-0836
(print)
nature.com footer links
About Nature Portfolio
About us
Press releases
Press office
Contact us
Discover content
Journals A-Z
Articles by subject
protocols.io
Nature Index
Publishing policies
Nature portfolio policies
Open access
Author & Researcher services
Reprints & permissions
Research data
Language editing
Scientific editing
Nature Masterclasses
Research Solutions
Libraries & institutions
Librarian service & tools
Librarian portal
Open research
Recommend to library
Advertising & partnerships
Advertising
Partnerships & Services
Media kits
Branded
content
Professional development
Nature Awards
Nature Careers
Nature
Conferences
Regional websites
Nature Africa
Nature China
Nature India
Nature Japan
Nature Middle East
Privacy
Policy
Use
of cookies
Your privacy choices/Manage cookies
Legal
notice
Accessibility
statement
Terms & Conditions
Your US state privacy rights
© 2026 Springer Nature Limited
Close banner
Close
Sign up for the
Nature Briefing
newsletter — what matters in science, free to your inbox daily.
Email address
Sign up
I agree my information will be processed in accordance with the
Nature
and Springer Nature Limited
Privacy Policy
.
Close banner
Close
Get the most important science stories of the day, free in your inbox.
Sign up for Nature Briefing


================================================================================

FILE: biorxiv_2024.05.21.589311v1_row4.txt
PATH: data/preprocessed/individual/CM4AI/biorxiv_2024.05.21.589311v1_row4.txt
SIZE: 53795 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: CM4AI
Source ID: biorxiv_preprint
Source type: preprint
Source URL: https://www.biorxiv.org/content/10.1101/2024.05.21.589311v1
Raw file: data/raw/CM4AI/biorxiv_2024.05.21.589311v1_row4.pdf
--------------------------------------------------------------------------------
bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture
from Disease-Relevant Cell Lines
Timothy Clark1†, Jillian Mohan2, Leah Schaffer2, Kirsten Obernier3, Sadnan Al Manir1,
Christopher P Churas2, Amir Dailamy2, Yesh Doctor2, Antoine Forget3, Jan Niklas
Hansen4, Mengzhou Hu2, Joanna Lenkiewicz2, Maxwell Adam Levinson1, Charlotte
Marquez2, Sami Nourreddine2, Justin Niestroy1, Dexter Pratt2, Gege Qian2, Swathi
Thaker5, Jean-Christophe Bélisle-Pipon6†, Cynthia Brandt7†, Jake Chen5†, Ying Ding8†,
Samah Fodeh7†, Nevan Krogan3†, Emma Lundberg4†, Prashant Mali2†, Pamela Payne-
Foster9†, Sarah Ratcliffe1†, Vardit Ravitsky10†, Andrej Sali3†, Wade Schulz7†, Trey
Ideker2*

1 - University of Virginia; 2 - University of California San Diego; 3 - University of
California San Francisco; 4 - Stanford University; 5 - University of Alabama at
Birmingham; 6 - Simon Fraser University; 7 - Yale University; 8 - University of Texas at
Austin; 9 - University of Alabama; 10 - University of Montreal
* corresponding author – tideker@health.ucsd.edu; † co-corresponding authors

Abstract

This article describes the Cell Maps for Artificial Intelligence (CM4AI) project and its
goals, methods, standards, current datasets, software tools , status, and future
directions. CM4AI is the Functional Genomics Data Generation Project in the U.S.
National Institute of Health’s (NIH) Bridge2AI program. Its overarching mission is to
produce ethical, AI-ready datasets of cell architecture, inferred from multimodal data
collected for human cell lines, to enable transformative biomedical AI research.

Introduction

1. Goals of CM4AI

Cell Maps for Artificial Intelligence (CM4AI) is an NIH-funded Bridge to Artificial
Intelligence (Bridge2AI)1 Data Generation Project (DGP) in the area of Functional
Genomics. It consists of three pillars – Data, People, and Ethics – organized into six
modules (Data: Data Acquisition, Tools, and Standards; People: Skills and Workforce
Development, and Teaming; and Ethics).

CM4AI’s objective is to deliver machine-readable hierarchical maps of cell architecture
as AI-Ready data produced from multimodal interrogation of 100 chromatin modifiers
and 100 metabolic enzymes involved in cancer, neuropsychiatric, and cardiac disorders
in disease-relevant cell lines under perturbed and unperturbed conditions, utilizing state-
of-the-art mass spectrometry based proteomics, spatial proteomics / cell imaging, and
genetic perturbations using CRISPR. The cell lines currently under investigation in
CM4AI are treated (with paclitaxel and vorinostat) versus untreated MDA-MB-468
breast cancer cell lines; and undifferentiated vs.  neurons and cardiomyocytes
generated from NIH-iPCS-1 KOLF2.1J induced pluripotent stem cells (IPSCs). CM4AI’s
software pipeline and AI-readiness framework provide important and reusable enabling
capabilities for this work.

CM4AI input data streams are generated using immunofluorescence (IF) subcellular
microscopy for spatial proteomics data; affinity purification mass spectroscopy (AP-MS)

1



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

and size exclusion mass spectroscopy (SEC-MS) methods for protein-protein
interaction (PPI) data; and single-cell CRISPR-Cas perturbation screens by cell type.
Input data streams are integrated via the multi-scale integrated cell (MuSIC) software
pipeline employing deep learning models and community detection algorithms2, and
output cell maps are packaged with provenance graphs and rich metadata as AI-Ready
datasets in RO-Crate format3,4 using an extended, client-server version of the
FAIRSCAPE framework5.

At a strategic level, CM4AI – as do other Bridge2AI DGPs – aims to ethically leverage
the most state-of-the art computational and experimental tools in biomedical science to
enable a new generation of  transformative artificial intelligence research to benefit
humanity. Its datasets and software are available to the research community under
minimally restrictive terms consistent with ethical guidelines, researcher attribution, and
research integrity.

2. What is a Cell Map?

Cell maps are hierarchical directed acyclic graphs (DAG), where each node represents
an assembly of proteins in proximity at a given scale, spanning from large assemblies
representing cell compartments (e.g. nucleus, mitochondria) to small assemblies of
proteins in close proximity (e.g. protein complexes or subunits). The data streams for
constructing these maps are generated from a panel of perturbed and unperturbed cell
lines, including treated and untreated cancer cell lines as well as differentiated and
naive induced pluripotent stem cells (iPSCs). Cell maps provide a foundation for
downstream applications in human genomics, including interpretation of genetic variants
and mutations. As an example, they have been used in AI tools for “visible machine
learning” or “visible neural networks” (VNNs) to interrogate how protein assemblies in
the cell affect cell-level phenotype predictions6–10.

In the CM4AI cell mapping process, affinity purification-mass spectrometry (AP-MS) and
size-exclusion mass spectroscopy (SEC-MS) techniques generate protein interaction
networks, while immunofluorescence (IF) staining and high-resolution microscopy
reveal protein localization and distribution within human cells11.

Cell maps are produced by integrating these localization and interaction data using self-
supervised embedding approaches from deep learning, followed by algorithmic
community detection at multiple resolutions to generate a hierarchy of protein
assemblies12. Cell maps are shared via the Network Data Exchange (NDEx)13  and can
be visualized in a web browser or accessed via tools such as Cytoscape14, HiView15,
and the Python ndex2 library16.
An example cell map from the Multi-Scale Integrated Cell (MuSIC)2 data set visualized
in HiView is shown in Figure 1.

3. What are Ethical, FAIR, AI-Ready Biomedical Data?

As defined by CM4AI, AI-Ready biomedical data are fully characterized FAIR data of
known provenance, which can be ethically and reliably processed by AI applications;
whose models and software are available and well-described for validation and re-use;
and whose predictions may be fully explained and interpreted to the user as needed.
CM4AI data are distinctive within Bridge2AI in that they are non-clinical (from tissue

2



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

Figure 1 – HiView visualization of a hierarchical cell map as a circle-packing diagram. (A)
A unified cell map of the interaction network in HEK293 cells generated by integrating
immunofluorescence images and affinity purification interactions for 661 proteins.2 The
hierarchical cell map contains 10 layers of depth representing communities at multiple
resolutions, such that communities at smaller distance (lighter circles) are nested inside larger
communities (darker circles) similar to physical compartments within a cell (the “root” structure).
HiView can be used to zoom in on nested communities, such as (B) the nuclear splicing speckle
(depth layer 4) and (C) the spliceosomal complex family (depth layer 5), to begin to resolve
smaller interaction communities and even individual proteins. For each community, HiView
provides the underlying interaction data that support it. For example, (D) shows the protein-
protein interaction network within the nuclear splicing speckle community, where yellow edges
represent similarities in the embedding from both protein image and affinity purification data.
HiView also provides a list of proteins within each community. For example, (E) shows the list of
proteins (15 total) located within the spliceosomal complex family.

cultures) and are considered to be de-identified as they cannot be matched, with current
knowledge, to a human subject.

AI-Readiness has been defined in multiple ways in the literature. Regardless of the
specific definitions adopted, in all cases significant requirements are placed on the
datasets used to train AI/ML models and on datasets analyzed using these AI/ML
models. Datasets that meet these requirements are called “AI-Ready”. We define the
following criteria for purposes of this discussion:

(cid:0)

  FAIR – Datasets, and the software used to prepare them for AI/ML analysis,

must comply with the FAIR Principles which have been outlined for data and for
software; interoperability is a particular requirement of AI-Readiness17. FAIRness
is an NIH requirement and is consistent with best current scholarly practice.

3




bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

(cid:0)

(cid:0)

(cid:0)

  Provenanced – Provenance Graphs of computations, data, software, and
models used to prepare the dataset must be available as metadata18–20.

(cid:0)
  Characterized – Complete schemas, validation procedures, and data sheets21
for all datasets in the provenance graph; and model cards22 for all models used
to prepare the dataset, are resolvable from the dataset’s metadata.

  Explainable – Dataset must be fully provenanced (see above) and statistically

characterized within any limitations and constraints, including ethical
considerations and limitations; and this characterization is available in the
metadata23–26;

  Ethical – The dataset was obtained, validated, documented, licensed, and

distributed ethically27. Ethical conditions will include but are not limited to: ethical
treatment of subjects including protection of human subjects data and ethical
treatment of animals; proper non-biasing data analysis and scientific integrity of
methods, conclusions, and publications; full documentation of scientific methods
and reagents; and licensing and distribution of data, software, and models with
openness to the scientific community for reuse consistent with protection of
human subjects privacy and adhering to responsible research conduct;
moreover, this requirement involves anticipating potential applications of the data
and derived AI/ML systems, integrating them into data governance frameworks to
optimize benefits and mitigate potential risks, and promoting their use for the
collective good.

These criteria are interdependent. For example, ethical validation capability depends
upon explainability, explainability depends upon provenance, and so forth.

CM4AI uses an expanded version of the FAIRSCAPE framework to establish a basis for
AI-readiness. As we progress through the project, our AI-readiness features will become
ever more complete. At present they comprise: (a) establishing FAIRness including rich
metadata and persistent, globally unique identifiers; (b) computing a machine-readable
provenance graph, resolvable as metadata for all results, including inputs,
computations, software, and outputs; and (c) characterizing and validating all datasets
and data elements, and mapping data elements to public ontology vocabularies, where
appropriate, using JSON-Schema mini-data-dictionary descriptions resolvable from the
provenance graph metadata. Further AI-Readiness capabilities on our near-term
research agenda are described in the Methods section.

Methods

1. Cell Lines
CM4AI has elected to use the cancer cell line MDA-MB-46828 (+/- treatment with
paclitaxel and vorinostat) and the iPSC line KOLF2.1J29 (+/- neuronal and
cardiomyocytic differentiation), both of which have been ethically sourced.

MDA-MB-468 (RRID:CVCL_0419) is a triple negative breast cancer cell line established
from a metastatic site pleural effusion of a 51-year-old black female with a metastatic
mammary adenocarcinoma, available from ATCC. This cell line has been extensively
used to study triple-negative breast cancer and is well characterized with data such as
transcriptomic, mutational profile and whole-genome sequencing available.

4



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

The KOLF2.1J (RRID:CVCL_B5P3) iPSCs cell line was derived from a healthy male
Northern European donor, available from the Human Induced Pluripotent Stem Cells
Initiative (HipSci) resource30. It is available for access by non-for-profit organizations via
a simple MTA.

2. Data Acquisition

2.1 Protein-Protein Interaction

To map protein-protein interactions (PPIs) of 100 chromatin regulators under different
conditions (cancer: no treatment, paclitaxel or vorinostat; iPSC: undifferentiated,
neurons and cardiomyocytes), we employed AP-MS on endogenously tagged cell lines
and size-exclusion chromatography coupled to mass spectrometry (SEC-MS) for
proteome-wide complex/interaction mapping as two orthogonal mass spectrometry-
based approaches. We have endogenously tagged 17 genes in MDA-MB-468 and
acquired AP-MS data under three conditions (untreated, paclitaxel and vorinostat
treated) . We are currently in the process of tagging 34 additional genes inMDA-MB-468
cells . In addition, we have performed SEC-MS on MDA-MB-468 cells under three
conditions (untreated, paclitaxel and vorinostat treated), which enabled us to identify
72/100 chromatin modifiers, with 52 of them being integral components of protein
complexes. In addition to chromatin modifiers of interest, SEC-MS allowed us to map
complexes in MDA-MB-468 cells and investigate the impact of treatment proteome-
wide. We detected PPI profiles of over a thousand complexes in MDA-MB468 cells, with
thousands of proteins exhibiting differential elution profiles between control and
paclitaxel or vorinostat treated cells. SEC-MS has also been performed in KOLF2 iPSCs
and derivatives, uncovering more than 700 protein complexes in both parental iPSCs
and differentiated neurons.

2.2 Spatial proteomics mapping

For Year 1, we proposed to map the spatial subcellular organization of key chromatin
modifiers, their interactors and key signaling molecules involved in cancer using the
Human Protein Atlas resource of antibodies. We established automated fixation and
permeabilization protocols for the pipetting robot for MDA-MB-468 and KOLF2 cell lines,
and completed spatial proteomics mapping of 100 chromatin regulators in the MDA-MB-
468 cells (+/- paclitaxel or vorinostat), with another 500 proteins pending significant hits
from the genetic perturbations and PPI results. The first set of images was released as
input for the CM4AI Tools Module’s MuSIC pipeline and processed to RO-Crate
outputs.

2.3 Genetic perturbation mapping

For Year 1, we performed single-cell CRISPR screens perturbing 100 chromatin
regulators in the MDA-MB-468 cells under 3 conditions (no treatment, paclitaxel or
vorinostat) and KOLF2 iPSC in the undifferentiated state. We have designed a CRISPR
lentiviral library targeting 100 chromatin factors with 6 guide RNAs per gene. We
generated and characterized MDA-MB-468 and KOLF2 CRISPR lines expressing
inducible dCas9. Single-cell CRISPR screens of 100 chromatin regulators (+/– drug) in
MDA-MB-468 cells and in undifferentiated KOLF2.1J iPSCs were done using the 10x
Genomics 3’HT kit. The resulting data are currently being QCed.

5



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

3. Tools: Cell Map Data Integration Pipeline and VNN Tools

The Tools Module of CM4AI is responsible for producing and maintaining a dataset
integration and map production system, the Multi-Scale Integrated Cell (MuSIC)
pipeline, which produces integrated cell maps from the multiple, multi-modal input data
streams. The MuSIC pipeline is divided into multiple segments, each of which calls the
FAIRSCAPE client package to validate inputs and create the output RO-Crate package
for that segment which, in turn, is ready for dispatch to the FAIRSCAPE server for PID
creation and provenance graph entailment computation. FAIRSCAPE-generated PIDs
resolve to complete human- and machine-readable metadata including licensing
information, and provide links to the underlying datasets.

MuSIC pipeline segments and their functions are shown below and in Figure 2. The
steps are:

a)  Download PPI and Image Data

- PPI networks for the cell line and conditions are downloaded from the Krogan
Laboratory’s deposition archive and made available for further processing.

-  IF images for the cell line and conditions are downloaded from the Lundberg
Laboratory’s deposition archive and made available for further processing.

b)  Generate embeddings
- Downloaded PPI networks are processed using the node2vec deep learning model31
to reduce dimensionality, producing a PPI embedding that contains information about
the protein’s interactions.
- IF images are processed using a Human Protein Atlas deep learning model32 to
reduce dimensionality, producing an image embedding that contains information about
protein localization.

c)  Co-Embedding

- The PPI and image embeddings are integrated to obtain a co-embedding for each
protein using contrastive deep learning. This integration model learns co-embeddings
such that the original embeddings can be reconstructed with minimal information loss
and proteins with similar PPI and image embeddings have similar co-embeddings33.

d)  Protein Community Detection and Hierarchy Creation

- Community detection is performed based on the all-by-all similarities of pairs of
proteins in the co-embedding space. The resulting hierarchy of protein communities is
output as the final cell map.

e)  Hierarchy evaluation

- Cell maps are annotated in a two pronged approach. First, cell maps are aligned to
known protein function and pathway resources, including the Gene Ontology (GO)34 and
Reactome35, to determine protein assemblies in the cell map with high overlap with
known cell biology. Second, assemblies in the cell maps are annotated using a large
language model (LLM) approach that we developed to name sets of proteins and assign
a name confidence score.

6



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

Figure 2 – Tools Module: Data Integration and Map Production Pipeline

With the Standards and Data Acquisition Modules, we developed an alpha release of
software that interfaces the MuSIC pipeline with the FAIRSCAPE infrastructure and
data repository. The system includes data dictionaries, formatting standards, and a
FAIRSCAPE metadata and provenance API, implemented as a command line software
package.

f) Integrative structure modeling

We are currently exploring the feasibility of integrative structure modeling of the MuSIC
communities. The resulting structural models will increase our understanding of the
MuSIC communities and help with planning future experiments.

To begin, we developed a bioinformatics pipeline for annotating MuSIC communities by
the available structural information about the community members and their
interactions. This structural information includes data from the Protein Data Bank
(PDB)36, AlphaFold Protein Structure Database (AlphaFoldDB)37, crosslinking mass
spectrometry, and prediction of disordered sequence segments. We then ranked the
MuSIC communities by the amount of available structural information, serving as a
proxy for the feasibility of integrative structure modeling. Finally, we performed
integrative structure modeling of a few top ranked communities. This process uses our
standard integrative structure modeling framework, which proceeds through the
following four stages38–40: (i) gathering input information, (ii) representing subunits and
translating data into spatial restraints, (iii) configurational sampling to produce an
ensemble of models that satisfy the restraints, and (iv) analyzing and validating input
information and models. The modeling protocol was scripted using the Python Modeling
Interface package, a library for modeling macromolecular complexes based on our
open-source Integrative Modeling Platform (IMP) package version 2.18
(https://integrativemodeling.org).

7






bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

4. Standards: AI-Readiness Packaging

The Standards Module of CM4AI is responsible for AI-Readiness packaging of data,
E
software, metadata and provenance graphs. This task is performed by the FAIRSCAPE
framework. The integration of FAIRSCAPE with CM4AI Data Acquisition and the Tools
data integration pipeline is shown in Figure 3.

Figure 3 – FAIRSCAPE AI-Readiness framework as applied to CM4AI datasets.

FAIRSCAPE consists of a client-side Python3 application, called either from the
command line or as a set of Python functions by the Tools Module’s data integration
pipeline, and a server application, also in Python3, which completes the packaging.

The client-side package, FAIRSCAPE-CLI, is called when any computation or coherent
set of computations in the pipeline is completed, and it is passed metadata which
defines schemas in JSON-Schema41 for the datasets in the computational unit, as well
as the inputs, computations, software, models, and outputs. FAIRSCAPE-CLI creates
an RO-Crate package with the datasets, metadata, and software – or resolvable
references to these components – and unique stubs for identifier creation on each of
these components.

The RO-Crates are then sent to the FAIRSCAPE server where they are registered and
assigned persistent, resolvable, globally unique persistent identifiers (PIDs). The RO-
Crates are then decomposed into their individual components – datasets, models,
software – which are also registered and assigned PIDs. The PID system currently in
use is the ARK scheme42 – with DOIs a future feature as supplementary PIDs for final-
state publishable work.

Lastly, the server computes end-to-end entailments on each RO-Crate’s provenance as
s
expressed in the EVI Evidence Graph Ontology and links them together where possible.
e.
g
A graphical view of an evidence graph presented as part of the human-readable landing
in
page for a CM4AI RO-Crate package, is shown in Figure 4. Alternate views serialized in
JSON-LD and in RDF-XML are also provided on the landing page. The complete RO-

8





bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

Crate is provided as Supplemental Data, and all RO-Crates for the current release are
available in the University of

Figure 4. Evidence graph provenance visualization from human-readable landing page.
This landing page is displayed on resolution of the RO-Crate’s ARK globally unique persistent
ID. It was generated on an RO-Crate package from CM4AI’s 0.5 alpha data release. Landing
page navigation tabs allow full metadata on the package to be displayed in human readable
form, or alternatively in JSON-LD serialization.

Virginia’s LibraData data archive, an instance of Harvard’s Dataverse. Links to the
archived RO-Crates are provided in the Data and Software Availability Statement.

PIDs generated by the server resolve to machine- and human-readable landing pages
containing the metadata, expressed in the JSON-LD graph language using vocabularies
from Schema.org, EVI, and other well-defined public ontologies. The complete RO-
Crate dataset is attached as Supplemental Data.

Both the client-side and sever-side FAIRSCAPE packages are PIP-installable and are
freely available, under the MIT open-source license (see Data and Software Availability
Statement).

5. Teaming

The CM4AI project's Teaming Module supports integration and expansion of technical /
scientific expertise in CM4AI, emphasizing ethical considerations and diverse
interdisciplinary perspectives. It facilitates effective communication and collaboration
among CM4AI investigators and personnel, of varied geographic, disciplinary, and
cultural backgrounds.

Teaming supports open dissemination of CM4AI-generated data, maps, and tools via
the CM4AI web portal (http://www.cm4ai.org), using the U-BRITE platform43, to support
open and trustworthy data management and sharing, ensuring the broad dissemination
of the datasets, cell maps, and tools developed. This effort aligns with our commitment

9




bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

to FAIR principles and high-performance access, making our deliverables both
accessible and useful to users.

Teaming has constructed an extensive scholarly network called a talent graph, based
initially on publications of CM4AI investigators, which links to each person’s peer-
reviewed research publications. Knowledge graphs involving this extended community
of experts, publications, data sets, and tools are captured through this effort.

Collaboration with the NIH Common Fund Data Ecosystem (CFDE) for data curation
helps integrate CM4AI contributions into the broader scientific community, and
engagement with the Bridge2AI Center helps us identify novel channels for
broadcasting our achievements, ensuring the outreach of our data, maps, and tools to
those who can utilize them fully.

The Teaming module fosters a culture of open science and community involvement with
mailing lists, shared online collaborating folders, shared documents, and remote
messaging and collaboration software within CM4AI, across the Bridge2AI program,
and in the broader scientific community, to promote an ethical, inclusive, and
collaborative scientific environment. In collaboration with the Skills and Workforce
Development module, it contributes significantly to enhancing the biomedical AI/ML
workforce and fostering innovation in the field.

6. Skills and Workforce Development

The Skills and Workforce Development module works to enhance the biomedical AI/ML
workforce by recruiting and training a diverse group of participants in the scientific
approach, data standards, and tools that can be used to accelerate innovation from
CM4AI. To recruit and prepare a biomedical workforce that is adept at both the data and
life sciences, new training that will allow researchers to leverage the data sets produced
by Bridge2AI and other programs is needed.  To achieve this goal, we aim to broadly
distribute the data and tools developed as part of the CM4AI project and provide broad
training in these components. We are leveraging a multimodal approach to content
delivery which includes asynchronous virtual training, hosted virtual events, and an in-
person internship being co-hosted at two of our sites.

A key focus of the Skills and Workforce Development module is to cultivate a diverse
biomedical AI/ML workforce.  Our recruitment approach has included communication
with investigators involved in the Bridge2AI program, integration with trainees at
organizations involved with Bridge2AI, and community outreach to other academic
organizations with an emphasis on those serving underrepresented communities. For
our inaugural virtual CodeFest event, we provided training to a total of 38 registrants
where 40% of the attendees identified as female and 30% came from underrepresented
communities. We have also developed a Diversity, Equity, and Inclusion (DEI)
Committee with representation from all Cores to direct initiatives which has at its core to
partner with historically black colleges and universities (Meharry Medical and
Morehouse School of Medicine) to increase the number of learners in the AI/ML field at
all levels:  graduate students:  masters and PhD, postdoctoral, and junior investigators.

Our in-person internship will be hosted at Yale University and the University of
California San Diego, with students jointly working on a final project between the two

10



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

sites and supervised by faculty from both institutions. This internship will provide
participants with the opportunity to learn 1) the impact of cell maps on biomedical
research applications, 2) how to interact with the data and standards used by CM4AI, 3)
how to develop a visible neural network using cell maps, and 4) how to integrate
external data into an existing cell map for the ethical development of AI/ML applications.

7. Ethics

The Ethics team worked closely with Data Acquisition, Tools, and Standards Modules to
develop a plan for ethical preparation, licensing, dissemination, and Data Access
supervision of CM4AI datasets. This plan balances the openness of the data produced,
with protection of the intellectual property, and supervision and monitoring of potential
efforts at commercialization. In addition to CM4AI data governance initiatives,
comprehensive guidelines and best practices are being established for the ethical
development of AI systems leveraging this data. The Ethics team employs a
methodology inspired by Value-Sensitive Design (VSD), facilitating the creation of
normative standards and expectations aimed at ensuring the responsible design of
datasets and AI technologies, in alignment with core values. This entails the Ethics
team engaging in conceptual efforts to craft an axiological repository—a repository of
values—to guide the articulation of design standards and expectations. Further
advancing this work, the Ethics team is instrumental in formulating the CM4AI Life
Cycle, a framework designed to elucidate governance milestones throughout the data
generation and AI system development processes. This lifecycle framework
underscores the commitment to embedding ethical considerations at every phase,
ensuring accountability and value alignment from inception through deployment.

Recognizing the importance of diverse perspectives in shaping ethical guidelines, the
Ethics team is dedicated to conducting mixed empirical research, combining qualitative
and quantitative methods. This research is critical for capturing a broad spectrum of
community insights, thus enriching the development of guidelines and best practices
with a multifaceted understanding of stakeholder values and concerns. Moreover, the
team is set to develop a suite of tools aimed at enhancing awareness and
comprehension of the ethical, legal, and social ramifications associated with CM4AI
data and the consequent AI systems. These tools are envisioned to bolster ethical
decision-making, providing vital support for navigating the complex landscape of AI
ethics. This comprehensive approach not only aims to elevate the standards of AI
development within CM4AI but also seeks to contribute to the broader discourse on
responsible AI, promoting a culture of ethical integrity, inclusivity, and societal benefit.

Results and Progress to Date

The CM4AI project has made significant progress with efforts by each of our six
modules (Data, Tools, Standards, Ethics, Skills, Teaming) relevant to each of the three
program pillars: Data – composed of the Data, Tools, and Standards Modules; People –
composed of the Skills and Teaming Modules; and Ethics – the Ethics Module.

The Data Module completed parallel maps of protein-protein interactions, protein
subcellular distributions, and transcriptional states associated with a near-
comprehensive set of chromatin regulators encoded by the human genome (Pillar:
Data). All of these data were generated in ethically sourced breast cancer cells (Pillar:

11



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

Ethics) and will be extended to ethically sourced human induced pluripotent stem cells
(iPS) in Year 2.

The Tools Module developed an early (alpha) release of the Multi-Scale Integrated Cell
(MuSIC) toolkit for assembling large-scale maps of human cells from fusion of protein
interaction and imaging data, including a highly automated pipeline that integrates with
Standards FAIRSCAPE toolkits. We expect to submit a mature version for publication
as a Protocols paper in early fall

The Standards Module established FAIR AI-Readiness metadata and validation toolkits,
which have been used to package and release our cell-line perturbation and
measurement data.

The Ethics Pillar established guidelines addressing socio-ethical challenges, positioning
us to now embark on stakeholder engagement and explainable AI development.
Preliminary work has been done to develop a first version of the Value Repository to
support the establishment of these guidelines44,45. The Ethics Pillar also supported the
work initiated by the Bridge2AI Ethics Working Group for supporting the Bridge2AI Data
Sharing and Dissemination Working Group in developing a Code of Conduct. The Code
of Conduct outlines basic requirements for participants in the Bridge2AI Open House to
access B2AI data (including CM4AI datasets), to attest and commit to it prior to
accessing the data.

The Workforce Development Module identified gaps in training of the functional
genomic and biomedical AI/ML workforce; and led the organization of a highly
successful CodeFest in March 2024, in collaboration with the Tools, Standards, Data,
and Teaming Modules. The CodeFest was attended by both within-project and external
participants. Our Workforce Development Module is now transitioning towards
collaborative content creation and delivery (Pillar: People).

Finally, the Teaming Module facilitated cross-disciplinary collaboration, completed the
open-access CM4AI web portal, and initiated a Diversity Equity and Inclusion (DEI)
Committee (Pillar: People). Beyond these specific activities, CM4AI has been actively
engaged in activities of the greater Bridge2AI Consortium, including a face-to-face
meeting in Pentagon City, Virginia; the overall Bridge2AI Steering Committee and
workgroups; and submission of numerous auxiliary concepts for supplemental funding
and a U-24 submission for cross-cutting integration across projects in the NIH Common
Fund Data Ecosystem.

Initial Datasets, Tools and Standards associated with all three parallel mapping
platforms have been released and are publicly available at our CM4AI web portal30.
These data were also made available at our CodeFest in March 2024, accompanied by
detailed tutorials explaining how the data were derived and integrated, and providing
production ready software for producing the datasets. This first data release was made
possible by intensive cross-module collaboration by dozens of staff as well as end-to-
end operation of all aspects of our CM4AI platform, from data generation to data
analysis by advanced toolsets to FAIR-compliant ethical AI-readiness standards for data
access and distribution.

12



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

Data packages from CM4AI including links to software used in preparing the data, input
datasets, dataset schemas, and deep provenance graphs, are available on the CM4AI
web portal and will soon be available on the Dataverse NIH approved generalist
repository.

All CM4AI integration and packaging software is freely available and licensed under
nonrestrictive open-source licenses.  Data packages are licensed under Creative
Commons CC-BY-NC-SA 4.0 license terms31, which include a requirement for
attribution to the dataset authors and the copyright holding institutions, and citation of
this article.  Commercial use requires a separate license negotiation with the copyright
holder (UCSD, Stanford, and/or UCSF depending upon the specific data package in
question). A Data Access Committee will supervise ethical matters related to dataset
distribution and potential dual licensing for commercial use.

Conclusions

The Bridge2AI Functional Genomics project, “Cell Maps for Artificial Intelligence”
(CM4AI), is a continuing effort which has already published valuable and re-usable
datasets and software, advanced the notions of ethical AI-readiness, developed a
strong teaming approach within the project including a shared portal, and provided
hands-on training to the research community in use of these tools.

This article presents the work as a whole in preprint form and will be continuously
updated as CM4AI progresses, with citations to detailed publications in each
contributing area.

With these measures, it is our intention that CM4AI become a highly transformative
platform for ethical, explainable, and interpretable biomedical AI, deepen our
understanding of processes of human disease and health, help to train a new
generation of researchers, enable development of novel cures, and assist researchers
and clinicians in equitably improving human lives.

Data and Software Availability Statement

The most recent data and metadata produced by CM4AI are licensed for reuse under
the Creative Commons Attribution Non-Commercial Share-Alike International 4.0
License  (https://creativecommons.org/licenses/by-nc-sa/4.0/) and are available in
LibraData, the University of Virginia’s archival data repository:

Clark T, Mohan J, Schaffer L, Obernier K, et al. Cell Maps for Artificial
Intelligence - Data Release", https://doi.org/10.18130/V3/DXWOS5, University of
Virginia Dataverse, V1

LibraData is the University of Virginia’s instance of Dataverse, an NIH-approved
generalist repository.

Attribution requirements for these datasets include attribution to the copyright holders
and the Cell Maps for Artificial Intelligence project, as referenced in the datasets, and
citation of the present article:

13



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

Clark T, Mohan J, Schaffer L, Obernier K, Al Manir S, Churas CP, Dailamy A,
Doctor Y, Forget A, Hansen JN, Hu M, Lenkiewicz J, Levinson MA, Marquez C,
Nourreddine S, Niestroy J, Pratt D, Qian G, Thaker S, Bélisle-Pipon J-C, Brandt
C, Chen J, Ding Y, Fodeh S, Krogan N, Lundberg E, Mali P, Payne-Foster P,
Ratcliffe S, Ravitsky V, Sali A, Schulz W, Ideker T. Cell Maps for Artificial
 Intelligence: AI-Ready Maps of Human Cell Architecture from  Disease-Relevant
Cell Lines. BioRXiv, May 2024.

Copyright © 2024 to these datasets is held by the Regents of the University of
California, except where otherwise indicated. Raw image data for spatial proteomics is
copyright © 2024 by The Board of Trustees of the Leland Stanford Junior University.

Software comprising the Tools data integration pipeline is available under BSD-3 open
source license in GitHub (for alpha level tools) or in the Zenodo long-term archive (for
production-ready tools). Links to these tools are packaged in RO-Crates as part of
CM4AI provenance representations.

The FAIRSCAPE AI-readiness framework is described with a tutorial and instructions
for installation at https://fairscape.github.io. Copyright © 2024 to this software is held by
the Rector and Board of Visitors of the University of Virginia. It is available under open
source MIT License.

The open source Integrative Modeling Platform (IMP) package is available at
http://integrativemodeling.org.

Software packages will be versioned as the project progresses, and the version used to
produce each dataset will be referenced in the dataset’s metadata.

Acknowledgements

This work was funded by the National Institutes of Health under awards
1OT2OD032742-01 (Bridge2AI Functional Genomics) and 5U54HG012513-02
(Bridge2AI Bridge Center), and by the Frederick Thomas Fund of the University of
Virginia.

We also acknowledge the helpful assistance of members of the other Bridge2AI
components, which helped to make this work possible, including but not limited to the
Bridge2AI Standards Working Group; the Bridge2AI Data Sharing and Dissemination
Working Group; and members of the Bridge2AI CHoRUS Critical Care Medicine Data
Generation Project.

REFERENCES

1.  National Institutes of Health. Bridge to Artificial Intelligence (Bridge2AI). Published

online February 7, 2023. Accessed February 9, 2023.
https://commonfund.nih.gov/bridge2ai

2.  Qin Y, Huttlin EL, Winsnes CF, et al. A multi-scale map of cell structure fusing

protein images and interactions. Nature. 2021;600(7889):536-542.
doi:10.1038/s41586-021-04115-9

14



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

3.  RO-Crate Community. Research Object Crate (RO-Crate). Published online October

26, 2023. https://www.researchobject.org/ro-crate/

4.  Soiland-Reyes S, Sefton P, Crosas M, et al. Packaging research artefacts with RO-

Crate. Published online August 13, 2021. doi:10.5281/ZENODO.5146227

5.  Levinson MA, Niestroy J, Al Manir S, et al. FAIRSCAPE: a Framework for FAIR and

Reproducible Biomedical Analytics. Neuroinform. 2022;20(1):187-202.
doi:10.1007/s12021-021-09529-4

6.  Ma J, Yu MK, Fong S, et al. Using deep learning to model the hierarchical structure
and function of a cell. Nat Methods. 2018;15(4):290-298. doi:10.1038/nmeth.4627

7.  Yu MK, Ma J, Fisher J, Kreisberg JF, Raphael BJ, Ideker T. Visible Machine

Learning for Biomedicine. Cell. 2018;173(7):1562-1565.
doi:10.1016/j.cell.2018.05.056

8.  Kuenzi BM, Park J, Fong SH, et al. Predicting Drug Response and Synergy Using a
Deep Learning Model of Human Cancer Cells. Cancer Cell. 2020;38(5):672-684.e6.
doi:10.1016/j.ccell.2020.09.014

9.  Zheng F, Kelly MR, Ramms DJ, et al. Interpretation of cancer mutations using a

multiscale map of protein systems. Science. 2021;374(6563):eabf3067.
doi:10.1126/science.abf3067

10. Ma J, Fong SH, Luo Y, et al. Few-shot learning creates predictive models of drug
response that translate from high-throughput screens to individual patients. Nat
Cancer. 2021;2(2):233-244. doi:10.1038/s43018-020-00169-2

11. Thul PJ, Åkesson L, Wiking M, et al. A subcellular map of the human proteome.

Science. 2017;356(6340):eaal3321. doi:10.1126/science.aal3321

12. Singhal A, Cao S, Churas C, et al. Multiscale community detection in Cytoscape.

Przytycka TM, ed. PLoS Comput Biol. 2020;16(10):e1008239.
doi:10.1371/journal.pcbi.1008239

13. Pratt D, Chen J, Pillich R, et al. NDEx 2.0: A Clearinghouse for Research on Cancer
Pathways. Cancer Research. 2017;77(21):e58-e61. doi:10.1158/0008-5472.CAN-
17-0606

14. Shannon P. Cytoscape: A Software Environment for Integrated Models of

Biomolecular Interaction Networks. Genome Research. 2003;13(11):2498-2504.
doi:10.1101/gr.1239303

15. Yu MK, Ma J, Ono K, et al. DDOT: A Swiss Army Knife for Investigating Data-Driven

Biological Ontologies. Cell Syst. 2019;8(3):267-273.e3.
doi:10.1016/j.cels.2019.02.003

15



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

16. Cytoscape Consortium. ndex2. Published online February 8, 2024.

https://ndex2.readthedocs.io/en/latest/?badge=latest

17. Wilkinson MD, Dumontier M, Aalbersberg IjJ, et al. The FAIR Guiding Principles for

scientific data management and stewardship. Sci Data. 2016;3:160018.
doi:10.1038/sdata.2016.18

18. Al Manir S, Niestroy J, Levinson MA, Clark T. Evidence Graphs: Supporting

Transparent and FAIR Computation, with Defeasible Reasoning on Data, Methods,
and Results. In: Glavic B, Braganholo V, Koop D, eds. Provenance and Annotation
of Data and Processes. Vol 12839. Lecture Notes in Computer Science. Springer
International Publishing; 2021:39-50. doi:10.1007/978-3-030-80960-7_3

19. Al Manir S, Niestroy J, Levinson M, Clark T. EVI: The Evidence Graph Ontology,

OWL 2 Vocabulary. Published online March 23, 2021.
https://doi.org/10.5281/zenodo.4630931

20. Lebo T, Sahoo S, McGuinness D, et al. PROV-O: The PROV Ontology W3C

Recommendation 30 April 2013. Published online 2013. http://www.w3.org/TR/prov-
o/

21. Gebru T, Morgenstern J, Vecchione B, et al. Datasheets for Datasets.

arXiv:180309010 [cs]. Published online March 19, 2020. Accessed August 4, 2021.
http://arxiv.org/abs/1803.09010

22. Mitchell M, Wu S, Zaldivar A, et al. Model Cards for Model Reporting. Proceedings
of the Conference on Fairness, Accountability, and Transparency. Published online
January 29, 2019:220-229. doi:10.1145/3287560.3287596

23. Arbelaez Ossa L, Starke G, Lorenzini G, Vogt JE, Shaw DM, Elger BS. Re-focusing

explainability in medicine. DIGITAL HEALTH. 2022;8:205520762210744.
doi:10.1177/20552076221074488

24. Combi C, Amico B, Bellazzi R, et al. A manifesto on explainability for artificial
intelligence in medicine. Artificial Intelligence in Medicine. 2022;133:102423.
doi:10.1016/j.artmed.2022.102423

25. Kundu S. AI in medicine must be explainable. Nat Med. 2021;27(8):1328-1328.

doi:10.1038/s41591-021-01461-z

26. Amann J, Blasimme A, Vayena E, Frey D, Madai VI, Precise4Q consortium.

Explainability for artificial intelligence in healthcare: a multidisciplinary perspective.
BMC Med Inform Decis Mak. 2020;20(1):310. doi:10.1186/s12911-020-01332-6

27. Floridi L. A Unified Framework of Ethical Principles for AI. In: The Ethics of Artificial

Intelligence. Oxford University Press.
https://doi.org/10.1093/oso/9780198883098.001.0001

16



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

28. Cailleau R, Olivé M, Cruciger QVJ. Long-term human breast carcinoma cell lines of

metastatic origin: Preliminary characterization. In Vitro. 1978;14(11):911-915.
doi:10.1007/BF02616120

29. Pantazis CB, Yang A, Lara E, et al. A reference human induced pluripotent stem cell

line for large-scale collaborative studies. Cell Stem Cell. 2022;29(12):1685-
1702.e22. doi:10.1016/j.stem.2022.11.004

30. Streeter I, Harrison PW, Faulconbridge A, et al. The human-induced pluripotent
stem cell initiative—data resources for cellular genetics. Nucleic Acids Res.
2017;45(D1):D691-D697. doi:10.1093/nar/gkw928

31. Grover A, Leskovec J. node2vec: Scalable Feature Learning for Networks. In:

Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge
Discovery and Data Mining. ACM; 2016:855-864. doi:10.1145/2939672.2939754

32. Le T, Winsnes CF, Axelsson U, et al. Analysis of the Human Protein Atlas Weakly
Supervised Single-Cell Classification competition. Nat Methods. 2022;19(10):1221-
1229. doi:10.1038/s41592-022-01606-z

33. Bao F, Deng Y, Wan S, et al. Integrative spatial analysis of cell morphologies and

transcriptional states with MUSE. Nat Biotechnol. 2022;40(8):1200-1209.
doi:10.1038/s41587-022-01251-z

34. Ashburner M, Ball C, Blake J, et al. Gene ontology: tool for the unification of biology.

The Gene Ontology Consortium. Nature Genetics. 2000;25(1):25-29.

35. Joshi-Tope G, Gillespie M, Vastrik I, et al. Reactome: a knowledgebase of biological

pathways. Nucleic Acids Res. 2005;33(Database issue):D428-32.
doi:10.1093/nar/gki072

36. Berman HM. The Protein Data Bank: a historical perspective. Acta crystallographica

Section A, Foundations of crystallography. 2008;64(Pt 1):88-95.
doi:10.1107/S0108767307035623

37. Varadi M, Anyango S, Deshpande M, et al. AlphaFold Protein Structure Database:
massively expanding the structural coverage of protein-sequence space with high-
accuracy models. Nucleic Acids Research. 2022;50(D1):D439-D444.
doi:10.1093/nar/gkab1061

38. Rout MP, Sali A. Principles for Integrative Structural Biology Studies. Cell.

2019;177(6):1384-1403. doi:10.1016/j.cell.2019.05.016

39. Sali A. From integrative structural biology to cell biology. Journal of Biological

Chemistry. 2021;296:100743. doi:10.1016/j.jbc.2021.100743

17



bioRxiv preprint

doi:

https://doi.org/10.1101/2024.05.21.589311

;

this version posted May 24, 2024.

The copyright holder for this preprint (which

was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made

available under a

CC-BY-NC-ND 4.0 International license
.

40. Russel D, Lasker K, Webb B, et al. Putting the pieces together: integrative modeling
platform software for structure determination of macromolecular assemblies. PLoS
Biol. 2012;10(1):e1001244. doi:10.1371/journal.pbio.1001244

41. Wright A, Andrews H, Hutton B, Dennis G. JSON Schema: A Media Type for
Describing JSON Documents. Internet Engineering Task Force (IETF); 2019.
https://datatracker.ietf.org/doc/html/draft-handrews-json-schema-02

42. Kunze J, Rodgers R. The ARK Identifier Scheme. Published online 2008.

https://escholarship.org/uc/item/9p9863nc

43. Cimino JJ, Liang WH, Wang J, et al. Empowering Team Science Across the
Translational Spectrum with the UAB Biomedical Research Infrastructure
Technology Enhancement (U-BRITE). In: 2020 IEEE 21st International Conference
on Information Reuse and Integration for Data Science (IRI). IEEE; 2020:194-200.
doi:10.1109/IRI49571.2020.00035

44. Victor G, Salem S, Bélisle-Pipon JC. A Scoping Review of Relevant Moral Values in

Health Sector AI Development. Published online 2023.
doi:10.17605/OSF.IO/BCVK3

45. Victor G, Bélisle-Pipon JC, Ravitsky V. Generative AI, Specific Moral Values: A

Closer Look at ChatGPT’s New Ethical Implications for Medical AI. The American
Journal of Bioethics. 2023;23(10):65-68. doi:10.1080/15265161.2023.2250311

18


================================================================================

FILE: reporter_nih_gov_project-details-11211616_row7.txt
PATH: data/preprocessed/individual/CM4AI/reporter_nih_gov_project-details-11211616_row7.txt
SIZE: 3260 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: CM4AI
Source ID: nih_reporter_project
Source type: NIH project page
Source URL: https://reporter.nih.gov/project-details/11211616
Raw file: data/raw/CM4AI/reporter_nih_gov_project-details-11211616_row7.txt
--------------------------------------------------------------------------------
NIH RePORTER Project
Source: https://reporter.nih.gov/project-details/11211616
Application ID: 11211616
Project number: 3OT2OD032742-01S2
Core project number: OT2OD032742
Title: Bridge2AI: Cell Maps for AI (CM4AI) Data Generation Project
Principal investigator: IDEKER, TREY
Organization: UNIVERSITY OF CALIFORNIA, SAN DIEGO
Fiscal year: 2025
Award amount: 5289382
Project start: 2022-09-01T00:00:00
Project end: 2026-08-31T00:00:00

As part of the NIH Common Fund’s Bridge2AI program, the CM4AI data generation
project seeks to map the spatiotemporal architecture of human cells and use these
maps toward the grand challenge of interpretable genotype-phenotype learning. In
genomics and precision medicine, machine learning models are often "black boxes,"
predicting phenotypes from genotypes without understanding the mechanisms by
which such translation occurs. To address this deficiency, project will launch a
coordinated effort involving three complementary mapping approaches – proteomic
mass spectrometry, cellular imaging, and genetic perturbation via CRISPR/Cas9 –
creating a library of large-scale maps of cellular structure/function across
demographic and disease contexts. These data will broadly stimulate research and
development in "visible" machine learning systems informed by multi-scale cell and
tissue architecture. In addition to data and tools, this project will implement a
standards data management approach based on FAIR access and software principles,
with deep provenance and replication packages for representation of cell maps and
their underlying datasets; initiate a research program in ethical AI, especially as it
relates to how maps will be used in genomic medicine and model interpretation; and
stimulate a diverse portfolio of training opportunities in the emerging field of
biomachine learning.

Machine learning (ML) models show great promise in analyzing the human genome
to make predictions, but the inner workings of these models are typically difficult-to-interpret "black boxes." To address this challenge, this Bridge2AI data generation
project will generate a resource of matched data and tools to enable the creation of
"visible" ML systems, which are not black boxes but are built directly on knowledge
maps of cell and tissue architecture.

Preferred terms:
Address;Architecture;Black Box;Bridge to Artificial Intelligence;CRISPR/Cas technology;Cells;Cellular Structures;Computer software;Data;Data Set;Disease;Funding;Generations;Genetic;Genomic medicine;Genotype;Human;Human Genome;Knowledge;Learning;Libraries;Machine Learning;Maps;Mass Spectrum Analysis;Modeling;Phenotype;Proteomics;Research;Resources;System;Text;Tissues;Translations;United States National Institutes of Health;cellular imaging;data management;data standards;machine learning model;precision medicine;programs;research and development;responsible artificial intelligence;spatiotemporal imaging;tool;training opportunity


================================================================================

FILE: cm4ai_org_row10.txt
PATH: data/preprocessed/individual/CM4AI/cm4ai_org_row10.txt
SIZE: 3544 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: CM4AI
Source ID: project_documentation
Source type: documentation
Source URL: https://cm4ai.org/
Raw file: data/raw/CM4AI/cm4ai_org_row10.html
--------------------------------------------------------------------------------
Cell Maps For AI (CM4AI) – Cell Maps For AI
Home
People
Project Overview
Tools
Data Acquisition
Teaming
Standards
Ethics
Skills and Workforce
Products
Data Releases
Tools
Product Documentation
Publications
Learning
Home
People
Project Overview
Tools
Data Acquisition
Teaming
Standards
Ethics
Skills and Workforce
Products
Data Releases
Tools
Product Documentation
Publications
Learning
Welcome, explore, discover.
Cell Maps For AI: Mapping Cellular Structure and Function
The
CM4AI
data generation project seeks to map the spatiotemporal architecture of human cells and use these maps toward the grand challenge of interpretable genotype-phenotype learning
Data Insights
Explore the numbers behind our innovative data generation project and see the impact we are making.
Data Releases
Protein Interactions
1,374
Immunofluorescent Images
53,788
Total Proteins Investigated
7,023
genes targeted
11,739
Data volume
21.4 TB
CM4AI Project Components
Tools
Read More
Data Acquisition
Read More
Teaming
Read More
Ethics
Read More
Standards
Read more
Skills & Workforce Development
Read More
Data Releases
CM4AI project is a coordinated effort involving three complementary mapping approaches – proteomic mass spectrometry, cellular imaging, and genetic perturbation via CRISPR/Cas9
Explore
Tools
CM4AI is developing tools to integrate multimodal data, with the goal of creating a library of large-scale maps of cellular structure and function across demographic and disease contexts
Discover
Publications
Check out our recent publications to learn more about how CM4AI is addressing the Bridge2AI’s Functional Genomics grand challenge
Read
Training
CM4AI offers a wealth of educational content and a robust hands-on training program to foster biomedical AI/ML skill development
Learn More
Mapping Approaches
Discover the different mapping approaches we use to generate valuable data for your research.
Learn More
Proteomic Mass Spectrometry
Our cutting-edge proteomic mass spectrometry technology allows for precise mapping of cellular structure and function.
Cellular Imaging
Through advanced cellular imaging techniques, we are able to reveal the intricate details of human cells.
Genetic Perturbation via CRISPR/Cas9
Our use of CRISPR/Cas9 technology allows for targeted genetic perturbation and further insights into cellular mechanisms.
CM4AI in Action
Through a growing portfolio of educational resources, CM4AI connects with the biomedical and AI/ML communities to advance collaboration and cultivate future talent.
LEARNING OPPORTUNITIES
CM4AI Institutions
Join the effort towards interpretable genotype-phenotype learning
Learn more about our project and the maps we are creating to revolutionize the fields of genomics and precision medicine, through the use of advanced technologies such as proteomic mass spectrometry, cellular imaging, and genetic perturbation via CRISPR/Cas9.
Learn More
Funding
NIH funding award number: 1OT2OD032742-01
Follow Us
Contact Us
Project-Related Inquiries:
Swathi Thaker, PhD, UAB, Program Manager
snthaker@uab.edu
Website Support & Updates:
Zhandos Sembay, UAB
zsembay8@uab.edu
© 2026 Cell Maps For AI (CM4AI)
Privacy Notice
Disclaimer
This repository is under review for potential modification in compliance with Administration directives.
Scroll Up


================================================================================

FILE: cm4ai_org_data-releases_row11.txt
PATH: data/preprocessed/individual/CM4AI/cm4ai_org_data-releases_row11.txt
SIZE: 4182 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: CM4AI
Source ID: data_release_documentation
Source type: documentation
Source URL: https://cm4ai.org/data-releases/
Raw file: data/raw/CM4AI/cm4ai_org_data-releases_row11.html
--------------------------------------------------------------------------------
Data Releases – Cell Maps For AI (CM4AI)
Home
People
Project Overview
Tools
Data Acquisition
Teaming
Standards
Ethics
Skills and Workforce
Products
Data Releases
Tools
Product Documentation
Publications
Learning
Home
People
Project Overview
Tools
Data Acquisition
Teaming
Standards
Ethics
Skills and Workforce
Products
Data Releases
Tools
Product Documentation
Publications
Learning
© 2019
Data Releases
Cell Maps for Artificial Intelligence (CM4AI) will deliver machine-readable hierarchical maps of cell architecture as AI-ready data, together with quarterly data releases of map-input data streams. CM4AI cell maps are produced from multimodal interrogation of chromatin modifiers, metabolic enzymes, and other proteins involved in cancer, neuropsychiatric, and cardiac disorders in disease-relevant cell lines under treated and untreated conditions, utilizing state-of-the-art mass spectrometry-based proteomics, spatial proteomics / cell imaging, and genetic perturbations via CRISPR/Cas9. CM4AI is a collaboration of UCSD, UCSF, Stanford, UVA, Yale, UT Austin, UA Birmingham, Simon Fraser University, and the Hastings Center, as part of the NIH Bridge2AI program.
Documentation
Data Insights
Explore the numbers behind our innovative data generation project and see the impact we are making.
Protein Interactions
1,374
Immunofluorescent Images
53,788
Total Proteins Investigated
7,023
genes targeted
11,739
Data volume
21.4 TB
CM4AI’s Flagship Datasets
We are generating deep multi-modal data for a shared list of proteins across all cell types and treatment conditions to achieve an unparallelled level of biological insight
Curated dataset for undifferentiated (parental) iPSCs and iPSC-derived neural progenitor cells (NPCs), neurons, and cardiomyocytes include:
• Perturb-seq of
>11,000 genes (whole-genome)
• SEC-MS capturing
~7,000 proteins
• IF images (
coming soon!
)
Curated datasets for TNBC cells under untreated (DMSO control), vorinostat, and paclitaxel conditions include:
•SEC-MS capturing
~7,000 proteins
•IF images for
523 proteins
•Perturb-seq of
200 genes
•AP-MS interactomes (
coming soon!
)
Our latest data release
Explore all
This repository is under review for potential modification in compliance with Administration directives.
June 2026 Data Release (Beta)
doi.org/10.18130/V3/HIGT4C
This dataset is the June 2026 Data Release of CM4AI, the Functional Genomics Grand Challenge in the NIH Bridge2AI program.
Released on:
June 17, 2025
Archive
May 2024 Data Release
March 2025 Data Release
June 2025 Data Release
October 2025 Data Release
Frequently Asked Questions
When will the next Data Release be available?
CM4AI releases data quarterly, ensuring continuous updates to the cell maps and input data streams.
What types of data does CM4AI provide?
CM4AI provides hierarchical cell maps derived from various data streams, including:
•
Mass spectrometry-based proteomics
•
Spatial proteomics and cell imaging
•
Genetic perturbations using CRISPR/Cas9
These datasets focus on chromatin modifiers, metabolic enzymes, and other proteins linked to cancer, neuropsychiatric, and cardiac disorders.
How are the data obtained and processed for the CM4AI project?
We use a combination of mapping techniques such as proteomic mass spectrometry, cellular imaging, and genetic perturbation via CRISPR/Cas9 to create a library of large-scale maps of cellular structure and function. These maps are then organized and made available for use in our Data Releases.
Funding
NIH funding award number: 1OT2OD032742-01
Follow Us
Contact Us
Project-Related Inquiries:
Swathi Thaker, PhD, UAB, Program Manager
snthaker@uab.edu
Website Support & Updates:
Zhandos Sembay, UAB
zsembay8@uab.edu
© 2026 Cell Maps For AI (CM4AI)
Privacy Notice
Disclaimer
This repository is under review for potential modification in compliance with Administration directives.
Scroll Up


================================================================================

FILE: creativecommons_org_licenses-by-nc-sa_row15.txt
PATH: data/preprocessed/individual/CM4AI/creativecommons_org_licenses-by-nc-sa_row15.txt
SIZE: 5747 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: CM4AI
Source ID: dataset_license
Source type: license
Source URL: https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en
Raw file: data/raw/CM4AI/creativecommons_org_licenses-by-nc-sa_row15.html
--------------------------------------------------------------------------------
Deed - Attribution-NonCommercial-ShareAlike 4.0 International - Creative
Commons
Skip to content
Languages available
aragonÃ©s
AzÉrbaycanca
Bahasa Indonesia
Basque
catalÃ
dansk
Deutsch
eesti
English
espaÃ±ol
Esperanto
franÃ§ais
frysk
Gaeilge
galego
Hrvatski
italiano
latvieÅ¡u
LietuviÅ¡kai
Magyar
Melayu
Nederlands
norsk
polski
PortuguÃªs
PortuguÃªs Brasileiro
RomÃ¢nÄ
Slovensky
SlovenÅ¡Äina
srpski (latinica)
suomi
svenska
TÃ¼rkÃ§e
Ãslenska
Äesky
ÎÎ»Î»Î·Î½Î¹ÎºÎ¬
Ð±ÐµÐ»Ð°ÑÑÑÐºÐ°Ñ
Ð±ÑÐ»Ð³Ð°ÑÑÐºÐ¸
Ð ÑÑÑÐºÐ¸Ð¹
Ð£ÐºÑÐ°ÑÐ½ÑÑÐºÐ°
Ø§ÙØ¹Ø±Ø¨ÙÙØ©
ÙØ§Ø±Ø³Û
à¤¹à¤¿à¤à¤¦à¥
à¦¬à¦¾à¦à¦²à¦¾
æ¥æ¬èª
ç®ä½ä¸­æ
ç¹é«ä¸­æ
íêµ­ì´
Creative Commons
Menu
Who We Are
Expand
Strategic Plan
Team
Governance
Opportunities
Annual Reports & Financials
History
Press
What We Do
Expand
Build
Open Infrastructure
Expand
CC Licenses
CC Signals
Public Domain
Chooser
FAQs
Implement
the Commons
Expand
Where CC Makes An Impact
Resources
Search the Commons
Engage
the People
Expand
Training + Webinars
Advocacy
Community
Events
Blog
Support Us
Expand
Make a Gift
Open Infrastructure Circle
Donor FAQ
Donate
Attribution-NonCommercial-ShareAlike 4.0 International
CC BY-NC-SA 4.0
Deed
Canonical URL
https://creativecommons.org/licenses/by-nc-sa/4.0/
See the legal code
You are free to:
Share
â copy and redistribute the material in any
medium or format
Adapt
â remix, transform, and build upon the
material
The licensor cannot revoke these freedoms as long as you follow
the license terms.
Under the following terms:
Attribution
â You must give
appropriate credit
, provide a link to the license, and
indicate if changes were made
. You may do so in any reasonable manner, but not in any way that
suggests the licensor endorses you or your use.
NonCommercial
â You may not use the material for
commercial purposes
.
ShareAlike
â If you remix, transform, or build
upon the material, you must distribute your contributions under
the
same license
as the original.
No additional restrictions
â You may not apply
legal terms or
technological measures
that legally restrict others from doing anything the license
permits.
Notices:
You do not have to comply with the license for elements of the
material in the public domain or where your use is permitted by an
applicable
exception or limitation
.
No warranties are given. The license may not give you all of the
permissions necessary for your intended use. For example, other
rights such as
publicity, privacy, or moral rights
may limit how you use the material.
Notice
This deed highlights only some of the key features and terms of the
actual license. It is not a license and has no legal value. You
should carefully review all of the terms and conditions of the
actual license before using the licensed material.
Creative Commons is not a law firm and does not provide legal
services. Distributing, displaying, or linking to this deed or the
license that it summarizes does not create a lawyer-client or any
other relationship.
Creative Commons is the nonprofit behind the open licenses and other
legal tools that allow creators to share their work. Our legal tools
are free to use.
Learn more about our work
Learn more about CC Licensing
Support our work
Use the license for your own material.
Licenses List
Public Domain List
Footnotes
return to reference
appropriate credit
â If supplied, you must
provide the name of the creator and attribution parties, a
copyright notice, a license notice, a disclaimer notice, and a
link to the material. CC licenses prior to Version 4.0 also
require you to provide the title of the material if supplied,
and may have other slight differences.
More info
return to reference
indicate if changes were made
â In 4.0, you
must indicate if you modified the material and retain an
indication of previous modifications. In 3.0 and earlier
license versions, the indication of changes is only required
if you create a derivative.
Marking guide
More info
return to reference
commercial purposes
â A commercial use is one
primarily intended for commercial advantage or monetary
compensation.
More info
return to reference
same license
â You may also use a license
listed as compatible at
https://creativecommons.org/compatiblelicenses
More info
return to reference
technological measures
â The license
prohibits application of effective technological measures,
defined with reference to Article 11 of the WIPO Copyright
Treaty.
More info
return to reference
exception or limitation
â The rights of users
under exceptions and limitations, such as fair use and fair
dealing, are not affected by the CC licenses.
More info
return to reference
publicity, privacy, or moral rights
â You may
need to get additional permissions before using the material
as you intend.
More info
Creative Commons
submit
Who we are
What we do
Blog
Support us
Store
Contact
Privacy
Policies
Terms
Contact Us
Creative Commons
PO Box 1866, Mountain View,
CA 94042
info@creativecommons.org
Bluesky
Mastodon
LinkedIn
Subscribe to our newsletter
Subscribe
Except where otherwise
noted
, content
on this site is licensed under a
Creative Commons Attribution 4.0 International license
. Icons by
Font Awesome
.


================================================================================

FILE: dataverse_10.18130_V3_B35XWX_row16.txt
PATH: data/preprocessed/individual/CM4AI/dataverse_10.18130_V3_B35XWX_row16.txt
SIZE: 27868 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: CM4AI
Source ID: march_2025_dataverse_release
Source type: historical data release
Source URL: https://dataverse.lib.virginia.edu/dataset.xhtml?persistentId=doi:10.18130/V3/B35XWX
Raw file: data/raw/CM4AI/dataverse_10.18130_V3_B35XWX_row16.html
--------------------------------------------------------------------------------
Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta) - Cell Maps for Artificial Intelligence
Skip to main content
Toggle navigation
Search
Search
About
User Guide
Support
Log In
Cell Maps for Artificial Intelligence
This collection is under review for potential modification in compliance with Administration directives.
University of Virginia Dataverse
>
LibraData: UVa's Scholarly Research
>
School of Medicine
>
Cell Maps for Artificial Intelligence
>
Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)
Version 1.4
Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Marquez C; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Polacco B; Pratt D; Qian G; Schaffer L; Sigaeva A; Thaker S; Zhang Y; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, "Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)",
https://doi.org/10.18130/V3/B35XWX
, University of Virginia Dataverse, V1
Cite Dataset
Download EndNote XML
Download RIS
Download BibTeX
View Styled Citation
Learn about
Data Citation Standards
.
Access Dataset
The dataset is too large to download. Please select the files you need from the files table.
Contact Owner
Share
Dataset Metrics
302 Downloads
Dataset Description
This dataset is the March 2025 Data Release of Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org), the Functional Genomics Grand Challenge in the NIH Bridge2AI program. This Beta release includes perturb-seq data in undifferentiated KOLF2.1J iPSCs; SEC-MS data in undifferentiated KOLF2.1J iPSCs and iPSC-derived NPCs, neurons, and cardiomyocytes; and IF images in MDA-MB-468 breast cancer cells in the presence and absence of chemotherapy (vorinostat and paclitaxel). CM4AI output data are packaged with provenance graphs and rich metadata as AI-ready datasets in RO-Crate format using the FAIRSCAPE framework. Data presented here will be augmented regularly through the end of the project. CM4AI is a collaboration of UCSD, UCSF, Stanford, UVA, Yale, UA Birmingham, Simon Fraser University, and the Hastings Center. This data is Copyright (c) 2025 The Regents of the University of California except where otherwise noted. Spatial proteomics raw image data is copyright (c) 2025 The Board of Trustees of the Leland Stanford Junior University. Dataset licensed for reuse under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International license (https://creativecommons.org/licenses/by-nc-sa/4.0/). Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived products should cite the Related Publication below, as well as directly citing this data collection (2025-03-04). (2025-03-07)
Subject
Medicine, Health and Life Sciences
Keyword
AI, affinity purification, AP-MS, artificial intelligence, breast cancer, Bridge2AI, cardiomyocyte, CM4AI, CRISPR/Cas9, induced pluripotent stem cell, iPSC, KOLF2.1J, machine learning, mass spectroscopy, MDA-MB-468, neural progenitor cell, NPC, neuron, paclitaxel, perturb-seq, perturbation sequencing, protein-protein interaction, protein localization, single-cell RNA sequencing, scRNAseq, SEC-MS, size exclusion chromatography, subcellular imaging, vorinostat
Related Publication
References: Clark T, Parker J, Schaffer L, Obernier K, Al Manir S, Churas CP, Dailamy A, Doctor Y, Forget A, Hansen JN, Hu M, Lenkiewicz J, Levinson MA, Marquez C, Nourreddine S, Niestroy J, Pratt D, Qian G, Thaker S, Bélisle-Pipon JC, Brandt C, Chen J, Ding Y, Fodeh S, Krogan N, Lundberg E, Mali P, Payne-Foster P, Ratcliffe S, Ravitsky V, Sali A, Schulz W, Ideker T. Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines. 2024.doi: http://doi.org/10.1101/2024.05.21.589311
License/Data Use Agreement
CC BY-NC-SA 4.0
ui-button
Files
Metadata
Terms
Versions
Change View
Table
Tree
Search
Filter by
File Type:
All
All
Archive (3)
Data (3)
Access:
All
All
Public (6)
Sort
Name (A-Z)
Name (Z-A)
Newest
Oldest
Size
Type
1 to 6 of 6 Files
Download
ro-crate-metadata.json
CRISPR Perturbation Cell Atlas/
JSON
- 31.1 KB
Published Mar 3, 2025
36 Downloads
MD5: cbdb263b1c099396d75e16f00a79a818
This dataset represents an expressed genome-scale CRISPRi Perturbation Cell Atlas in undifferentiated KOLF2.1J human induced pluripotent stem cells (hiPSCs) mapping transcriptional and fitness phenotypes associated with 11,739 targeted genes, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) Functional Genomics Grand Challenge, a component of the U.S. National Institute of Health’s (NIH) Bridge2AI program. We validated these findings via phenotypic, protein-interaction, and metabolic tracing assays.
Access File
File Access
Public
Download Options
JSON
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
ro-crate-metadata.json
CRISPR Perturbation RNA Sequences - Raw Sequences/
JSON
- 1000.1 KB
Published Mar 3, 2025
19 Downloads
MD5: 1cafefa32a897998e3e2ba0a29a3ef5c
This dataset represents raw sequence data from an expressed genome-scale CRISPRi Perturbation Cell Atlas in KOLF2.1J human induced pluripotent stem cells (hiPSCs) mapping transcriptional and fitness phenotypes associated with 11,739 targeted genes, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) Functional Genomics Grand Challenge, a component of the U.S. National Institute of Health’s (NIH) Bridge2AI program.
Access File
File Access
Public
Download Options
JSON
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai-v0.6-beta-if-images-paclitaxel.zip
Protein Localization Subcellular Images/
ZIP Archive
- 2.6 GB
Published Mar 3, 2025
95 Downloads
MD5: 9422486c80bc9e1d35b2fbbc72a5f043
This data set displays the spatial localization of 563 proteins of interest in cells of the breast cancer cell line MDA-MB-468 treated with paclitaxel as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).
Preview "Protein Localization Subcellular Images/cm4ai-v0.6-beta-if-images-paclitaxel.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai-v0.6-beta-if-images-untreated.zip
Protein Localization Subcellular Images/
ZIP Archive
- 3.2 GB
Published Mar 3, 2025
70 Downloads
MD5: 0b4d129f5fbc3bb7f7ea564cd032cef7
This data set displays the spatial localization of 563 proteins of interest in untreated cells of the breast cancer cell line MDA-MB-468 as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).
Preview "Protein Localization Subcellular Images/cm4ai-v0.6-beta-if-images-untreated.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai-v0.6-beta-if-images-vorinostat.zip
Protein Localization Subcellular Images/
ZIP Archive
- 2.8 GB
Published Mar 3, 2025
59 Downloads
MD5: ac577109a41a9806978461157b777d52
This data set displays the spatial localization of 563 proteins of interest in cells of the breast cancer cell line MDA-MB-468 treated with vorinostat as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).
Preview "Protein Localization Subcellular Images/cm4ai-v0.6-beta-if-images-vorinostat.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
ro-crate-metadata.json
Protein-protein Interaction SEC-MS/
JSON
- 2.9 KB
Published Mar 3, 2025
23 Downloads
MD5: cb67e7749b15ce87b9042a9feba9d032
This dataset was generated by size exclusion chromatography-mass spectroscopy (SEC-MS) on undifferentiated KOLF2.1J human induced pluripotent stem cells (hiPSCs), in the Nevan Krogan laboratory at the University of California San Francisco, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) Functional Genomics Grand Challenge, a component of the U.S. National Institute of Health’s (NIH) Bridge2AI program. The data will be uploaded to Pride when available.
Access File
File Access
Public
Download Options
JSON
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
Export Metadata
OAI_ORE
DataCite
OpenAIRE
Schema.org JSON-LD
DDI Codebook v2
Dublin Core
DDI HTML Codebook
JSON
Citation Metadata
Persistent Identifier
doi:10.18130/V3/B35XWX
Publication Date
2025-03-03
Title
Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)
Author
Clark T (University of Virginia) - ORCID:
https://orcid.org/0000-0003-4060-7360
Parker J (University of California, San Diego) - ORCID:
https://orcid.org/0000-0003-4535-3486
Al Manir S (University of Virginia) - ORCID:
https://orcid.org/0000-0003-4647-3877
Axelsson U (KTH Royal Institute of Technology,)
Ballllosero Navarro F (Stanford University) - ORCID:
https://orcid.org/0000-0002-4180-422X
Chinn B (University of California San Diego)
Churas CP (University of California San Diego) https://orcid.org/0000-0001-9998-705X
Dailamy A (University of California, San Diego) - ORCID:
https://orcid.org/0000-0002-6711-8260
Doctor Y (University of California, San Diego) - ORCID:
https://orcid.org/0009-0009-0483-7506
Fall J (KTH - Royal Institute of Technology)
Forget A (University of California San Francisco) - ORCID:
https://orcid.org/0000-0003-0223-0312
Gao J (University of California San Diego) - ORCID:
https://orcid.org/0000-0002-6311-3526
Hansen JN (Stanford University) - ORCID:
https://orcid.org/0000-0002-4650-9094
Hu M (University of California San Diego) https://orcid.org/0000-0002-1571-8029
Johannesson A (KTH - Royal Institute of Technology)
Khaliq H (University of California San Diego)
Lee YH (University of California San Diego) - ORCID:
https://orcid.org/0000-0003-0917-355X
Lenkiewicz J (University of California San Diego) https://orcid.org/0000-0001-7252-8638
Levinson MA (University of Virginia) - ORCID:
https://orcid.org/0000-0003-0384-8499
Marquez C (University of California San Diego) - ORCID:
0000-0003-3960-420X
Metallo C (University of California San Diego) - ORCID:
https://orcid.org/0000-0003-2404-3040
Muralidharan M (University of California San Francisco)
Nourreddine S (University of California San Diego) https://orcid.org/0000-0003-3881-7588
Niestroy J (University of Virginia) - ORCID:
https://orcid.org/0000-0002-1103-3882
Obernier K (University of California San Francisco) - ORCID:
https://orcid.org/0000-0002-4025-1299
Pan E (University of California San Diego)
Polacco B (University of California San Francisco)
Pratt D (University of California San Diego) - ORCID:
https://orcid.org/0000-0002-1471-9513
Qian G (University of California San Diego) - ORCID:
https://orcid.org/0009-0005-4217-2745
Schaffer L (University of California San Diego) - ORCID:
https://orcid.org/0000-0001-6339-9141
Sigaeva A (KTH Royal Institute of Technology) - ORCID:
https://orcid.org/0000-0003-3361-3797
Thaker S (University of Alabama at Birmingham) - ORCID:
https://orcid.org/0000-0001-6730-2773
Zhang Y (University of California San Diego)
Bélisle-Pipon JC (Simon Fraser University) - ORCID:
https://orcid.org/0000-0002-8965-8153
Brandt C (Yale University) - ORCID:
https://orcid.org/0000-0001-8179-1796
Chen JY (The University of Alabama at Birmingham) - ORCID:
https://orcid.org/0000-0002-6112-415X
Ding Y (University of Texas at Austin) - ORCID:
https://orcid.org/0000-0003-2567-2009
Fodeh S (Yale University) - ORCID:
https://orcid.org/0000-0003-4664-3143
Krogan N (University of California San Francisco) - ORCID:
https://orcid.org/0000-0003-4902-337X
Lundberg E (Stanford University) - ORCID:
https://orcid.org/0000-0001-7034-0850
Mali P (University of California San Diego) https://orcid.org/0000-0002-3383-1287
Payne-Foster P (University of Alabama) - ORCID:
https://orcid.org/0000-0002-3508-3577
Ratcliffe S (University of Virginia) - ORCID:
https://orcid.org/0000-0002-6644-8284
Ravitsky V (University of Montreal) - ORCID:
https://orcid.org/0000-0002-7080-8801
Sali A (University of California San Diego) - ORCID:
https://orcid.org/0000-0003-0435-6197
Schulz W (Yale University) - ORCID:
https://orcid.org/0000-0002-2048-4028
Ideker T (University of California San Diego) - ORCID:
https://orcid.org/0000-0002-1708-8454
Point of Contact
Use email button above to contact.
Ideker Trey (University of California San Diego)
Dataset Description
This dataset is the March 2025 Data Release of Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org), the Functional Genomics Grand Challenge in the NIH Bridge2AI program. This Beta release includes perturb-seq data in undifferentiated KOLF2.1J iPSCs; SEC-MS data in undifferentiated KOLF2.1J iPSCs and iPSC-derived NPCs, neurons, and cardiomyocytes; and IF images in MDA-MB-468 breast cancer cells in the presence and absence of chemotherapy (vorinostat and paclitaxel). CM4AI output data are packaged with provenance graphs and rich metadata as AI-ready datasets in RO-Crate format using the FAIRSCAPE framework. Data presented here will be augmented regularly through the end of the project. CM4AI is a collaboration of UCSD, UCSF, Stanford, UVA, Yale, UA Birmingham, Simon Fraser University, and the Hastings Center. This data is Copyright (c) 2025 The Regents of the University of California except where otherwise noted. Spatial proteomics raw image data is copyright (c) 2025 The Board of Trustees of the Leland Stanford Junior University. Dataset licensed for reuse under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International license (https://creativecommons.org/licenses/by-nc-sa/4.0/). Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived products should cite the Related Publication below, as well as directly citing this data collection (2025-03-04). (2025-03-07)
Subject
Medicine, Health and Life Sciences
Keyword
AI
http://purl.obolibrary.org/obo/NCIT_C16309
(NCI Thesaurus)
affinity purification
http://www.bioassayontology.org/bao#BAO_0002603
(BioAssay Ontology (BAO))
AP-MS
http://www.ebi.ac.uk/swo/SWO_1100012
(Software Ontology)
artificial intelligence
http://purl.obolibrary.org/obo/NCIT_C16309
(NCI Thesaurus)
breast cancer
http://purl.bioontology.org/ontology/LNC/LA14283-8
(LOINC)
Bridge2AI
cardiomyocyte
http://purl.obolibrary.org/obo/CL_0000746
CM4AI
CRISPR/Cas9
http://www.bioassayontology.org/bao#BAO_0010249
(Bioassay Ontology (BAO))
induced pluripotent stem cell
http://www.ebi.ac.uk/efo/EFO_0004905
(Experimental Factor Ontology (EFO))
iPSC
http://www.ebi.ac.uk/efo/EFO_0004905
(Experimental Factor Ontology (EFO))
KOLF2.1J
machine learning
http://purl.obolibrary.org/obo/OBI_0002587
(Ontology of Biomedical Investigations (OBI))
http://purl.obolibrary.org/obo/obi.owl
mass spectroscopy
http://purl.bioontology.org/ontology/MESH/D013058
(Medical Subject Headings (MeSH))
MDA-MB-468
neural progenitor cell
http://purl.obolibrary.org/obo/CL_0011020
(Cell Ontology (CL))
NPC
http://purl.obolibrary.org/obo/CL_0011020
(Cell Ontology (CL))
http://purl.obolibrary.org/obo/cl.owl
neuron
http://purl.obolibrary.org/obo/CL_0000540
(Cell Ontology (CL))
http://purl.obolibrary.org/obo/cl.owl
paclitaxel
http://purl.obolibrary.org/obo/CHEBI_45863
(Chemical Entitites of Biological Interest (CHEBI))
perturb-seq
http://www.ebi.ac.uk/efo/EFO_0008860
(Experimental Factor Ontology (EFO))
perturbation sequencing
http://www.ebi.ac.uk/efo/EFO_0008860
(Experimental Factor Ontology (EFO))
protein-protein interaction
http://purl.obolibrary.org/obo/NCIT_C18469
(NCI Thesaurus (NCIT))
protein localization
http://purl.obolibrary.org/obo/GO_0008104
(Gene Ontology (GO))
http://purl.obolibrary.org/obo/go/extensions/go-plus.owl
single-cell RNA sequencing
http://www.ebi.ac.uk/efo/EFO_0008913
(Experimental Factor Ontology (EFO))
scRNAseq
http://www.ebi.ac.uk/efo/EFO_0008913
(Experimental Factor Ontology (EFO))
SEC-MS
size exclusion chromatography
subcellular imaging
vorinostat
http://purl.obolibrary.org/obo/CHEBI_45716
(Chemical Entitites of Biological Interest (CHEBI))
Related Publication
References: Clark T, Parker J, Schaffer L, Obernier K, Al Manir S, Churas CP, Dailamy A, Doctor Y, Forget A, Hansen JN, Hu M, Lenkiewicz J, Levinson MA, Marquez C, Nourreddine S, Niestroy J, Pratt D, Qian G, Thaker S, Bélisle-Pipon JC, Brandt C, Chen J, Ding Y, Fodeh S, Krogan N, Lundberg E, Mali P, Payne-Foster P, Ratcliffe S, Ravitsky V, Sali A, Schulz W, Ideker T. Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines. 2024. doi http://doi.org/10.1101/2024.05.21.589311
Nourreddine S, Doctor Y, Dailamy A, Forget A, Lee YH, Chinn B, Khaliq H, Polacco B, Muralidharan M, Pan E, Zhang Y, Sigaeva A, Hansen JN, Gao J, Parker JA, Obernier K, Clark T, Chen JY, Metallo C, Lundberg E, Ideker T, Krogan N, Mali P. A PERTURBATION CELL ATLAS OF HUMAN INDUCED PLURIPOTENT STEM CELLS. bioRxiv. 2024 Nov 4;2024.11.03.621734. PMCID: PMC11580897 doi https://doi.org/10.1101/2024.11.03.621734
Data Creation Date
2025-02-27
Production Location
University of California San Diego; University of California San Francisco; University of California San Francisco; Stanford University; University of Virginia
Funding Information
National Institutes of Health: 1OT2OD032742-01
Depositor
Niestroy, Justin
Deposit Date
2025-02-27
Dataset Terms
License/Data Use Agreement
Our
Community Norms
as well as good scientific practices expect that proper credit is given via citation. Please use the data citation shown on the dataset page.
CC BY-NC-SA 4.0
View Differences
Direct
Dataset Version
Summary
Version Note
Contributors
Published on
No records found.
Edit File
This file has already been deleted (or replaced) in the current version. It may not be edited.
Close
Restrict Access
Restricting limits access to published files. People who want to use the restricted files can request access by default.
If you disable request access, you must add information about access to the Terms of Access field.
Learn about restricting files and dataset access in the
User Guide
.
Request Access
Enable access request
You must enable request access or add terms of access to restrict file access.
Terms of Access for Restricted Files
Save Changes
Cancel
Edit Embargo
The selected file or files have already been published. Contact an administrator to change the embargo date or reason of the file or files.
Cancel
Edit Retention Period
The selected file or files have already been published. Contact an administrator to change the retention period date or reason of the file or files.
Cancel
Delete Files
The file will be deleted after you click on the Delete button.
Files will not be removed from previously published versions of the dataset.
Delete
Cancel
Continue
Cancel
Select File(s)
Please select one or more files.
Close
Share Dataset
Share this dataset on your favorite social media networks.
Close
Continue
Cancel
Dataset Citations
Citations for this dataset are retrieved from Crossref via DataCite using Make Data Count standards. For more information about dataset metrics, please refer to the
User Guide
.
Sorry, no citations were found.
Close
Inaccessible Files Selected
The selected file(s) may not be downloaded because you have not been granted access or the file(s) have a retention period that has expired or the files can only be transferred via Globus.
You may request access to any restricted file(s) by clicking the Request Access button.
Close
Ineligible Files Selected
The selected file(s) may not be transferred because you have not been granted access or the file(s) have a retention period that has expired or the files are not Globus accessible.
You may request access to any restricted file(s) by clicking the Request Access button.
Close
Download Options
The files selected are too large to download as a ZIP.
You can select individual files that are below the 953.7 MB download limit from the files table, or use the
Data Access API
for programmatic access to the files.
Select File(s)
Please select a file or files to be downloaded.
Close
Inaccessible Files Selected
The selected file(s) may not be downloaded because you have not been granted access or the file(s) have a retention period that has expired.
Click Continue to download the files you have access to download.
Continue
Cancel
Ineligible Files Selected
Some file(s) cannot be transferred. (They are restricted, embargoed, with an expired retention period, or not Globus accessible.)
Click Continue to transfer the elligible files.
Continue
Cancel
Delete Dataset
Are you sure you want to delete this dataset and all of its files? You cannot undelete this dataset.
Continue
Cancel
Delete Draft Version
Are you sure you want to delete this draft version? Files will be reverted to the most recently published version. You cannot undelete this draft.
Continue
Cancel
Unpublished Dataset Preview URL
You can create a Preview URL to copy and share with others who will not need a repository account to review this unpublished dataset version. Once the dataset is published or if the URL is disabled, the URL will no longer work and will point to a "Page not found" page.
Only one Preview URL can be active for a single draft dataset.
To cite this data in publications, use the datasets persistent ID instead of this URL. For more information about the Preview URL feature, please refer to the
User Guide
.
General Preview
Create a URL that others can use to review this draft dataset version before it is published. They will be able to access all files in the dataset and see all metadata, including metadata that may identify the dataset's authors.
Create General Preview URL
Anonymous Preview
Create a URL that others can use to access an anonymized view of this unpublished dataset version. Metadata that could identify the dataset's author will not be displayed. (See Tool Tip for the list of withheld metadata fields.) Non-identifying metadata will be visible.
The dataset's files are not changed and users of the Anonymous Preview URL will be able to access them. Users of the Anonymous Preview URL will not be able to see the name of the Dataverse that this dataset is in but will be able to see the name of the repository, which might expose the dataset authors' identities.
To verify that all identifying information has been removed or anonymized, it is recommended that you logout and review the dataset as as it would be seen by an Anonymous Preview URL user.
See
User Guide
for more information.
You won't be able to create an Anonymous Preview URL once a version of this dataset has been published.
Create Anonymous Preview URL
Close
Unpublished Dataset Preview URL
Are you sure you want to disable the Preview URL? If you have shared the Preview URL with others they will no longer be able to use it to access your unpublished dataset.
Yes, Disable General Preview URL
Cancel
Delete Files
The file(s) will be deleted after you click on the Delete button.
Files will not be removed from previously published versions of the dataset.
Delete
Cancel
Compute
This dataset contains restricted files you may not compute on because you have not been granted access.
Close
Deaccession Dataset
Are you sure you want to deaccession? This is permanent and the selected version(s) will no longer be viewable by the public.
No
Deaccession Dataset
Are you sure you want to deaccession this dataset? This is permanent an it will no longer be viewable by the public.
No
Version Differences Details
Please select two versions to view the differences.
Close
Version Differences Details
Version:
Last Updated:
Version:
Last Updated:
Done
Select File(s)
Please select a file or files for access request.
Close
Select File(s)
Embargoed files cannot be accessed. Please select an unembargoed file or files for your access request.
Close
Edit Tags
Select existing file tags or create new tags to describe your files. Each file can have more than one tag.
Save Changes
Cancel
Request Access
You need to
Log In
to request access.
Close
Dataset Terms
Please confirm and/or complete the information needed below in order to request access to files in this dataset.
This dataset is made available under the following terms. Please confirm and/or complete the information needed below in order to continue.
License/Data Use Agreement
Our
Community Norms
as well as good scientific practices expect that proper credit is given via citation. Please use the data citation shown on the dataset page.
CC BY-NC-SA 4.0
Preview Guestbook
Upon downloading files the guestbook asks for the following information.
Guestbook Name
Collected Data
Account Information
Close
Package File Download
Use the Download URL in a Wget command or a download manager to download this package file. Download via web browser is not recommended.
User Guide - Downloading a Dataverse Package via URL
Download URL
https://dataverse.lib.virginia.edu/api/access/datafile/
Close
Compute Batch
Clear Batch
ui-button
Dataset
Persistent Identifier
Change Compute Batch
Compute Batch
Cancel
Submit for Review
You will not be able to make changes to this dataset while it is in review.
Submit
Cancel
Publish Dataset
Are you sure you want to republish this dataset?
Select if this is a minor or major version update.
Minor Release (1.5)
Major Release (2.0)
Continue
Cancel
Publish Dataset
This dataset cannot be published until
Cell Maps for Artificial Intelligence
is published by its administrator.
Close
Publish Dataset
This dataset cannot be published until
Cell Maps for Artificial Intelligence
and
School of Medicine
are published.
Close
Return to Author
Return this dataset to contributor for modification. The reason for return entered below will be sent by email to the author.
Continue
Cancel
Add/Edit a Version Note
Styled Citation
Copyright © 2025, by the Rector and Visitors of the University of Virginia |
Terms of Use
|
Privacy Policy
Powered by
v. 6.6 build 1829-192cdc4
Contact University of Virginia Dataverse Support
To
University of Virginia Dataverse Support
From
Subject
Message
Please fill this out to prove you are not a robot.
2 + 7 =
Send Message
Cancel


================================================================================

FILE: dataverse_10.18130_V3_F3TD5R_row19.txt
PATH: data/preprocessed/individual/CM4AI/dataverse_10.18130_V3_F3TD5R_row19.txt
SIZE: 29599 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: CM4AI
Source ID: june_2025_dataverse_release
Source type: historical data release
Source URL: https://dataverse.lib.virginia.edu/dataset.xhtml?persistentId=doi:10.18130/V3/F3TD5R
Raw file: data/raw/CM4AI/dataverse_10.18130_V3_F3TD5R_row19.html
--------------------------------------------------------------------------------
Cell Maps for Artificial Intelligence - June 2025 Data Release (Beta) - Cell Maps for Artificial Intelligence
Skip to main content
Toggle navigation
Search
Search
About
User Guide
Support
Log In
Cell Maps for Artificial Intelligence
This collection is under review for potential modification in compliance with Administration directives.
University of Virginia Dataverse
>
LibraData: UVa's Scholarly Research
>
School of Medicine
>
Cell Maps for Artificial Intelligence
>
Info
– The "DRAFT" version was not found. This is version "2.1".
Cell Maps for Artificial Intelligence - June 2025 Data Release (Beta)
Version 2.1
Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Marquez C; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Polacco B; Pratt D; Qian G; Schaffer L; Sigaeva A; Thaker S; Zhang Y; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, "Cell Maps for Artificial Intelligence - June 2025 Data Release (Beta)",
https://doi.org/10.18130/V3/F3TD5R
, University of Virginia Dataverse, V2
Cite Dataset
Download EndNote XML
Download RIS
Download BibTeX
View Styled Citation
Learn about
Data Citation Standards
.
Access Dataset
The dataset is too large to download. Please select the files you need from the files table.
Contact Owner
Share
Dataset Metrics
256 Downloads
Dataset Description
Description
This dataset is a revision of the June 2025 Data Release of Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org), the Functional Genomics Grand Challenge in the NIH Bridge2AI program. This revision includes adding RGB immunofluorescent images, corrections to ro-crate metadata, and changes to naming conventions.
This Beta release includes perturb-seq data in undifferentiated KOLF2.1J iPSCs; SEC-MS data in undifferentiated KOLF2.1J iPSCs, iPSC-derived NPCs, neurons, cardiomyocytes, and treated and untreated MDA-MB468 breast cancer cells; and IF images in MDA-MB-468 breast cancer cells in the presence and absence of chemotherapy (vorinostat and paclitaxel).
External Data Links
Access external data resources related to this dataset:
Sequence Read Archive (SRA) Data:
NCBI BioProject
Mass Spectrometry Data (Human iPSCs):
MassIVE Repository
Mass Spectrometry Data (Human Cancer Cells):
MassIVE Repository
Data Governance & Ethics
Human Subjects:
No
De-identified Samples:
Yes
FDA Regulated:
No
Data Governance Committee:
Jillian Parker (jillianparker@health.ucsd.edu)
Ethical Review:
Vardit Ravitsky (ravitskyv@thehastingscenter.org) and Jean-Christophe Belisle-Pipon (jean-christophe_belisle-pipon@sfu.ca)
Completeness
These data are not yet in completed final form:
Some datasets are under temporary pre-publication embargo
Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap
Computed cell maps not included in this release
Maintenance Plan
Dataset will be regularly updated and augmented through the end of the project in November 2026
Updates on a quarterly basis
Long term preservation in the University of Virginia Dataverse, supported by committed institutional funds
Intended Use
This dataset is intended for:
AI-ready datasets to support research in functional genomics
AI model training
Cellular process analysis
Cell architectural changes and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations
Limitations
Researchers should be aware of inherent limitations:
This is an interim release
Does not contain predicted cell maps, which will be added in future releases
The current release is most suitable for bioinformatics analysis of the individual datasets
Requires domain expertise for meaningful analysis
Prohibited Uses
These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval
Potential Sources of Bias
Users should be aware of potential biases:
Data in this release was derived from commercially available de-identified human cell lines
Does not represent all biological variants which may be seen in the population at large
(2025-06-30)
Subject
Medicine, Health and Life Sciences
Keyword
AI, affinity purification, AP-MS, artificial intelligence, breast cancer, Bridge2AI, cardiomyocyte, CM4AI, CRISPR/Cas9, induced pluripotent stem cell, iPSC, KOLF2.1J, machine learning, mass spectroscopy, MDA-MB-468, neural progenitor cell, NPC, neuron, paclitaxel, perturb-seq, perturbation sequencing, protein-protein interaction, protein localization, single-cell RNA sequencing, scRNAseq, SEC-MS, size exclusion chromatography, subcellular imaging, vorinostat
Related Publication
References: Clark T, Parker J, Schaffer L, Obernier K, Al Manir S, Churas CP, Dailamy A, Doctor Y, Forget A, Hansen JN, Hu M, Lenkiewicz J, Levinson MA, Marquez C, Nourreddine S, Niestroy J, Pratt D, Qian G, Thaker S, Bélisle-Pipon JC, Brandt C, Chen J, Ding Y, Fodeh S, Krogan N, Lundberg E, Mali P, Payne-Foster P, Ratcliffe S, Ravitsky V, Sali A, Schulz W, Ideker T. Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines. 2024.doi: http://doi.org/10.1101/2024.05.21.589311
License/Data Use Agreement
CC BY-NC-SA 4.0
ui-button
Files
Metadata
Terms
Versions
Change View
Table
Tree
Search
Filter by
File Type:
All
All
Text (12)
Data (5)
Archive (3)
Metadata (1)
Access:
All
All
Public (21)
Sort
Name (A-Z)
Name (Z-A)
Newest
Oldest
Size
Type
1 to 10 of 21 Files
Download
release-ro-crate-datasheet.html
HTML
- 89.5 KB
Published Jul 1, 2025
20 Downloads
MD5: 599c9ece9b88b3ce797b82463b4a1eb4
HTML datasheet summarizing key release information.
Preview "release-ro-crate-datasheet.html"
Access File
File Access
Public
Download Options
HTML
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
Explore Options
HTML Previewer
release-ro-crate-metadata.json
RO-Crate metadata
- 35.0 KB
Published Jul 1, 2025
24 Downloads
MD5: 99f9e00053bff3020fd9832a3a518bbb
Release RO-Crate with pointers to sub ro-crates.
Preview "release-ro-crate-metadata.json"
Access File
File Access
Public
Download Options
RO-Crate metadata
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_ifimages_MDA-MB-468_paclitaxel.zip
Images/
ZIP Archive
- 3.8 GB
Published Oct 22, 2025
1 Download
MD5: 0d972b80744344ddeede516a0cf6e3d7
This data set displays the spatial localization of 464 proteins of interest in cells of the breast cancer cell line MDA-MB-468 treated with paclitaxel as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).
Preview "Images/cm4ai_ifimages_MDA-MB-468_paclitaxel.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_ifimages_MDA-MB-468_untreated.zip
Images/
ZIP Archive
- 4.6 GB
Published Oct 22, 2025
0 Downloads
MD5: a98affcc05429650c6bb3906cd836d55
This data set displays the spatial localization of 464 proteins of interest in cells of the breast cancer cell line MDA-MB-468 treated as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).
Preview "Images/cm4ai_ifimages_MDA-MB-468_untreated.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_ifimages_MDA-MB-468_vorinostat.zip
Images/
ZIP Archive
- 4.2 GB
Published Oct 22, 2025
0 Downloads
MD5: ad4e68ccc14b0f3349dad3321e7b81b2
This data set displays the spatial localization of 464 proteins of interest in cells of the breast cancer cell line MDA-MB-468 treated with vorinostat as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).
Preview "Images/cm4ai_ifimages_MDA-MB-468_vorinostat.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
Images-paclitaxel-provenance-graph.html
Images/paclitaxel/
HTML
- 39.2 KB
Published Jul 1, 2025
7 Downloads
MD5: e38e63e4c8dfc5808a5ffa2d7829fc38
The .html provenance graph files require download to view properly. They will not display in the Dataverse preview.
Preview "Images/paclitaxel/Images-paclitaxel-provenance-graph.html"
Access File
File Access
Public
Download Options
HTML
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
Explore Options
HTML Previewer
Images-untreated-provenance-graph.html
Images/untreated/
HTML
- 39.2 KB
Published Jul 1, 2025
1 Download
MD5: 1a3b510f74d3f8647e07c6559ce64ee8
The .html provenance graph files require download to view properly. They will not display in the Dataverse preview.
Preview "Images/untreated/Images-untreated-provenance-graph.html"
Access File
File Access
Public
Download Options
HTML
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
Explore Options
HTML Previewer
Images-vorinostat-provenance-graph.html
Images/vorinostat/
HTML
- 39.2 KB
Published Jul 1, 2025
3 Downloads
MD5: 58935fe4e254b31d33fed019f24c7668
The .html provenance graph files require download to view properly. They will not display in the Dataverse preview.
Preview "Images/vorinostat/Images-vorinostat-provenance-graph.html"
Access File
File Access
Public
Download Options
HTML
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
Explore Options
HTML Previewer
mass-spec-cancer-cells-provenance-graph.html
mass-spec/cancer-cells/
HTML
- 41.7 KB
Published Jul 1, 2025
9 Downloads
MD5: 931ad9b552562024cb84ebe62d1f1838
The .html provenance graph files require download to view properly.
They will not display in the Dataverse preview.
Preview "mass-spec/cancer-cells/mass-spec-cancer-cells-provenance-graph.html"
Access File
File Access
Public
Download Options
HTML
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
Explore Options
HTML Previewer
mass-spec-cancer-cells-ro-crate-metadata.json
mass-spec/cancer-cells/
JSON
- 343.4 KB
Published Jul 1, 2025
6 Downloads
MD5: 3a7063bb391ea5e05a32ba5da5f4b2f8
Access File
File Access
Public
Download Options
JSON
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
F
P
1
2
3
N
E
Files Per Page
Rows Per Page
10
25
50
Export Metadata
OAI_ORE
DataCite
OpenAIRE
Schema.org JSON-LD
DDI Codebook v2
Dublin Core
DDI HTML Codebook
JSON
Citation Metadata
Persistent Identifier
doi:10.18130/V3/F3TD5R
Publication Date
2025-07-01
Title
Cell Maps for Artificial Intelligence - June 2025 Data Release (Beta)
Author
Clark T (University of Virginia) - ORCID:
https://orcid.org/0000-0003-4060-7360
Parker J (University of California, San Diego) - ORCID:
https://orcid.org/0000-0003-4535-3486
Al Manir S (University of Virginia) - ORCID:
https://orcid.org/0000-0003-4647-3877
Axelsson U (KTH Royal Institute of Technology,)
Ballllosero Navarro F (Stanford University) - ORCID:
https://orcid.org/0000-0002-4180-422X
Chinn B (University of California San Diego)
Churas CP (University of California San Diego) https://orcid.org/0000-0001-9998-705X
Dailamy A (University of California, San Diego) - ORCID:
https://orcid.org/0000-0002-6711-8260
Doctor Y (University of California, San Diego) - ORCID:
https://orcid.org/0009-0009-0483-7506
Fall J (KTH - Royal Institute of Technology)
Forget A (University of California San Francisco) - ORCID:
https://orcid.org/0000-0003-0223-0312
Gao J (University of California San Diego) - ORCID:
https://orcid.org/0000-0002-6311-3526
Hansen JN (Stanford University) - ORCID:
https://orcid.org/0000-0002-4650-9094
Hu M (University of California San Diego) https://orcid.org/0000-0002-1571-8029
Johannesson A (KTH - Royal Institute of Technology)
Khaliq H (University of California San Diego)
Lee YH (University of California San Diego) - ORCID:
https://orcid.org/0000-0003-0917-355X
Lenkiewicz J (University of California San Diego) https://orcid.org/0000-0001-7252-8638
Levinson MA (University of Virginia) - ORCID:
https://orcid.org/0000-0003-0384-8499
Marquez C (University of California San Diego) - ORCID:
0000-0003-3960-420X
Metallo C (University of California San Diego) - ORCID:
https://orcid.org/0000-0003-2404-3040
Muralidharan M (University of California San Francisco)
Nourreddine S (University of California San Diego) https://orcid.org/0000-0003-3881-7588
Niestroy J (University of Virginia) - ORCID:
https://orcid.org/0000-0002-1103-3882
Obernier K (University of California San Francisco) - ORCID:
https://orcid.org/0000-0002-4025-1299
Pan E (University of California San Diego)
Polacco B (University of California San Francisco)
Pratt D (University of California San Diego) - ORCID:
https://orcid.org/0000-0002-1471-9513
Qian G (University of California San Diego) - ORCID:
https://orcid.org/0009-0005-4217-2745
Schaffer L (University of California San Diego) - ORCID:
https://orcid.org/0000-0001-6339-9141
Sigaeva A (KTH Royal Institute of Technology) - ORCID:
https://orcid.org/0000-0003-3361-3797
Thaker S (University of Alabama at Birmingham) - ORCID:
https://orcid.org/0000-0001-6730-2773
Zhang Y (University of California San Diego)
Bélisle-Pipon JC (Simon Fraser University) - ORCID:
https://orcid.org/0000-0002-8965-8153
Brandt C (Yale University) - ORCID:
https://orcid.org/0000-0001-8179-1796
Chen JY (The University of Alabama at Birmingham) - ORCID:
https://orcid.org/0000-0002-6112-415X
Ding Y (University of Texas at Austin) - ORCID:
https://orcid.org/0000-0003-2567-2009
Fodeh S (Yale University) - ORCID:
https://orcid.org/0000-0003-4664-3143
Krogan N (University of California San Francisco) - ORCID:
https://orcid.org/0000-0003-4902-337X
Lundberg E (Stanford University) - ORCID:
https://orcid.org/0000-0001-7034-0850
Mali P (University of California San Diego) https://orcid.org/0000-0002-3383-1287
Payne-Foster P (University of Alabama) - ORCID:
https://orcid.org/0000-0002-3508-3577
Ratcliffe S (University of Virginia) - ORCID:
https://orcid.org/0000-0002-6644-8284
Ravitsky V (University of Montreal) - ORCID:
https://orcid.org/0000-0002-7080-8801
Sali A (University of California San Diego) - ORCID:
https://orcid.org/0000-0003-0435-6197
Schulz W (Yale University) - ORCID:
https://orcid.org/0000-0002-2048-4028
Ideker T (University of California San Diego) - ORCID:
https://orcid.org/0000-0002-1708-8454
Point of Contact
Use email button above to contact.
Ideker Trey (University of California San Diego)
Dataset Description
Description
This dataset is a revision of the June 2025 Data Release of Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org), the Functional Genomics Grand Challenge in the NIH Bridge2AI program. This revision includes adding RGB immunofluorescent images, corrections to ro-crate metadata, and changes to naming conventions.
This Beta release includes perturb-seq data in undifferentiated KOLF2.1J iPSCs; SEC-MS data in undifferentiated KOLF2.1J iPSCs, iPSC-derived NPCs, neurons, cardiomyocytes, and treated and untreated MDA-MB468 breast cancer cells; and IF images in MDA-MB-468 breast cancer cells in the presence and absence of chemotherapy (vorinostat and paclitaxel).
External Data Links
Access external data resources related to this dataset:
Sequence Read Archive (SRA) Data:
NCBI BioProject
Mass Spectrometry Data (Human iPSCs):
MassIVE Repository
Mass Spectrometry Data (Human Cancer Cells):
MassIVE Repository
Data Governance & Ethics
Human Subjects:
No
De-identified Samples:
Yes
FDA Regulated:
No
Data Governance Committee:
Jillian Parker (jillianparker@health.ucsd.edu)
Ethical Review:
Vardit Ravitsky (ravitskyv@thehastingscenter.org) and Jean-Christophe Belisle-Pipon (jean-christophe_belisle-pipon@sfu.ca)
Completeness
These data are not yet in completed final form:
Some datasets are under temporary pre-publication embargo
Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap
Computed cell maps not included in this release
Maintenance Plan
Dataset will be regularly updated and augmented through the end of the project in November 2026
Updates on a quarterly basis
Long term preservation in the University of Virginia Dataverse, supported by committed institutional funds
Intended Use
This dataset is intended for:
AI-ready datasets to support research in functional genomics
AI model training
Cellular process analysis
Cell architectural changes and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations
Limitations
Researchers should be aware of inherent limitations:
This is an interim release
Does not contain predicted cell maps, which will be added in future releases
The current release is most suitable for bioinformatics analysis of the individual datasets
Requires domain expertise for meaningful analysis
Prohibited Uses
These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval
Potential Sources of Bias
Users should be aware of potential biases:
Data in this release was derived from commercially available de-identified human cell lines
Does not represent all biological variants which may be seen in the population at large
(2025-06-30)
Subject
Medicine, Health and Life Sciences
Keyword
AI
http://purl.obolibrary.org/obo/NCIT_C16309
(NCI Thesaurus)
affinity purification
http://www.bioassayontology.org/bao#BAO_0002603
(BioAssay Ontology (BAO))
AP-MS
http://www.ebi.ac.uk/swo/SWO_1100012
(Software Ontology)
artificial intelligence
http://purl.obolibrary.org/obo/NCIT_C16309
(NCI Thesaurus)
breast cancer
http://purl.bioontology.org/ontology/LNC/LA14283-8
(LOINC)
Bridge2AI
cardiomyocyte
http://purl.obolibrary.org/obo/CL_0000746
CM4AI
CRISPR/Cas9
http://www.bioassayontology.org/bao#BAO_0010249
(Bioassay Ontology (BAO))
induced pluripotent stem cell
http://www.ebi.ac.uk/efo/EFO_0004905
(Experimental Factor Ontology (EFO))
iPSC
http://www.ebi.ac.uk/efo/EFO_0004905
(Experimental Factor Ontology (EFO))
KOLF2.1J
machine learning
http://purl.obolibrary.org/obo/OBI_0002587
(Ontology of Biomedical Investigations (OBI))
http://purl.obolibrary.org/obo/obi.owl
mass spectroscopy
http://purl.bioontology.org/ontology/MESH/D013058
(Medical Subject Headings (MeSH))
MDA-MB-468
neural progenitor cell
http://purl.obolibrary.org/obo/CL_0011020
(Cell Ontology (CL))
NPC
http://purl.obolibrary.org/obo/CL_0011020
(Cell Ontology (CL))
http://purl.obolibrary.org/obo/cl.owl
neuron
http://purl.obolibrary.org/obo/CL_0000540
(Cell Ontology (CL))
http://purl.obolibrary.org/obo/cl.owl
paclitaxel
http://purl.obolibrary.org/obo/CHEBI_45863
(Chemical Entitites of Biological Interest (CHEBI))
perturb-seq
http://www.ebi.ac.uk/efo/EFO_0008860
(Experimental Factor Ontology (EFO))
perturbation sequencing
http://www.ebi.ac.uk/efo/EFO_0008860
(Experimental Factor Ontology (EFO))
protein-protein interaction
http://purl.obolibrary.org/obo/NCIT_C18469
(NCI Thesaurus (NCIT))
protein localization
http://purl.obolibrary.org/obo/GO_0008104
(Gene Ontology (GO))
http://purl.obolibrary.org/obo/go/extensions/go-plus.owl
single-cell RNA sequencing
http://www.ebi.ac.uk/efo/EFO_0008913
(Experimental Factor Ontology (EFO))
scRNAseq
http://www.ebi.ac.uk/efo/EFO_0008913
(Experimental Factor Ontology (EFO))
SEC-MS
size exclusion chromatography
subcellular imaging
vorinostat
http://purl.obolibrary.org/obo/CHEBI_45716
(Chemical Entitites of Biological Interest (CHEBI))
Related Publication
References: Clark T, Parker J, Schaffer L, Obernier K, Al Manir S, Churas CP, Dailamy A, Doctor Y, Forget A, Hansen JN, Hu M, Lenkiewicz J, Levinson MA, Marquez C, Nourreddine S, Niestroy J, Pratt D, Qian G, Thaker S, Bélisle-Pipon JC, Brandt C, Chen J, Ding Y, Fodeh S, Krogan N, Lundberg E, Mali P, Payne-Foster P, Ratcliffe S, Ravitsky V, Sali A, Schulz W, Ideker T. Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines. 2024. doi http://doi.org/10.1101/2024.05.21.589311
Nourreddine S, Doctor Y, Dailamy A, Forget A, Lee YH, Chinn B, Khaliq H, Polacco B, Muralidharan M, Pan E, Zhang Y, Sigaeva A, Hansen JN, Gao J, Parker JA, Obernier K, Clark T, Chen JY, Metallo C, Lundberg E, Ideker T, Krogan N, Mali P. A PERTURBATION CELL ATLAS OF HUMAN INDUCED PLURIPOTENT STEM CELLS. bioRxiv. 2024 Nov 4;2024.11.03.621734. PMCID: PMC11580897 doi https://doi.org/10.1101/2024.11.03.621734
Data Creation Date
2025-02-27
Production Location
University of California San Diego; University of California San Francisco; Stanford University; University of Virginia
Funding Information
National Institutes of Health: 1OT2OD032742-01
Depositor
Niestroy, Justin
Deposit Date
2025-02-27
Dataset Terms
License/Data Use Agreement
Our
Community Norms
as well as good scientific practices expect that proper credit is given via citation. Please use the data citation shown on the dataset page.
CC BY-NC-SA 4.0
View Differences
Direct
Dataset Version
Summary
Version Note
Contributors
Published on
No records found.
Edit File
This file has already been deleted (or replaced) in the current version. It may not be edited.
Close
Restrict Access
Restricting limits access to published files. People who want to use the restricted files can request access by default.
If you disable request access, you must add information about access to the Terms of Access field.
Learn about restricting files and dataset access in the
User Guide
.
Request Access
Enable access request
You must enable request access or add terms of access to restrict file access.
Terms of Access for Restricted Files
Save Changes
Cancel
Edit Embargo
The selected file or files have already been published. Contact an administrator to change the embargo date or reason of the file or files.
Cancel
Edit Retention Period
The selected file or files have already been published. Contact an administrator to change the retention period date or reason of the file or files.
Cancel
Delete Files
The file will be deleted after you click on the Delete button.
Files will not be removed from previously published versions of the dataset.
Delete
Cancel
Continue
Cancel
Select File(s)
Please select one or more files.
Close
Share Dataset
Share this dataset on your favorite social media networks.
Close
Continue
Cancel
Dataset Citations
Citations for this dataset are retrieved from Crossref via DataCite using Make Data Count standards. For more information about dataset metrics, please refer to the
User Guide
.
Sorry, no citations were found.
Close
Inaccessible Files Selected
The selected file(s) may not be downloaded because you have not been granted access or the file(s) have a retention period that has expired or the files can only be transferred via Globus.
You may request access to any restricted file(s) by clicking the Request Access button.
Close
Ineligible Files Selected
The selected file(s) may not be transferred because you have not been granted access or the file(s) have a retention period that has expired or the files are not Globus accessible.
You may request access to any restricted file(s) by clicking the Request Access button.
Close
Download Options
The files selected are too large to download as a ZIP.
You can select individual files that are below the 953.7 MB download limit from the files table, or use the
Data Access API
for programmatic access to the files.
Select File(s)
Please select a file or files to be downloaded.
Close
Inaccessible Files Selected
The selected file(s) may not be downloaded because you have not been granted access or the file(s) have a retention period that has expired.
Click Continue to download the files you have access to download.
Continue
Cancel
Ineligible Files Selected
Some file(s) cannot be transferred. (They are restricted, embargoed, with an expired retention period, or not Globus accessible.)
Click Continue to transfer the elligible files.
Continue
Cancel
Delete Dataset
Are you sure you want to delete this dataset and all of its files? You cannot undelete this dataset.
Continue
Cancel
Delete Draft Version
Are you sure you want to delete this draft version? Files will be reverted to the most recently published version. You cannot undelete this draft.
Continue
Cancel
Unpublished Dataset Preview URL
Preview URL can only be used with unpublished versions of datasets.
Cancel
Unpublished Dataset Preview URL
Are you sure you want to disable the Preview URL? If you have shared the Preview URL with others they will no longer be able to use it to access your unpublished dataset.
Yes, Disable General Preview URL
Cancel
Delete Files
The file(s) will be deleted after you click on the Delete button.
Files will not be removed from previously published versions of the dataset.
Delete
Cancel
Compute
This dataset contains restricted files you may not compute on because you have not been granted access.
Close
Deaccession Dataset
Are you sure you want to deaccession? This is permanent and the selected version(s) will no longer be viewable by the public.
No
Deaccession Dataset
Are you sure you want to deaccession this dataset? This is permanent an it will no longer be viewable by the public.
No
Version Differences Details
Please select two versions to view the differences.
Close
Version Differences Details
Version:
Last Updated:
Version:
Last Updated:
Done
Select File(s)
Please select a file or files for access request.
Close
Select File(s)
Embargoed files cannot be accessed. Please select an unembargoed file or files for your access request.
Close
Edit Tags
Select existing file tags or create new tags to describe your files. Each file can have more than one tag.
Save Changes
Cancel
Request Access
You need to
Log In
to request access.
Close
Dataset Terms
Please confirm and/or complete the information needed below in order to request access to files in this dataset.
This dataset is made available under the following terms. Please confirm and/or complete the information needed below in order to continue.
License/Data Use Agreement
Our
Community Norms
as well as good scientific practices expect that proper credit is given via citation. Please use the data citation shown on the dataset page.
CC BY-NC-SA 4.0
Preview Guestbook
Upon downloading files the guestbook asks for the following information.
Guestbook Name
Collected Data
Account Information
Close
Package File Download
Use the Download URL in a Wget command or a download manager to download this package file. Download via web browser is not recommended.
User Guide - Downloading a Dataverse Package via URL
Download URL
https://dataverse.lib.virginia.edu/api/access/datafile/
Close
Compute Batch
Clear Batch
ui-button
Dataset
Persistent Identifier
Change Compute Batch
Compute Batch
Cancel
Submit for Review
You will not be able to make changes to this dataset while it is in review.
Submit
Cancel
Publish Dataset
Are you sure you want to republish this dataset?
Select if this is a minor or major version update.
Minor Release (2.2)
Major Release (3.0)
Continue
Cancel
Publish Dataset
This dataset cannot be published until
Cell Maps for Artificial Intelligence
is published by its administrator.
Close
Publish Dataset
This dataset cannot be published until
Cell Maps for Artificial Intelligence
and
School of Medicine
are published.
Close
Return to Author
Return this dataset to contributor for modification. The reason for return entered below will be sent by email to the author.
Continue
Cancel
Add/Edit a Version Note
Styled Citation
Copyright © 2025, by the Rector and Visitors of the University of Virginia |
Terms of Use
|
Privacy Policy
Powered by
v. 6.6 build 1829-192cdc4
Contact University of Virginia Dataverse Support
To
University of Virginia Dataverse Support
From
Subject
Message
Please fill this out to prove you are not a robot.
8 + 3 =
Send Message
Cancel


================================================================================

FILE: dataverse_10.18130_V3_K7TGEM_row16.txt
PATH: data/preprocessed/individual/CM4AI/dataverse_10.18130_V3_K7TGEM_row16.txt
SIZE: 29429 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: CM4AI
Source ID: october_2025_dataverse_release
Source type: data resource
Source URL: https://dataverse.lib.virginia.edu/dataset.xhtml?persistentId=doi:10.18130/V3/K7TGEM
Raw file: data/raw/CM4AI/dataverse_10.18130_V3_K7TGEM_row16.html
--------------------------------------------------------------------------------
Cell Maps for Artificial Intelligence - October 2025 Data Release (Beta) - Cell Maps for Artificial Intelligence
Skip to main content
Toggle navigation
Search
Search
About
User Guide
Support
Log In
Cell Maps for Artificial Intelligence
This collection is under review for potential modification in compliance with Administration directives.
University of Virginia Dataverse
>
LibraData: UVa's Scholarly Research
>
School of Medicine
>
Cell Maps for Artificial Intelligence
>
Cell Maps for Artificial Intelligence - October 2025 Data Release (Beta)
Version 2.1
Clark, T; Parker, J; Al Manir, S; Axelsson, U; Ballllosero Navarro, F; Chinn, B; Churas, CP; Dailamy, A; Doctor, Y; Fall, J; Forget, A; Gao, J; Hansen, JN; Hu, M; Johannesson, A; Khaliq, H; Lee, YH; Lenkiewicz, J; Levinson, MA; Marquez, C; Metallo, C; Muralidharan, M; Nourreddine, S; Niestroy, J; Obernier, K; Pan, E; Polacco, B; Pratt, D; Qian, G; Schaffer, L; Sigaeva, A; Thaker, S; Zhang, Y; Bélisle-Pipon, JC; Brandt, C; Chen, JY; Ding, Y; Fodeh, S; Krogan, N; Lundberg, E; Mali, P; Payne-Foster, P; Ratcliffe, S; Ravitsky, V; Sali, A; Schulz, W; Ideker, T, 2025, "Cell Maps for Artificial Intelligence - October 2025 Data Release (Beta)",
https://doi.org/10.18130/V3/K7TGEM
, University of Virginia Dataverse, V2
Cite Dataset
Download EndNote XML
Download RIS
Download BibTeX
View Styled Citation
Learn about
Data Citation Standards
.
Access Dataset
The dataset is too large to download. Please select the files you need from the files table.
Contact Owner
Share
Dataset Metrics
405 Downloads
Dataset Description
Description
This dataset is the October 2025 Data Release of Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org), the Functional Genomics Grand Challenge in the NIH Bridge2AI program. CM4AI is generating multi-modal data including protein-protein interaction (PPI), spatial localization, and genetic perturbation data in MDA-MB-468 breast cancer cells (+/- paclitaxel or vorinostat) and iPSCs (+/- differentiation). This Beta release includes:
Perturb-seq data for MDA-MB-468 breast cancer cells +/- treatment and undifferentiated (parental) KOLF2.1J iPSCs
SEC-MS data for MDA-MB-468 breast cancer cells +/- treatment, undifferentiated KOLF2.1J iPSCs, and iPSC-derived neuron progenitor cells (NPCs), neurons, and cardiomyocytes
IF images in MDA-MB-468 breast cancer cells +/- treatment
External Data Links
Access external data resources related to this dataset:
Perturb-seq data in KOLF2.1J iPSCs (undifferentiated):
Embargoed
Perturb-seq data in MDA-MB-468 breast cancer cells (+/- treatment):
Embargoed
SEC-MS data in KOLF2.1J iPSCs (undifferentiated, NPC, neuron, and cardiomyocyte):
MassIVE Repository
SEC-MS data in MDA-MB-468 breast cancer cells (+/- treatment):
MassIVE Repository
Data Governance & Ethics
Human Subjects:
No
De-identified Samples:
Yes
FDA Regulated:
No
Data Governance Committee:
Jillian Parker (jillianparker@health.ucsd.edu)
Ethical Review:
Vardit Ravitsky (ravitskyv@thehastingscenter.org) and Jean-Christophe Belisle-Pipon (jean-christophe_belisle-pipon@sfu.ca)
Completeness
These data are not yet in completed final form:
Some datasets are under temporary pre-publication embargo
Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap
Computed cell maps not included in this release
Maintenance Plan
Dataset will be regularly updated and augmented through the end of the project in November 2026
Updates on a quarterly basis
Long term preservation in the University of Virginia Dataverse, supported by committed institutional funds
Intended Use
This dataset is intended for:
AI-ready datasets to support research in functional genomics
AI model training
Cellular process analysis
Cell architectural changes and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations
Limitations
Researchers should be aware of inherent limitations:
This is an interim release
Does not contain predicted cell maps, which will be added in future releases
The current release is most suitable for bioinformatics analysis of the individual datasets
Requires domain expertise for meaningful analysis
Prohibited Uses
These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval
Potential Sources of Bias
Users should be aware of potential biases:
Data in this release was derived from commercially available de-identified human cell lines
Does not represent all biological variants which may be seen in the population at large
(2025-06-30)
Subject
Medicine, Health and Life Sciences
Keyword
AI, affinity purification, AP-MS, artificial intelligence, breast cancer, Bridge2AI, cardiomyocyte, CM4AI, CRISPR/Cas9, induced pluripotent stem cell, iPSC, KOLF2.1J, machine learning, mass spectroscopy, MDA-MB-468, neural progenitor cell, NPC, neuron, paclitaxel, perturb-seq, perturbation sequencing, protein-protein interaction, protein localization, single-cell RNA sequencing, scRNAseq, SEC-MS, size exclusion chromatography, subcellular imaging, vorinostat
Related Publication
References: Clark T, Parker J, Schaffer L, Obernier K, Al Manir S, Churas CP, Dailamy A, Doctor Y, Forget A, Hansen JN, Hu M, Lenkiewicz J, Levinson MA, Marquez C, Nourreddine S, Niestroy J, Pratt D, Qian G, Thaker S, Bélisle-Pipon JC, Brandt C, Chen J, Ding Y, Fodeh S, Krogan N, Lundberg E, Mali P, Payne-Foster P, Ratcliffe S, Ravitsky V, Sali A, Schulz W, Ideker T. Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines. 2024.doi: http://doi.org/10.1101/2024.05.21.589311
License/Data Use Agreement
CC BY-NC-SA 4.0
ui-button
Files
Metadata
Terms
Versions
Change View
Table
Tree
Search
Filter by
File Type:
All
All
Archive (8)
Access:
All
All
Public (8)
Sort
Name (A-Z)
Name (Z-A)
Newest
Oldest
Size
Type
1 to 8 of 8 Files
Download
cm4ai_mass-spec_KOLF2.zip
ZIP Archive
- 23.8 MB
Published Oct 31, 2025
51 Downloads
MD5: fb04933a21e4395cce56930a014c6b4e
This dataset was generated by size exclusion chromatography-mass spectroscopy (SEC-MS) on undifferentiated KOLF2.1J human induced pluripotent stem cells (hiPSCs), in the Nevan Krogan laboratory at the University of California San Francisco.
Preview "cm4ai_mass-spec_KOLF2.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_mass-spec_MDA-MB-468.zip
ZIP Archive
- 23.0 MB
Published Oct 31, 2025
43 Downloads
MD5: 662d62ced9d379f7024e5c6d55859fbb
This dataset was generated by size exclusion chromatography-mass spectroscopy (SEC-MS) following the treatment of vorinostat or paclitaxel on MDA-MB468 human breast cancer cells, in the Nevan Krogan laboratory at the University of California San Francisco.
Preview "cm4ai_mass-spec_MDA-MB-468.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_perturb-seq_KOLF2_cell_atlas.zip
ZIP Archive
- 24.6 KB
Published Oct 31, 2025
49 Downloads
MD5: 15dc59312466dbfac85f59b57d27eee6
This dataset represents an expressed genome-scale CRISPRi Perturbation Cell Atlas in KOLF2.1J human induced pluripotent stem cells (hiPSCs) mapping transcriptional and fitness phenotypes associated with 11,739 targeted genes.
Preview "cm4ai_perturb-seq_KOLF2_cell_atlas.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_perturb-seq_KOLF2_raw_sra.zip
ZIP Archive
- 79.6 KB
Published Dec 22, 2025
32 Downloads
MD5: 1cfc4e8e2ead7513a035cc7730046ebb
This dataset represents raw sequence data from an expressed genome-scale CRISPRi Perturbation Cell Atlas in KOLF2.1J human induced pluripotent stem cells (hiPSCs) mapping transcriptional and fitness phenotypes associated with 11,739 targeted genes.
Preview "cm4ai_perturb-seq_KOLF2_raw_sra.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_release_metadata.zip
ZIP Archive
- 203.5 KB
Published Dec 22, 2025
47 Downloads
MD5: c14cc7aaaa0e7231a0d3ec296ff58518
Preview "cm4ai_release_metadata.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_ifimages_MDA-MB-468_paclitaxel.zip
Images/
ZIP Archive
- 3.8 GB
Published Oct 31, 2025
53 Downloads
MD5: 0d972b80744344ddeede516a0cf6e3d7
This data set displays the spatial localization of 464 proteins of interest in cells of the breast cancer cell line MDA-MB-468 treated with paclitaxel as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).
Preview "Images/cm4ai_ifimages_MDA-MB-468_paclitaxel.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_ifimages_MDA-MB-468_untreated.zip
Images/
ZIP Archive
- 4.6 GB
Published Oct 31, 2025
59 Downloads
MD5: a98affcc05429650c6bb3906cd836d55
This data set displays the spatial localization of 464 proteins of interest in cells of the breast cancer cell line MDA-MB-468 treated as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).
Preview "Images/cm4ai_ifimages_MDA-MB-468_untreated.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_ifimages_MDA-MB-468_vorinostat.zip
Images/
ZIP Archive
- 4.2 GB
Published Oct 31, 2025
44 Downloads
MD5: ad4e68ccc14b0f3349dad3321e7b81b2
This data set displays the spatial localization of 464 proteins of interest in cells of the breast cancer cell line MDA-MB-468 treated with vorinostat as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).
Preview "Images/cm4ai_ifimages_MDA-MB-468_vorinostat.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
Export Metadata
OAI_ORE
DataCite
OpenAIRE
Schema.org JSON-LD
DDI Codebook v2
Dublin Core
Croissant
DDI HTML Codebook
JSON
Citation Metadata
Persistent Identifier
doi:10.18130/V3/K7TGEM
Publication Date
2025-10-31
Title
Cell Maps for Artificial Intelligence - October 2025 Data Release (Beta)
Author
Clark, T
University of Virginia
ORCID
https://orcid.org/0000-0003-4060-7360
Parker, J
University of California, San Diego
ORCID
https://orcid.org/0000-0003-4535-3486
Al Manir, S
University of Virginia
ORCID
https://orcid.org/0000-0003-4647-3877
Axelsson, U
KTH Royal Institute of Technology,
Ballllosero Navarro, F
Stanford University
ORCID
https://orcid.org/0000-0002-4180-422X
Chinn, B
University of California San Diego
Churas, CP
University of California San Diego
https://orcid.org/0000-0001-9998-705X
Dailamy, A
University of California, San Diego
ORCID
https://orcid.org/0000-0002-6711-8260
Doctor, Y
University of California, San Diego
ORCID
https://orcid.org/0009-0009-0483-7506
Fall, J
KTH - Royal Institute of Technology
Forget, A
University of California San Francisco
ORCID
https://orcid.org/0000-0003-0223-0312
Gao, J
University of California San Diego
ORCID
https://orcid.org/0000-0002-6311-3526
Hansen, JN
Stanford University
ORCID
https://orcid.org/0000-0002-4650-9094
Hu, M
University of California San Diego
https://orcid.org/0000-0002-1571-8029
Johannesson, A
KTH - Royal Institute of Technology
Khaliq, H
University of California San Diego
Lee, YH
University of California San Diego
ORCID
https://orcid.org/0000-0003-0917-355X
Lenkiewicz, J
University of California San Diego
https://orcid.org/0000-0001-7252-8638
Levinson, MA
University of Virginia
ORCID
https://orcid.org/0000-0003-0384-8499
Marquez, C
University of California San Diego
ORCID
0000-0003-3960-420X
Metallo, C
University of California San Diego
ORCID
https://orcid.org/0000-0003-2404-3040
Muralidharan, M
University of California San Francisco
Nourreddine, S
University of California San Diego
https://orcid.org/0000-0003-3881-7588
Niestroy, J
University of Virginia
ORCID
https://orcid.org/0000-0002-1103-3882
Obernier, K
University of California San Francisco
ORCID
https://orcid.org/0000-0002-4025-1299
Pan, E
University of California San Diego
Polacco, B
University of California San Francisco
Pratt, D
University of California San Diego
ORCID
https://orcid.org/0000-0002-1471-9513
Qian, G
University of California San Diego
ORCID
https://orcid.org/0009-0005-4217-2745
Schaffer, L
University of California San Diego
ORCID
https://orcid.org/0000-0001-6339-9141
Sigaeva, A
KTH Royal Institute of Technology
ORCID
https://orcid.org/0000-0003-3361-3797
Thaker, S
University of Alabama at Birmingham
ORCID
https://orcid.org/0000-0001-6730-2773
Zhang, Y
University of California San Diego
Bélisle-Pipon, JC
Simon Fraser University
ORCID
https://orcid.org/0000-0002-8965-8153
Brandt, C
Yale University
ORCID
https://orcid.org/0000-0001-8179-1796
Chen, JY
The University of Alabama at Birmingham
ORCID
https://orcid.org/0000-0002-6112-415X
Ding, Y
University of Texas at Austin
ORCID
https://orcid.org/0000-0003-2567-2009
Fodeh, S
Yale University
ORCID
https://orcid.org/0000-0003-4664-3143
Krogan, N
University of California San Francisco
ORCID
https://orcid.org/0000-0003-4902-337X
Lundberg, E
Stanford University
ORCID
https://orcid.org/0000-0001-7034-0850
Mali, P
University of California San Diego
https://orcid.org/0000-0002-3383-1287
Payne-Foster, P
University of Alabama
ORCID
https://orcid.org/0000-0002-3508-3577
Ratcliffe, S
University of Virginia
ORCID
https://orcid.org/0000-0002-6644-8284
Ravitsky, V
University of Montreal
ORCID
https://orcid.org/0000-0002-7080-8801
Sali, A
University of California San Diego
ORCID
https://orcid.org/0000-0003-0435-6197
Schulz, W
Yale University
ORCID
https://orcid.org/0000-0002-2048-4028
Ideker, T
University of California San Diego
ORCID
https://orcid.org/0000-0002-1708-8454
Point of Contact
Use email button above to contact.
Ideker, Trey (University of California San Diego)
Dataset Description
Description
This dataset is the October 2025 Data Release of Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org), the Functional Genomics Grand Challenge in the NIH Bridge2AI program. CM4AI is generating multi-modal data including protein-protein interaction (PPI), spatial localization, and genetic perturbation data in MDA-MB-468 breast cancer cells (+/- paclitaxel or vorinostat) and iPSCs (+/- differentiation). This Beta release includes:
Perturb-seq data for MDA-MB-468 breast cancer cells +/- treatment and undifferentiated (parental) KOLF2.1J iPSCs
SEC-MS data for MDA-MB-468 breast cancer cells +/- treatment, undifferentiated KOLF2.1J iPSCs, and iPSC-derived neuron progenitor cells (NPCs), neurons, and cardiomyocytes
IF images in MDA-MB-468 breast cancer cells +/- treatment
External Data Links
Access external data resources related to this dataset:
Perturb-seq data in KOLF2.1J iPSCs (undifferentiated):
Embargoed
Perturb-seq data in MDA-MB-468 breast cancer cells (+/- treatment):
Embargoed
SEC-MS data in KOLF2.1J iPSCs (undifferentiated, NPC, neuron, and cardiomyocyte):
MassIVE Repository
SEC-MS data in MDA-MB-468 breast cancer cells (+/- treatment):
MassIVE Repository
Data Governance & Ethics
Human Subjects:
No
De-identified Samples:
Yes
FDA Regulated:
No
Data Governance Committee:
Jillian Parker (jillianparker@health.ucsd.edu)
Ethical Review:
Vardit Ravitsky (ravitskyv@thehastingscenter.org) and Jean-Christophe Belisle-Pipon (jean-christophe_belisle-pipon@sfu.ca)
Completeness
These data are not yet in completed final form:
Some datasets are under temporary pre-publication embargo
Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap
Computed cell maps not included in this release
Maintenance Plan
Dataset will be regularly updated and augmented through the end of the project in November 2026
Updates on a quarterly basis
Long term preservation in the University of Virginia Dataverse, supported by committed institutional funds
Intended Use
This dataset is intended for:
AI-ready datasets to support research in functional genomics
AI model training
Cellular process analysis
Cell architectural changes and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations
Limitations
Researchers should be aware of inherent limitations:
This is an interim release
Does not contain predicted cell maps, which will be added in future releases
The current release is most suitable for bioinformatics analysis of the individual datasets
Requires domain expertise for meaningful analysis
Prohibited Uses
These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval
Potential Sources of Bias
Users should be aware of potential biases:
Data in this release was derived from commercially available de-identified human cell lines
Does not represent all biological variants which may be seen in the population at large
(2025-06-30)
Subject
Medicine, Health and Life Sciences
Keyword
AI
http://purl.obolibrary.org/obo/NCIT_C16309
(NCI Thesaurus)
affinity purification
http://www.bioassayontology.org/bao#BAO_0002603
(BioAssay Ontology (BAO))
AP-MS
http://www.ebi.ac.uk/swo/SWO_1100012
(Software Ontology)
artificial intelligence
http://purl.obolibrary.org/obo/NCIT_C16309
(NCI Thesaurus)
breast cancer
http://purl.bioontology.org/ontology/LNC/LA14283-8
(LOINC)
Bridge2AI
cardiomyocyte
http://purl.obolibrary.org/obo/CL_0000746
CM4AI
CRISPR/Cas9
http://www.bioassayontology.org/bao#BAO_0010249
(Bioassay Ontology (BAO))
induced pluripotent stem cell
http://www.ebi.ac.uk/efo/EFO_0004905
(Experimental Factor Ontology (EFO))
iPSC
http://www.ebi.ac.uk/efo/EFO_0004905
(Experimental Factor Ontology (EFO))
KOLF2.1J
machine learning
http://purl.obolibrary.org/obo/OBI_0002587
(Ontology of Biomedical Investigations (OBI))
http://purl.obolibrary.org/obo/obi.owl
mass spectroscopy
http://purl.bioontology.org/ontology/MESH/D013058
(Medical Subject Headings (MeSH))
MDA-MB-468
neural progenitor cell
http://purl.obolibrary.org/obo/CL_0011020
(Cell Ontology (CL))
NPC
http://purl.obolibrary.org/obo/CL_0011020
(Cell Ontology (CL))
http://purl.obolibrary.org/obo/cl.owl
neuron
http://purl.obolibrary.org/obo/CL_0000540
(Cell Ontology (CL))
http://purl.obolibrary.org/obo/cl.owl
paclitaxel
http://purl.obolibrary.org/obo/CHEBI_45863
(Chemical Entitites of Biological Interest (CHEBI))
perturb-seq
http://www.ebi.ac.uk/efo/EFO_0008860
(Experimental Factor Ontology (EFO))
perturbation sequencing
http://www.ebi.ac.uk/efo/EFO_0008860
(Experimental Factor Ontology (EFO))
protein-protein interaction
http://purl.obolibrary.org/obo/NCIT_C18469
(NCI Thesaurus (NCIT))
protein localization
http://purl.obolibrary.org/obo/GO_0008104
(Gene Ontology (GO))
http://purl.obolibrary.org/obo/go/extensions/go-plus.owl
single-cell RNA sequencing
http://www.ebi.ac.uk/efo/EFO_0008913
(Experimental Factor Ontology (EFO))
scRNAseq
http://www.ebi.ac.uk/efo/EFO_0008913
(Experimental Factor Ontology (EFO))
SEC-MS
size exclusion chromatography
subcellular imaging
vorinostat
http://purl.obolibrary.org/obo/CHEBI_45716
(Chemical Entitites of Biological Interest (CHEBI))
Related Publication
References: Clark T, Parker J, Schaffer L, Obernier K, Al Manir S, Churas CP, Dailamy A, Doctor Y, Forget A, Hansen JN, Hu M, Lenkiewicz J, Levinson MA, Marquez C, Nourreddine S, Niestroy J, Pratt D, Qian G, Thaker S, Bélisle-Pipon JC, Brandt C, Chen J, Ding Y, Fodeh S, Krogan N, Lundberg E, Mali P, Payne-Foster P, Ratcliffe S, Ravitsky V, Sali A, Schulz W, Ideker T. Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines. 2024. doi http://doi.org/10.1101/2024.05.21.589311
Nourreddine S, Doctor Y, Dailamy A, Forget A, Lee YH, Chinn B, Khaliq H, Polacco B, Muralidharan M, Pan E, Zhang Y, Sigaeva A, Hansen JN, Gao J, Parker JA, Obernier K, Clark T, Chen JY, Metallo C, Lundberg E, Ideker T, Krogan N, Mali P. A PERTURBATION CELL ATLAS OF HUMAN INDUCED PLURIPOTENT STEM CELLS. bioRxiv. 2024 Nov 4;2024.11.03.621734. PMCID: PMC11580897 doi https://doi.org/10.1101/2024.11.03.621734
Data Creation Date
2025-02-27
Production Location
University of California San Diego; University of California San Francisco; Stanford University; University of Virginia
Funding Information
National Institutes of Health: 1OT2OD032742-01
Depositor
Niestroy, Justin
Deposit Date
2025-02-27
Dataset Terms
License/Data Use Agreement
Our
Community Norms
as well as good scientific practices expect that proper credit is given via citation. Please use the data citation shown on the dataset page.
CC BY-NC-SA 4.0
View Differences
Direct
Dataset Version
Summary
Contributors
Published on
No records found.
Edit File
This file has already been deleted (or replaced) in the current version. It may not be edited.
Close
Restrict Access
Restricting limits access to published files. People who want to use the restricted files can request access by default.
If you disable request access, you must add information about access to the Terms of Access field.
Learn about restricting files and dataset access in the
User Guide
.
Request Access
Enable access request
You must enable request access or add terms of access to restrict file access.
Terms of Access for Restricted Files
Save Changes
Cancel
Edit Embargo
The selected file or files have already been published. Contact an administrator to change the embargo date or reason of the file or files.
Cancel
Edit Retention Period
The selected file or files have already been published. Contact an administrator to change the retention period date or reason of the file or files.
Cancel
Delete Files
The file will be deleted after you click on the Delete button.
Files will not be removed from previously published versions of the dataset.
Delete
Cancel
Continue
Cancel
Select File(s)
Please select one or more files.
Close
Share Dataset
Share this dataset on your favorite social media networks.
Close
Continue
Cancel
Dataset Citations
Citations for this dataset are retrieved from Crossref via DataCite using Make Data Count standards. For more information about dataset metrics, please refer to the
User Guide
.
Sorry, no citations were found.
Close
Inaccessible Files Selected
The selected file(s) may not be downloaded because you have not been granted access or the file(s) have a retention period that has expired or the files can only be transferred via Globus.
You may request access to any restricted file(s) by clicking the Request Access button.
Close
Ineligible Files Selected
The selected file(s) may not be transferred because you have not been granted access or the file(s) have a retention period that has expired or the files are not Globus accessible.
You may request access to any restricted file(s) by clicking the Request Access button.
Close
Download Options
The files selected are too large to download as a ZIP.
You can select individual files that are below the 1.9 GB download limit from the files table, or use the
Data Access API
for programmatic access to the files.
Select File(s)
Please select a file or files to be downloaded.
Close
Inaccessible Files Selected
The selected file(s) may not be downloaded because you have not been granted access or the file(s) have a retention period that has expired.
Click Continue to download the files you have access to download.
Continue
Cancel
Ineligible Files Selected
Some file(s) cannot be transferred. (They are restricted, embargoed, with an expired retention period, or not Globus accessible.)
Click Continue to transfer the elligible files.
Continue
Cancel
Delete Dataset
Are you sure you want to delete this dataset and all of its files? You cannot undelete this dataset.
Continue
Cancel
Delete Draft Version
Are you sure you want to delete this draft version? Files will be reverted to the most recently published version. You cannot undelete this draft.
Continue
Cancel
Unpublished Dataset Preview URL
Preview URL can only be used with unpublished versions of datasets.
Cancel
Unpublished Dataset Preview URL
Are you sure you want to disable the Preview URL? If you have shared the Preview URL with others they will no longer be able to use it to access your unpublished dataset.
Yes, Disable General Preview URL
Cancel
Delete Files
The file(s) will be deleted after you click on the Delete button.
Files will not be removed from previously published versions of the dataset.
Delete
Cancel
Compute
This dataset contains restricted files you may not compute on because you have not been granted access.
Close
Deaccession Dataset
Are you sure you want to deaccession? This is permanent and the selected version(s) will no longer be viewable by the public.
No
Deaccession Dataset
Are you sure you want to deaccession this dataset? This is permanent an it will no longer be viewable by the public.
No
Version Differences Details
Please select two versions to view the differences.
Close
Version Differences Details
Version:
Last Updated:
Version:
Last Updated:
Done
Select File(s)
Please select a file or files for access request.
Close
Select File(s)
Embargoed files cannot be accessed. Please select an unembargoed file or files for your access request.
Close
Edit Tags
Select existing file tags or create new tags to describe your files. Each file can have more than one tag.
Save Changes
Cancel
Request Access
You need to
Log In
to request access.
Close
Dataset Terms
Please confirm and/or complete the information needed below in order to request access to files in this dataset.
This dataset is made available under the following terms. Please confirm and/or complete the information needed below in order to continue.
License/Data Use Agreement
Our
Community Norms
as well as good scientific practices expect that proper credit is given via citation. Please use the data citation shown on the dataset page.
CC BY-NC-SA 4.0
Preview Guestbook
Upon downloading files the guestbook asks for the following information.
Guestbook Name
Collected Data
Account Information
Close
Package File Download
Use the Download URL in a Wget command or a download manager to download this package file. Download via web browser is not recommended.
User Guide - Downloading a Dataverse Package via URL
Download URL
https://dataverse.lib.virginia.edu/api/access/datafile/
Close
Compute Batch
Clear Batch
ui-button
Dataset
Persistent Identifier
Change Compute Batch
Compute Batch
Cancel
Submit for Review
You will not be able to make changes to this dataset while it is in review.
Submit
Cancel
Publish Dataset
Are you sure you want to republish this dataset?
Select if this is a minor or major version update.
Minor Release (2.2)
Major Release (3.0)
Continue
Cancel
Publish Dataset
This dataset cannot be published until
Cell Maps for Artificial Intelligence
is published by its administrator.
Close
Publish Dataset
This dataset cannot be published until
Cell Maps for Artificial Intelligence
and
School of Medicine
are published.
Close
Return to Author
Return this dataset to contributor for modification. The reason for return entered below will be sent by email to the author.
Continue
Cancel
Curation Status History
Status
Date
Assigner
No records found.
Add/Edit a Version Note
Styled Citation
Copyright © 2026, by the Rector and Visitors of the University of Virginia |
Terms of Use
|
Privacy Policy
Powered by
v. 6.9 build 2027-e2021d3
Contact University of Virginia Dataverse Support
To
University of Virginia Dataverse Support
From
Subject
Message
Please fill this out to prove you are not a robot.
8 + 9 =
Send Message
Cancel


================================================================================

FILE: dataverse_10.18130_V3_HIGT4C_2026-07-24.txt
PATH: data/preprocessed/individual/CM4AI/dataverse_10.18130_V3_HIGT4C_2026-07-24.txt
SIZE: 27784 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: CM4AI
Source ID: june_2026_dataverse_release
Source type: data resource
Source URL: https://dataverse.lib.virginia.edu/dataset.xhtml?persistentId=doi:10.18130/V3/HIGT4C
Raw file: data/raw/CM4AI/dataverse_10.18130_V3_HIGT4C_2026-07-24.html
--------------------------------------------------------------------------------
Cell Maps for Artificial Intelligence - June 2026 Data Release (Beta) - Cell Maps for Artificial Intelligence
Skip to main content
Toggle navigation
Search
Search
About
User Guide
Support
Log In
Cell Maps for Artificial Intelligence
This collection is under review for potential modification in compliance with Administration directives.
University of Virginia Dataverse
>
LibraData: UVa's Scholarly Research
>
School of Medicine
>
Cell Maps for Artificial Intelligence
>
Cell Maps for Artificial Intelligence - June 2026 Data Release (Beta)
Version 2.0
Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Marquez C; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Polacco B; Pratt D; Qian G; Schaffer L; Sigaeva A; Thaker S; Zhang Y; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2026, "Cell Maps for Artificial Intelligence - June 2026 Data Release (Beta)",
https://doi.org/10.18130/V3/HIGT4C
, University of Virginia Dataverse, V2
Cite Dataset
Download EndNote XML
Download RIS
Download BibTeX
View Styled Citation
Learn about
Data Citation Standards
.
Access Dataset
The dataset is too large to download. Please select the files you need from the files table.
Contact Owner
Share
Dataset Metrics
181 Downloads
Dataset Description
Description
This dataset is the June 2026 Data Release of Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org), the Functional Genomics Grand Challenge in the NIH Bridge2AI program. CM4AI is generating multi-modal data including protein-protein interaction (PPI), spatial localization, and genetic perturbation data in MDA-MB-468 breast cancer cells (+/- paclitaxel or vorinostat) and iPSCs (+/- differentiation). This Beta release includes:
Perturb-seq data for MDA-MB-468 breast cancer cells +/- treatment and undifferentiated (parental) KOLF2.1J iPSCs
SEC-MS data for MDA-MB-468 breast cancer cells +/- treatment, undifferentiated KOLF2.1J iPSCs, and iPSC-derived neuron progenitor cells (NPCs), neurons, and cardiomyocytes
AP-MS data for MDA-MB-468 breast cancer cells + treatment
IF images in MDA-MB-468 breast cancer cells +/- treatment
External Data Links
Access external data resources related to this dataset:
Perturb-seq data in KOLF2.1J iPSCs (undifferentiated):
Embargoed
Perturb-seq data in MDA-MB-468 breast cancer cells (+/- treatment):
Embargoed
CRISPRi Perturbation Atlas of iPSCs:
Figshare
SEC-MS data in KOLF2.1J iPSCs (undifferentiated, NPC, neuron, and cardiomyocyte):
MassIVE Repository
SEC-MS data in MDA-MB-468 breast cancer cells (+/- treatment):
MassIVE Repository
AP-MS data in MDA-MB-468 breast cancer cells (treated with paclitaxel):
MassIVE Repository
AP-MS data in MDA-MB-468 breast cancer cells (treated with vorinostat):
MassIVE Repository
Data Governance & Ethics
Human Subjects:
No
De-identified Samples:
Yes
FDA Regulated:
No
Data Governance Committee:
Jillian Parker (jillianparker@health.ucsd.edu)
Ethical Review:
Vardit Ravitsky (ravitskyv@thehastingscenter.org) and Jean-Christophe Belisle-Pipon (jean-christophe_belisle-pipon@sfu.ca)
Completeness
These data are not yet in completed final form:
Some datasets are under temporary pre-publication embargo
Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap
Computed cell maps not included in this release
Maintenance Plan
Dataset will be regularly updated and augmented through the end of the project in November 2026
Updates on a quarterly basis
Long term preservation in the University of Virginia Dataverse, supported by committed institutional funds
Intended Use
This dataset is intended for:
AI-ready datasets to support research in functional genomics
AI model training
Cellular process analysis
Cell architectural changes and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations
Limitations
Researchers should be aware of inherent limitations:
This is an interim release
Does not contain predicted cell maps, which will be added in future releases
The current release is most suitable for bioinformatics analysis of the individual datasets
Requires domain expertise for meaningful analysis
Prohibited Uses
These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval
Potential Sources of Bias
Users should be aware of potential biases:
Data in this release was derived from commercially available de-identified human cell lines
Does not represent all biological variants which may be seen in the population at large
(2025-06-30)
Subject
Medicine, Health and Life Sciences
Keyword
AI, affinity purification, AP-MS, artificial intelligence, breast cancer, Bridge2AI, cardiomyocyte, CM4AI, CRISPR/Cas9, induced pluripotent stem cell, iPSC, KOLF2.1J, machine learning, mass spectroscopy, MDA-MB-468, neural progenitor cell, NPC, neuron, paclitaxel, perturb-seq, perturbation sequencing, protein-protein interaction, protein localization, single-cell RNA sequencing, scRNAseq, SEC-MS, size exclusion chromatography, subcellular imaging, vorinostat
Related Publication
References: Clark T, Parker J, Schaffer L, Obernier K, Al Manir S, Churas CP, Dailamy A, Doctor Y, Forget A, Hansen JN, Hu M, Lenkiewicz J, Levinson MA, Marquez C, Nourreddine S, Niestroy J, Pratt D, Qian G, Thaker S, Bélisle-Pipon JC, Brandt C, Chen J, Ding Y, Fodeh S, Krogan N, Lundberg E, Mali P, Payne-Foster P, Ratcliffe S, Ravitsky V, Sali A, Schulz W, Ideker T. Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines. 2024.doi: http://doi.org/10.1101/2024.05.21.589311
License/Data Use Agreement
CC BY-NC-SA 4.0
ui-button
Files
Metadata
Terms
Versions
Search
Filter by
File Type:
All
All
Archive (10)
Access:
All
All
Public (10)
Sort
Name (A-Z)
Name (Z-A)
Newest
Oldest
Size
Type
1 to 10 of 10 Files
Download
cm4ai_apms_MDA-MB-468_paclitaxel.zip
ZIP Archive
- 113.3 KB
Published Jun 17, 2026
22 Downloads
MD5: 8642317c321126213240af4837030a53
Preview "cm4ai_apms_MDA-MB-468_paclitaxel.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_apms_MDA-MB-468_vorinostat.zip
ZIP Archive
- 135.8 KB
Published Jun 17, 2026
12 Downloads
MD5: b941f28d6f19c563d1222f7aff0c57b6
Preview "cm4ai_apms_MDA-MB-468_vorinostat.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_ifimages_MDA-MB-468_paclitaxel.zip
ZIP Archive
- 3.8 GB
Published Jul 15, 2026
13 Downloads
MD5: 6c1a86520eec2696ec19444eb8a8b428
Preview "cm4ai_ifimages_MDA-MB-468_paclitaxel.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_ifimages_MDA-MB-468_untreated.zip
ZIP Archive
- 4.6 GB
Published Jul 15, 2026
13 Downloads
MD5: 6d066e6b9d762284c6aa5d0b44565568
Preview "cm4ai_ifimages_MDA-MB-468_untreated.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_ifimages_MDA-MB-468_vorinostat.zip
ZIP Archive
- 4.2 GB
Published Jul 15, 2026
14 Downloads
MD5: df79632753989ef23b0a3df6604d4af1
Preview "cm4ai_ifimages_MDA-MB-468_vorinostat.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_mass-spec_KOLF2.zip
ZIP Archive
- 171.8 KB
Published Jun 17, 2026
11 Downloads
MD5: f250bf0b04f516f7462d30432441fc0e
Preview "cm4ai_mass-spec_KOLF2.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_mass-spec_MDA-MB-468.zip
ZIP Archive
- 93.9 KB
Published Jun 17, 2026
12 Downloads
MD5: 9aed30b6e8a91b151155316aae657ea5
Preview "cm4ai_mass-spec_MDA-MB-468.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_perturb-seq_KOLF2_cell_atlas.zip
ZIP Archive
- 30.2 KB
Published Jun 17, 2026
15 Downloads
MD5: 291a362842a176a4fcccaf2cdc1dedde
Preview "cm4ai_perturb-seq_KOLF2_cell_atlas.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_perturb-seq_KOLF2_raw_sra.zip
ZIP Archive
- 73.3 KB
Published Jun 17, 2026
8 Downloads
MD5: 8bb3f365f01fb8bbbd2ac6a535059305
Preview "cm4ai_perturb-seq_KOLF2_raw_sra.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
cm4ai_release_metadata.zip
ZIP Archive
- 1.1 MB
Published Jun 17, 2026
22 Downloads
MD5: 318deb7c5edab0d2fa55e26fcb67b440
Preview "cm4ai_release_metadata.zip"
Access File
File Access
Public
Download Options
ZIP Archive
Download Metadata
Data File Citation
Download EndNote XML
Download RIS
Download BibTeX
Export Metadata
OAI_ORE
DataCite
OpenAIRE
Schema.org JSON-LD
DDI Codebook v2
Dublin Core
Croissant
DDI HTML Codebook
JSON
Citation Metadata
Persistent Identifier
doi:10.18130/V3/HIGT4C
Publication Date
2026-06-17
Title
Cell Maps for Artificial Intelligence - June 2026 Data Release (Beta)
Author
Clark T
https://ror.org/0153tk833
ORCID
https://orcid.org/0000-0003-4060-7360
Parker J
University of California, San Diego
ORCID
https://orcid.org/0000-0003-4535-3486
Al Manir S
https://ror.org/0153tk833
ORCID
https://orcid.org/0000-0003-4647-3877
Axelsson U
KTH Royal Institute of Technology,
Ballllosero Navarro F
Stanford University
ORCID
https://orcid.org/0000-0002-4180-422X
Chinn B
University of California San Diego
Churas CP
University of California San Diego
https://orcid.org/0000-0001-9998-705X
Dailamy A
University of California, San Diego
ORCID
https://orcid.org/0000-0002-6711-8260
Doctor Y
University of California, San Diego
ORCID
https://orcid.org/0009-0009-0483-7506
Fall J
KTH - Royal Institute of Technology
Forget A
University of California San Francisco
ORCID
https://orcid.org/0000-0003-0223-0312
Gao J
University of California San Diego
ORCID
https://orcid.org/0000-0002-6311-3526
Hansen JN
Stanford University
ORCID
https://orcid.org/0000-0002-4650-9094
Hu M
University of California San Diego
https://orcid.org/0000-0002-1571-8029
Johannesson A
KTH - Royal Institute of Technology
Khaliq H
University of California San Diego
Lee YH
University of California San Diego
ORCID
https://orcid.org/0000-0003-0917-355X
Lenkiewicz J
University of California San Diego
https://orcid.org/0000-0001-7252-8638
Levinson MA
https://ror.org/0153tk833
ORCID
https://orcid.org/0000-0003-0384-8499
Marquez C
University of California San Diego
ORCID
0000-0003-3960-420X
Metallo C
University of California San Diego
ORCID
https://orcid.org/0000-0003-2404-3040
Muralidharan M
University of California San Francisco
Nourreddine S
University of California San Diego
https://orcid.org/0000-0003-3881-7588
Niestroy J
https://ror.org/0153tk833
ORCID
https://orcid.org/0000-0002-1103-3882
Obernier K
University of California San Francisco
ORCID
https://orcid.org/0000-0002-4025-1299
Pan E
University of California San Diego
Polacco B
University of California San Francisco
Pratt D
University of California San Diego
ORCID
https://orcid.org/0000-0002-1471-9513
Qian G
University of California San Diego
ORCID
https://orcid.org/0009-0005-4217-2745
Schaffer L
University of California San Diego
ORCID
https://orcid.org/0000-0001-6339-9141
Sigaeva A
KTH Royal Institute of Technology
ORCID
https://orcid.org/0000-0003-3361-3797
Thaker S
University of Alabama at Birmingham
ORCID
https://orcid.org/0000-0001-6730-2773
Zhang Y
University of California San Diego
Bélisle-Pipon JC
Simon Fraser University
ORCID
https://orcid.org/0000-0002-8965-8153
Brandt C
Yale University
ORCID
https://orcid.org/0000-0001-8179-1796
Chen JY
The University of Alabama at Birmingham
ORCID
https://orcid.org/0000-0002-6112-415X
Ding Y
University of Texas at Austin
ORCID
https://orcid.org/0000-0003-2567-2009
Fodeh S
Yale University
ORCID
https://orcid.org/0000-0003-4664-3143
Krogan N
University of California San Francisco
ORCID
https://orcid.org/0000-0003-4902-337X
Lundberg E
Stanford University
ORCID
https://orcid.org/0000-0001-7034-0850
Mali P
University of California San Diego
https://orcid.org/0000-0002-3383-1287
Payne-Foster P
University of Alabama
ORCID
https://orcid.org/0000-0002-3508-3577
Ratcliffe S
https://ror.org/0153tk833
ORCID
https://orcid.org/0000-0002-6644-8284
Ravitsky V
University of Montreal
ORCID
https://orcid.org/0000-0002-7080-8801
Sali A
University of California San Diego
ORCID
https://orcid.org/0000-0003-0435-6197
Schulz W
Yale University
ORCID
https://orcid.org/0000-0002-2048-4028
Ideker T
University of California San Diego
ORCID
https://orcid.org/0000-0002-1708-8454
Point of Contact
Use email button above to contact.
Ideker Trey (University of California San Diego)
Dataset Description
Description
This dataset is the June 2026 Data Release of Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org), the Functional Genomics Grand Challenge in the NIH Bridge2AI program. CM4AI is generating multi-modal data including protein-protein interaction (PPI), spatial localization, and genetic perturbation data in MDA-MB-468 breast cancer cells (+/- paclitaxel or vorinostat) and iPSCs (+/- differentiation). This Beta release includes:
Perturb-seq data for MDA-MB-468 breast cancer cells +/- treatment and undifferentiated (parental) KOLF2.1J iPSCs
SEC-MS data for MDA-MB-468 breast cancer cells +/- treatment, undifferentiated KOLF2.1J iPSCs, and iPSC-derived neuron progenitor cells (NPCs), neurons, and cardiomyocytes
AP-MS data for MDA-MB-468 breast cancer cells + treatment
IF images in MDA-MB-468 breast cancer cells +/- treatment
External Data Links
Access external data resources related to this dataset:
Perturb-seq data in KOLF2.1J iPSCs (undifferentiated):
Embargoed
Perturb-seq data in MDA-MB-468 breast cancer cells (+/- treatment):
Embargoed
CRISPRi Perturbation Atlas of iPSCs:
Figshare
SEC-MS data in KOLF2.1J iPSCs (undifferentiated, NPC, neuron, and cardiomyocyte):
MassIVE Repository
SEC-MS data in MDA-MB-468 breast cancer cells (+/- treatment):
MassIVE Repository
AP-MS data in MDA-MB-468 breast cancer cells (treated with paclitaxel):
MassIVE Repository
AP-MS data in MDA-MB-468 breast cancer cells (treated with vorinostat):
MassIVE Repository
Data Governance & Ethics
Human Subjects:
No
De-identified Samples:
Yes
FDA Regulated:
No
Data Governance Committee:
Jillian Parker (jillianparker@health.ucsd.edu)
Ethical Review:
Vardit Ravitsky (ravitskyv@thehastingscenter.org) and Jean-Christophe Belisle-Pipon (jean-christophe_belisle-pipon@sfu.ca)
Completeness
These data are not yet in completed final form:
Some datasets are under temporary pre-publication embargo
Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap
Computed cell maps not included in this release
Maintenance Plan
Dataset will be regularly updated and augmented through the end of the project in November 2026
Updates on a quarterly basis
Long term preservation in the University of Virginia Dataverse, supported by committed institutional funds
Intended Use
This dataset is intended for:
AI-ready datasets to support research in functional genomics
AI model training
Cellular process analysis
Cell architectural changes and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations
Limitations
Researchers should be aware of inherent limitations:
This is an interim release
Does not contain predicted cell maps, which will be added in future releases
The current release is most suitable for bioinformatics analysis of the individual datasets
Requires domain expertise for meaningful analysis
Prohibited Uses
These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval
Potential Sources of Bias
Users should be aware of potential biases:
Data in this release was derived from commercially available de-identified human cell lines
Does not represent all biological variants which may be seen in the population at large
(2025-06-30)
Subject
Medicine, Health and Life Sciences
Keyword
AI
http://purl.obolibrary.org/obo/NCIT_C16309
(NCI Thesaurus)
affinity purification
http://www.bioassayontology.org/bao#BAO_0002603
(BioAssay Ontology (BAO))
AP-MS
http://www.ebi.ac.uk/swo/SWO_1100012
(Software Ontology)
artificial intelligence
http://purl.obolibrary.org/obo/NCIT_C16309
(NCI Thesaurus)
breast cancer
http://purl.bioontology.org/ontology/LNC/LA14283-8
(LOINC)
Bridge2AI
cardiomyocyte
http://purl.obolibrary.org/obo/CL_0000746
CM4AI
CRISPR/Cas9
http://www.bioassayontology.org/bao#BAO_0010249
(Bioassay Ontology (BAO))
induced pluripotent stem cell
http://www.ebi.ac.uk/efo/EFO_0004905
(Experimental Factor Ontology (EFO))
iPSC
http://www.ebi.ac.uk/efo/EFO_0004905
(Experimental Factor Ontology (EFO))
KOLF2.1J
machine learning
http://purl.obolibrary.org/obo/OBI_0002587
(Ontology of Biomedical Investigations (OBI))
http://purl.obolibrary.org/obo/obi.owl
mass spectroscopy
http://purl.bioontology.org/ontology/MESH/D013058
(Medical Subject Headings (MeSH))
MDA-MB-468
neural progenitor cell
http://purl.obolibrary.org/obo/CL_0011020
(Cell Ontology (CL))
NPC
http://purl.obolibrary.org/obo/CL_0011020
(Cell Ontology (CL))
http://purl.obolibrary.org/obo/cl.owl
neuron
http://purl.obolibrary.org/obo/CL_0000540
(Cell Ontology (CL))
http://purl.obolibrary.org/obo/cl.owl
paclitaxel
http://purl.obolibrary.org/obo/CHEBI_45863
(Chemical Entitites of Biological Interest (CHEBI))
perturb-seq
http://www.ebi.ac.uk/efo/EFO_0008860
(Experimental Factor Ontology (EFO))
perturbation sequencing
http://www.ebi.ac.uk/efo/EFO_0008860
(Experimental Factor Ontology (EFO))
protein-protein interaction
http://purl.obolibrary.org/obo/NCIT_C18469
(NCI Thesaurus (NCIT))
protein localization
http://purl.obolibrary.org/obo/GO_0008104
(Gene Ontology (GO))
http://purl.obolibrary.org/obo/go/extensions/go-plus.owl
single-cell RNA sequencing
http://www.ebi.ac.uk/efo/EFO_0008913
(Experimental Factor Ontology (EFO))
scRNAseq
http://www.ebi.ac.uk/efo/EFO_0008913
(Experimental Factor Ontology (EFO))
SEC-MS
size exclusion chromatography
subcellular imaging
vorinostat
http://purl.obolibrary.org/obo/CHEBI_45716
(Chemical Entitites of Biological Interest (CHEBI))
Related Publication
References: Clark T, Parker J, Schaffer L, Obernier K, Al Manir S, Churas CP, Dailamy A, Doctor Y, Forget A, Hansen JN, Hu M, Lenkiewicz J, Levinson MA, Marquez C, Nourreddine S, Niestroy J, Pratt D, Qian G, Thaker S, Bélisle-Pipon JC, Brandt C, Chen J, Ding Y, Fodeh S, Krogan N, Lundberg E, Mali P, Payne-Foster P, Ratcliffe S, Ravitsky V, Sali A, Schulz W, Ideker T. Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines. 2024. doi http://doi.org/10.1101/2024.05.21.589311
Nourreddine S, Doctor Y, Dailamy A, Forget A, Lee YH, Chinn B, Khaliq H, Polacco B, Muralidharan M, Pan E, Zhang Y, Sigaeva A, Hansen JN, Gao J, Parker JA, Obernier K, Clark T, Chen JY, Metallo C, Lundberg E, Ideker T, Krogan N, Mali P. A PERTURBATION CELL ATLAS OF HUMAN INDUCED PLURIPOTENT STEM CELLS. bioRxiv. 2024 Nov 4;2024.11.03.621734. PMCID: PMC11580897 doi https://doi.org/10.1101/2024.11.03.621734
Data Creation Date
2025-02-27
Production Location
University of California San Diego; University of California San Francisco; Stanford University; University of Virginia
Funding Information
National Institutes of Health: 1OT2OD032742-01
Depositor
Niestroy, Justin
Deposit Date
2025-02-27
Dataset Terms
License/Data Use Agreement
Our
Community Norms
as well as good scientific practices expect that proper credit is given via citation. Please use the data citation shown on the dataset page.
CC BY-NC-SA 4.0
Direct
Dataset Version
Summary
Contributors
Published on
No records found.
Edit File
This file has already been deleted (or replaced) in the current version. It may not be edited.
Close
Restrict Access
Restricting limits access to published files. People who want to use the restricted files can request access by default.
If you disable request access, you must add information about access to the Terms of Access field.
Learn about restricting files and dataset access in the
User Guide
.
Request Access
Enable access request
You must enable request access or add terms of access to restrict file access.
Terms of Access for Restricted Files
Save Changes
Cancel
Edit Embargo
The selected file or files have already been published. Contact an administrator to change the embargo date or reason of the file or files.
Cancel
Edit Retention Period
The selected file or files have already been published. Contact an administrator to change the retention period date or reason of the file or files.
Cancel
Delete Files
The file will be deleted after you click on the Delete button.
Files will not be removed from previously published versions of the dataset.
Delete
Cancel
Continue
Cancel
Select File(s)
Please select one or more files.
Close
Share Dataset
Share this dataset on your favorite social media networks.
Close
Continue
Cancel
Dataset Citations
Citations for this dataset are retrieved from Crossref via DataCite using Make Data Count standards. For more information about dataset metrics, please refer to the
User Guide
.
Sorry, no citations were found.
Close
Inaccessible Files Selected
The selected file(s) may not be downloaded because you have not been granted access or the file(s) have a retention period that has expired or the files can only be transferred via Globus.
You may request access to any restricted file(s) by clicking the Request Access button.
Close
Ineligible Files Selected
The selected file(s) may not be transferred because you have not been granted access or the file(s) have a retention period that has expired or the files are not Globus accessible.
You may request access to any restricted file(s) by clicking the Request Access button.
Close
Download Options
The files selected are too large to download as a ZIP.
You can select individual files that are below the 1.9 GB download limit from the files table, or use the
Data Access API
for programmatic access to the files.
Select File(s)
Please select a file or files to be downloaded.
Close
Inaccessible Files Selected
The selected file(s) may not be downloaded because you have not been granted access or the file(s) have a retention period that has expired.
Click Continue to download the files you have access to download.
Continue
Cancel
Ineligible Files Selected
Some file(s) cannot be transferred. (They are restricted, embargoed, with an expired retention period, or not Globus accessible.)
Click Continue to transfer the elligible files.
Continue
Cancel
Delete Dataset
Are you sure you want to delete this dataset and all of its files? You cannot undelete this dataset.
Continue
Cancel
Delete Draft Version
Are you sure you want to delete this draft version? Files will be reverted to the most recently published version. You cannot undelete this draft.
Continue
Cancel
Unpublished Dataset Preview URL
Preview URL can only be used with unpublished versions of datasets.
Cancel
Unpublished Dataset Preview URL
Are you sure you want to disable the Preview URL? If you have shared the Preview URL with others they will no longer be able to use it to access your unpublished dataset.
Yes, Disable General Preview URL
Cancel
Delete Files
The file(s) will be deleted after you click on the Delete button.
Files will not be removed from previously published versions of the dataset.
Delete
Cancel
Compute
This dataset contains restricted files you may not compute on because you have not been granted access.
Close
Deaccession Dataset
Are you sure you want to deaccession? This is permanent and the selected version(s) will no longer be viewable by the public.
No
Deaccession Dataset
Are you sure you want to deaccession this dataset? This is permanent an it will no longer be viewable by the public.
No
Version Differences Details
Please select two versions to view the differences.
Close
Version Differences Details
Version:
Last Updated:
Version:
Last Updated:
Done
Select File(s)
Please select a file or files for access request.
Close
Select File(s)
Embargoed files cannot be accessed. Please select an unembargoed file or files for your access request.
Close
Edit Tags
Select existing file tags or create new tags to describe your files. Each file can have more than one tag.
Save Changes
Cancel
Request Access
You need to
Log In
to request access.
Close
Dataset Terms
Please confirm and/or complete the information needed below in order to request access to files in this dataset.
This dataset is made available under the following terms. Please confirm and/or complete the information needed below in order to continue.
License/Data Use Agreement
Our
Community Norms
as well as good scientific practices expect that proper credit is given via citation. Please use the data citation shown on the dataset page.
CC BY-NC-SA 4.0
Preview Guestbook
Upon downloading files the guestbook asks for the following information.
Guestbook Name
Collected Data
Account Information
Close
Package File Download
Use the Download URL in a Wget command or a download manager to download this package file. Download via web browser is not recommended.
User Guide - Downloading a Dataverse Package via URL
Download URL
https://dataverse.lib.virginia.edu/api/access/datafile/
Close
Compute Batch
Clear Batch
ui-button
Dataset
Persistent Identifier
Change Compute Batch
Compute Batch
Cancel
Submit for Review
You will not be able to make changes to this dataset while it is in review.
Submit
Cancel
Publish Dataset
Are you sure you want to republish this dataset?
Select if this is a minor or major version update.
Minor Release (2.1)
Major Release (3.0)
Continue
Cancel
Publish Dataset
This dataset cannot be published until
Cell Maps for Artificial Intelligence
is published by its administrator.
Close
Publish Dataset
This dataset cannot be published until
Cell Maps for Artificial Intelligence
and
School of Medicine
are published.
Close
Return to Author
Return this dataset to contributor for modification. The reason for return entered below will be sent by email to the author.
Continue
Cancel
Curation Status History
Status
Date
Assigner
No records found.
Add/Edit a Version Note
Styled Citation
Copyright © 2026, by the Rector and Visitors of the University of Virginia |
Terms of Use
|
Privacy Policy
Powered by
v. 6.9 build 2027-e2021d3
Contact University of Virginia Dataverse Support
To
University of Virginia Dataverse Support
From
Subject
Message
Please fill this out to prove you are not a robot.
1 + 8 =
Send Message
Cancel


================================================================================
CRATE EVIDENCE
================================================================================

================================================================================
FILE: CM4AI_crate_metadata_reduced.json
ROLE: crate JSON-LD with file inventories collapsed; the substantive evidence (rai:* fields, ethics, access, provenance)
SIZE: 107,351 characters
--------------------------------------------------------------------------------
{
 "@context": {
  "@vocab": "https://schema.org/",
  "EVI": "https://w3id.org/EVI#"
 },
 "@graph": [
  {
   "@id": "ro-crate-metadata.json",
   "@type": "CreativeWork",
   "conformsTo": {
    "@id": "https://w3id.org/ro/crate/1.2"
   },
   "about": {
    "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release"
   },
   "fairscapeVersion": "1.1.3"
  },
  {
   "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release",
   "@type": [
    "Dataset",
    "https://w3id.org/EVI#ROCrate"
   ],
   "name": "Cell Maps for Artificial Intelligence - June 2026 Data Release (Beta)",
   "description": "This dataset is the June 2026 Data Release of Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org), the Functional Genomics Grand Challenge in the NIH Bridge2AI program. This Beta release includes perturb-seq data in undifferentiated KOLF2.1J iPSCs; SEC-MS data in undifferentiated KOLF2.1J iPSCs and iPSC-derived NPCs, neurons, and cardiomyocytes; AP-MS data in MDA-MB-468 breast cancer cells in the presence of chemotherapy (vorinostat and paclitaxel); and IF images in MDA-MB-468 breast cancer cells in the presence and absence of chemotherapy (vorinostat and paclitaxel). CM4AI output data are packaged with provenance graphs and rich metadata as AI-ready datasets in RO-Crate format using the FAIRSCAPE framework. Data presented here will be augmented regularly through the end of the project. CM4AI is a collaboration of UCSD, UCSF, Stanford, UVA, Yale, UA Birmingham, Simon Fraser University, and the Hastings Center.",
   "keywords": [
    "AI",
    "affinity purification",
    "AP-MS",
    "artificial intelligence",
    "breast cancer",
    "Bridge2AI",
    "cardiomyocyte",
    "CM4AI",
    "CRISPR/Cas9",
    "induced pluripotent stem cell",
    "iPSC",
    "KOLF2.1J",
    "machine learning",
    "mass spectroscopy",
    "MDA-MB-468",
    "neural progenitor cell",
    "NPC",
    "neuron",
    "paclitaxel",
    "perturb-seq",
    "perturbation sequencing",
    "protein-protein interaction",
    "protein localization",
    "single-cell RNA sequencing",
    "scRNAseq",
    "SEC-MS",
    "size exclusion chromatography",
    "subcellular imaging",
    "vorinostat",
    "Artificial intelligence",
    "Breast cancer",
    "CRISPR perturbation",
    "Cell maps",
    "IPSC",
    "Machine learning",
    "Mass spectroscopy",
    "Perturb-seq",
    "Protein-protein interaction",
    "cell maps"
   ],
   "version": "1.0",
   "datePublished": "2026-06-30",
   "isPartOf": [
    {
     "@id": "https://fairscape.net/api/ark:59853/organization-university-of-california-san-diego-T4a649a1RtE"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/project-cell-maps-for-artificial-intelligence-qTdsTBd3FtA"
    }
   ],
   "hasPart": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 60,
    "id_families": {
     "0000": 36,
     "rocrate": 9,
     "0009": 2,
     "ui?ui=D001943": 1,
     "ui?ui=D057026": 1,
     "ui?ui=D064113": 1,
     "ui?ui=D013058": 1,
     "ui?ui=D005453": 1,
     "ui?ui=D017239": 1,
     "ui?ui=D000077337": 1,
     "topic_0121": 1,
     "topic_3170": 1
    },
    "sample_ids": [
     "https://orcid.org/0000-0003-4060-7360",
     "https://orcid.org/0000-0003-4535-3486",
     "https://orcid.org/0000-0003-4647-3877",
     "https://orcid.org/0000-0002-4180-422X",
     "https://orcid.org/0000-0001-9998-705X"
    ]
   },
   "author": [
    {
     "@id": "https://orcid.org/0000-0003-4060-7360"
    },
    {
     "@id": "https://orcid.org/0000-0003-4535-3486"
    },
    {
     "@id": "https://orcid.org/0000-0003-4647-3877"
    },
    "Axelsson, U",
    {
     "@id": "https://orcid.org/0000-0002-4180-422X"
    },
    "Chinn, B",
    {
     "@id": "https://orcid.org/0000-0001-9998-705X"
    },
    {
     "@id": "https://orcid.org/0000-0002-6711-8260"
    },
    {
     "@id": "https://orcid.org/0009-0009-0483-7506"
    },
    "Fall, J",
    {
     "@id": "https://orcid.org/0000-0003-0223-0312"
    },
    {
     "@id": "https://orcid.org/0000-0002-6311-3526"
    },
    {
     "@id": "https://orcid.org/0000-0002-4650-9094"
    },
    {
     "@id": "https://orcid.org/0000-0002-1571-8029"
    },
    "Johannesson, A",
    "Khaliq, H",
    {
     "@id": "https://orcid.org/0000-0003-0917-355X"
    },
    {
     "@id": "https://orcid.org/0000-0001-7252-8638"
    },
    {
     "@id": "https://orcid.org/0000-0003-0384-8499"
    },
    {
     "@id": "https://orcid.org/0000-0003-3960-420X"
    },
    {
     "@id": "https://orcid.org/0000-0003-2404-3040"
    },
    "Muralidharan, M",
    {
     "@id": "https://orcid.org/0000-0003-3881-7588"
    },
    {
     "@id": "https://orcid.org/0000-0002-1103-3882"
    },
    {
     "@id": "https://orcid.org/0000-0002-4025-1299"
    },
    "Pan, E",
    "Polacco, B",
    {
     "@id": "https://orcid.org/0000-0002-1471-9513"
    },
    {
     "@id": "https://orcid.org/0009-0005-4217-2745"
    },
    {
     "@id": "https://orcid.org/0000-0001-6339-9141"
    },
    {
     "@id": "https://orcid.org/0000-0003-3361-3797"
    },
    {
     "@id": "https://orcid.org/0000-0001-6730-2773"
    },
    "Zhang, Y",
    {
     "@id": "https://orcid.org/0000-0002-8965-8153"
    },
    {
     "@id": "https://orcid.org/0000-0001-8179-1796"
    },
    {
     "@id": "https://orcid.org/0000-0002-6112-415X"
    },
    {
     "@id": "https://orcid.org/0000-0003-2567-2009"
    },
    {
     "@id": "https://orcid.org/0000-0003-4664-3143"
    },
    {
     "@id": "https://orcid.org/0000-0003-4902-337X"
    },
    {
     "@id": "https://orcid.org/0000-0001-7034-0850"
    },
    {
     "@id": "https://orcid.org/0000-0002-3383-1287"
    },
    {
     "@id": "https://orcid.org/0000-0002-3508-3577"
    },
    {
     "@id": "https://orcid.org/0000-0002-6644-8284"
    },
    {
     "@id": "https://orcid.org/0000-0002-7080-8801"
    },
    {
     "@id": "https://orcid.org/0000-0003-0435-6197"
    },
    {
     "@id": "https://orcid.org/0000-0002-2048-4028"
    },
    {
     "@id": "https://orcid.org/0000-0002-1708-8454"
    }
   ],
   "publisher": "https://dataverse.lib.virginia.edu/",
   "principalInvestigator": "Trey Ideker",
   "funder": "National Institutes of Health: 1OT2OD032742-01, R01HG012351, R01NS131560, U54CA274502, #S10 OD026929. Department of Defense: W81XWH-22-1-0401. CIRM training: EDUC4-12804. Dutch Research Council: NWO, 019.231EN.013. National Cancer Institute: P30CA023100",
   "contactEmail": "tideker@health.ucsd.edu",
   "citation": "Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Park, S; Polacco B; Pratt D; Qian G; Schaffer, LV; Sigaeva A; Thaker S; Zhang Y; Zhao, X; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, \"Cell Maps for Artificial Intelligence - June 2026 Data Release (Beta)\", https://doi.org/10.18130/V3/HIGT4C , https://dataverse.lib.virginia.edu/, V1",
   "associatedPublication": [
    "Clark T, Parker J, Al Manir S, et al. (2024) Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines. bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311",
    "Nourreddine S, et al. (2024) A Perturbation Cell Atlas of Human Induced Pluripotent Stem Cells. bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.11.03.621734",
    "Qin, Y., Huttlin, E.L., Winsnes, C.F. et al. A multi-scale map of cell structure fusing protein images and interactions. Nature 600, 536–542 (2021). https://doi.org/10.1038/s41586-021-04115-9",
    "Schaffer LV, Hu M, Qian G, et al. Multimodal cell maps as a foundation for structural and functional genomics. Nature [Internet]. 2025 Apr 9; Available from: https://www.nature.com/articles/s41586-025-08878-3"
   ],
   "identifier": "https://doi.org/10.18130/V3/HIGT4C",
   "license": "https://creativecommons.org/licenses/by-nc-sa/4.0/",
   "conditionsOfAccess": "Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived data products should cite the Related Publications below, as well as directly citing this data collection.",
   "copyrightNotice": "Copyright (c) 2026 The Regents of the University of California except where otherwise noted. Spatial proteomics raw image data is copyright (c) 2026 The Board of Trustees of the Leland Stanford Junior University.",
   "contentSize": "19.9 TB",
   "usageInfo": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "ethicalReview": "Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.",
   "confidentialityLevel": "Unrestricted",
   "humanSubjectExemption": "Exempt — research with commercially available de-identified human cell lines does not constitute human subjects research.",
   "humanSubjectResearch": "None - data collected from commercially available cell lines",
   "dataGovernanceCommittee": "Jilian Parker",
   "about": [
    {
     "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D001943"
    },
    {
     "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D057026"
    },
    {
     "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D064113"
    },
    {
     "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D013058"
    },
    {
     "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D005453"
    },
    {
     "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D017239"
    },
    {
     "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D000077337"
    },
    {
     "@id": "http://edamontology.org/topic_0121"
    },
    {
     "@id": "http://edamontology.org/topic_3170"
    },
    {
     "@id": "http://edamontology.org/topic_3320"
    },
    {
     "@id": "http://edamontology.org/topic_3474"
    },
    {
     "@id": "https://www.cellosaurus.org/CVCL_0419"
    },
    {
     "@id": "https://www.cellosaurus.org/CVCL_B5P3"
    }
   ],
   "rai:dataLimitations": "This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.",
   "rai:dataBiases": "Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.",
   "rai:dataUseCases": "AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable biologically-driven, interpretable ML applications, for example as proposed in Ma et al. 2018 (PMID: 29505029) and Kuenzi et al. 2020 (PMID: 33096023).",
   "rai:dataReleaseMaintenancePlan": "Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.",
   "rai:dataCollection": "Data collection processes are generally described in Clark T et al. (2024) \"Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines\" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized. ",
   "rai:dataCollectionType": [
    "Perturb-seq; IF imaging; SEC-MS; AP-MS"
   ],
   "rai:dataCollectionMissingData": "Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "rai:dataCollectionTimeframe": [
    "9/1/2022",
    "6/1/2026"
   ],
   "completeness": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "prohibitedUses": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "evi:datasetCount": 53877,
   "evi:computationCount": 1976,
   "evi:softwareCount": 6,
   "evi:schemaCount": 20,
   "evi:totalContentSizeBytes": 21051331945400,
   "evi:entitiesWithSummaryStats": 0,
   "evi:entitiesWithChecksums": 8,
   "evi:totalEntities": 55859,
   "evi:formats": [
    ".d",
    ".d directory group",
    ".tsv",
    ".xml",
    "TSV",
    "csv",
    "executable",
    "fastq.gz",
    "h5",
    "h5ad",
    "image/jpeg",
    "pdf",
    "unknown"
   ],
   "d4d:informedConsent": "Not applicable — data collected from commercially available de-identified human cell lines (KOLF2.1J iPSC and MDA-MB-468). No primary human subjects research conducted under this release.",
   "d4d:atRiskPopulations": "None — no human subjects involved; commercially sourced de-identified cell lines only."
  },
  {
   "@id": "https://orcid.org/0000-0003-4060-7360",
   "@type": "Person",
   "name": "Clark, T",
   "identifier": "https://orcid.org/0000-0003-4060-7360",
   "affiliation": {
    "@type": "Organization",
    "name": "University of Virginia"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-4535-3486",
   "@type": "Person",
   "name": "Parker, J",
   "identifier": "https://orcid.org/0000-0003-4535-3486",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California, San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-4647-3877",
   "@type": "Person",
   "name": "Al Manir, S",
   "identifier": "https://orcid.org/0000-0003-4647-3877",
   "affiliation": {
    "@type": "Organization",
    "name": "University of Virginia"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-4180-422X",
   "@type": "Person",
   "name": "Ballllosero Navarro, F",
   "identifier": "https://orcid.org/0000-0002-4180-422X",
   "affiliation": {
    "@type": "Organization",
    "name": "Stanford University"
   }
  },
  {
   "@id": "https://orcid.org/0000-0001-9998-705X",
   "@type": "Person",
   "name": "Churas, CP",
   "identifier": "https://orcid.org/0000-0001-9998-705X",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-6711-8260",
   "@type": "Person",
   "name": "Dailamy, A",
   "identifier": "https://orcid.org/0000-0002-6711-8260",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California, San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0009-0009-0483-7506",
   "@type": "Person",
   "name": "Doctor, Y",
   "identifier": "https://orcid.org/0009-0009-0483-7506",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California, San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-0223-0312",
   "@type": "Person",
   "name": "Forget, A",
   "identifier": "https://orcid.org/0000-0003-0223-0312",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Francisco"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-6311-3526",
   "@type": "Person",
   "name": "Gao, J",
   "identifier": "https://orcid.org/0000-0002-6311-3526",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-4650-9094",
   "@type": "Person",
   "name": "Hansen, JN",
   "identifier": "https://orcid.org/0000-0002-4650-9094",
   "affiliation": {
    "@type": "Organization",
    "name": "Stanford University"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-1571-8029",
   "@type": "Person",
   "name": "Hu, M",
   "identifier": "https://orcid.org/0000-0002-1571-8029",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-0917-355X",
   "@type": "Person",
   "name": "Lee, YH",
   "identifier": "https://orcid.org/0000-0003-0917-355X",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0001-7252-8638",
   "@type": "Person",
   "name": "Lenkiewicz, J",
   "identifier": "https://orcid.org/0000-0001-7252-8638",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-0384-8499",
   "@type": "Person",
   "name": "Levinson, MA",
   "identifier": "https://orcid.org/0000-0003-0384-8499",
   "affiliation": {
    "@type": "Organization",
    "name": "University of Virginia"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-3960-420X",
   "@type": "Person",
   "name": "Marquez, C",
   "identifier": "https://orcid.org/0000-0003-3960-420X",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-2404-3040",
   "@type": "Person",
   "name": "Metallo, C",
   "identifier": "https://orcid.org/0000-0003-2404-3040",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-3881-7588",
   "@type": "Person",
   "name": "Nourreddine, S",
   "identifier": "https://orcid.org/0000-0003-3881-7588",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-1103-3882",
   "@type": "Person",
   "name": "Niestroy, J",
   "identifier": "https://orcid.org/0000-0002-1103-3882",
   "affiliation": {
    "@type": "Organization",
    "name": "University of Virginia"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-4025-1299",
   "@type": "Person",
   "name": "Obernier, K",
   "identifier": "https://orcid.org/0000-0002-4025-1299",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Francisco"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-1471-9513",
   "@type": "Person",
   "name": "Pratt, D",
   "identifier": "https://orcid.org/0000-0002-1471-9513",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0009-0005-4217-2745",
   "@type": "Person",
   "name": "Qian, G",
   "identifier": "https://orcid.org/0009-0005-4217-2745",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0001-6339-9141",
   "@type": "Person",
   "name": "Schaffer, L",
   "identifier": "https://orcid.org/0000-0001-6339-9141",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-3361-3797",
   "@type": "Person",
   "name": "Sigaeva, A",
   "identifier": "https://orcid.org/0000-0003-3361-3797",
   "affiliation": {
    "@type": "Organization",
    "name": "KTH Royal Institute of Technology"
   }
  },
  {
   "@id": "https://orcid.org/0000-0001-6730-2773",
   "@type": "Person",
   "name": "Thaker, S",
   "identifier": "https://orcid.org/0000-0001-6730-2773",
   "affiliation": {
    "@type": "Organization",
    "name": "University of Alabama at Birmingham"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-8965-8153",
   "@type": "Person",
   "name": "Bélisle-Pipon, JC",
   "identifier": "https://orcid.org/0000-0002-8965-8153",
   "affiliation": {
    "@type": "Organization",
    "name": "Simon Fraser University"
   }
  },
  {
   "@id": "https://orcid.org/0000-0001-8179-1796",
   "@type": "Person",
   "name": "Brandt, C",
   "identifier": "https://orcid.org/0000-0001-8179-1796",
   "affiliation": {
    "@type": "Organization",
    "name": "Yale University"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-6112-415X",
   "@type": "Person",
   "name": "Chen, JY",
   "identifier": "https://orcid.org/0000-0002-6112-415X",
   "affiliation": {
    "@type": "Organization",
    "name": "The University of Alabama at Birmingham"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-2567-2009",
   "@type": "Person",
   "name": "Ding, Y",
   "identifier": "https://orcid.org/0000-0003-2567-2009",
   "affiliation": {
    "@type": "Organization",
    "name": "University of Texas at Austin"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-4664-3143",
   "@type": "Person",
   "name": "Fodeh, S",
   "identifier": "https://orcid.org/0000-0003-4664-3143",
   "affiliation": {
    "@type": "Organization",
    "name": "Yale University"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-4902-337X",
   "@type": "Person",
   "name": "Krogan, N",
   "identifier": "https://orcid.org/0000-0003-4902-337X",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Francisco"
   }
  },
  {
   "@id": "https://orcid.org/0000-0001-7034-0850",
   "@type": "Person",
   "name": "Lundberg, E",
   "identifier": "https://orcid.org/0000-0001-7034-0850",
   "affiliation": {
    "@type": "Organization",
    "name": "Stanford University"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-3383-1287",
   "@type": "Person",
   "name": "Mali, P",
   "identifier": "https://orcid.org/0000-0002-3383-1287",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-3508-3577",
   "@type": "Person",
   "name": "Payne-Foster, P",
   "identifier": "https://orcid.org/0000-0002-3508-3577",
   "affiliation": {
    "@type": "Organization",
    "name": "University of Alabama"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-6644-8284",
   "@type": "Person",
   "name": "Ratcliffe, S",
   "identifier": "https://orcid.org/0000-0002-6644-8284",
   "affiliation": {
    "@type": "Organization",
    "name": "University of Virginia"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-7080-8801",
   "@type": "Person",
   "name": "Ravitsky, V",
   "identifier": "https://orcid.org/0000-0002-7080-8801",
   "affiliation": {
    "@type": "Organization",
    "name": "University of Montreal"
   }
  },
  {
   "@id": "https://orcid.org/0000-0003-0435-6197",
   "@type": "Person",
   "name": "Sali, A",
   "identifier": "https://orcid.org/0000-0003-0435-6197",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-2048-4028",
   "@type": "Person",
   "name": "Schulz, W",
   "identifier": "https://orcid.org/0000-0002-2048-4028",
   "affiliation": {
    "@type": "Organization",
    "name": "Yale University"
   }
  },
  {
   "@id": "https://orcid.org/0000-0002-1708-8454",
   "@type": "Person",
   "name": "Ideker, T",
   "identifier": "https://orcid.org/0000-0002-1708-8454",
   "affiliation": {
    "@type": "Organization",
    "name": "University of California San Diego"
   }
  },
  {
   "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D001943",
   "@type": "DefinedTerm",
   "name": "Breast Neoplasms",
   "termCode": "D001943",
   "inDefinedTermSet": {
    "@id": "https://www.nlm.nih.gov/mesh/",
    "name": "MeSH"
   }
  },
  {
   "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D057026",
   "@type": "DefinedTerm",
   "name": "Induced Pluripotent Stem Cells",
   "termCode": "D057026",
   "inDefinedTermSet": {
    "@id": "https://www.nlm.nih.gov/mesh/",
    "name": "MeSH"
   }
  },
  {
   "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D064113",
   "@type": "DefinedTerm",
   "name": "CRISPR-Cas Systems",
   "termCode": "D064113",
   "inDefinedTermSet": {
    "@id": "https://www.nlm.nih.gov/mesh/",
    "name": "MeSH"
   }
  },
  {
   "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D013058",
   "@type": "DefinedTerm",
   "name": "Mass Spectrometry",
   "termCode": "D013058",
   "inDefinedTermSet": {
    "@id": "https://www.nlm.nih.gov/mesh/",
    "name": "MeSH"
   }
  },
  {
   "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D005453",
   "@type": "DefinedTerm",
   "name": "Fluorescent Antibody Technique",
   "termCode": "D005453",
   "inDefinedTermSet": {
    "@id": "https://www.nlm.nih.gov/mesh/",
    "name": "MeSH"
   }
  },
  {
   "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D017239",
   "@type": "DefinedTerm",
   "name": "Paclitaxel",
   "termCode": "D017239",
   "inDefinedTermSet": {
    "@id": "https://www.nlm.nih.gov/mesh/",
    "name": "MeSH"
   }
  },
  {
   "@id": "https://meshb.nlm.nih.gov/record/ui?ui=D000077337",
   "@type": "DefinedTerm",
   "name": "Vorinostat",
   "termCode": "D000077337",
   "inDefinedTermSet": {
    "@id": "https://www.nlm.nih.gov/mesh/",
    "name": "MeSH"
   }
  },
  {
   "@id": "http://edamontology.org/topic_0121",
   "@type": "DefinedTerm",
   "name": "Proteomics",
   "termCode": "topic_0121",
   "inDefinedTermSet": {
    "@id": "http://edamontology.org",
    "name": "EDAM"
   }
  },
  {
   "@id": "http://edamontology.org/topic_3170",
   "@type": "DefinedTerm",
   "name": "RNA-Seq",
   "termCode": "topic_3170",
   "inDefinedTermSet": {
    "@id": "http://edamontology.org",
    "name": "EDAM"
   }
  },
  {
   "@id": "http://edamontology.org/topic_3320",
   "@type": "DefinedTerm",
   "name": "Functional genomics",
   "termCode": "topic_3320",
   "inDefinedTermSet": {
    "@id": "http://edamontology.org",
    "name": "EDAM"
   }
  },
  {
   "@id": "http://edamontology.org/topic_3474",
   "@type": "DefinedTerm",
   "name": "Machine learning",
   "termCode": "topic_3474",
   "inDefinedTermSet": {
    "@id": "http://edamontology.org",
    "name": "EDAM"
   }
  },
  {
   "@id": "https://www.cellosaurus.org/CVCL_0419",
   "@type": "DefinedTerm",
   "name": "MDA-MB-468",
   "termCode": "CVCL_0419",
   "inDefinedTermSet": {
    "@id": "https://www.cellosaurus.org/",
    "name": "Cellosaurus"
   }
  },
  {
   "@id": "https://www.cellosaurus.org/CVCL_B5P3",
   "@type": "DefinedTerm",
   "name": "KOLF2.1J",
   "termCode": "CVCL_B5P3",
   "inDefinedTermSet": {
    "@id": "https://www.cellosaurus.org/",
    "name": "Cellosaurus"
   }
  },
  {
   "@id": "https://fairscape.net/api/ark:59853/rocrate-endotag-ap-ms-in-mda-mb-468-cells-paclitaxel",
   "@type": [
    "Dataset",
    "https://w3id.org/EVI#ROCrate"
   ],
   "name": "EndoTag AP-MS Profiling of Chromatin Modifier Interactome Rewiring in MDA-MB-468 Cells Upon Paclitaxel Perturbation",
   "description": "This submission comprises affinity purification mass spectrometry (AP-MS) raw data generated from endogenously tagged MDA-MB-468 cells under Paclitaxel perturbation. The dataset focuses on drug-induced interactome rewiring of chromatin modifiers relevant to the triple-negative breast cancer context. All AP-MS experiments were performed in four biological replicates. Each AP-MS batch includes an untagged parental control, 10 chromatin modifier tagged cell lines, and a positive control tagged line. DMSO-treated samples serve as vehicle control across all AP-MS experiments (Unpublished data).",
   "keywords": [
    "proteomics",
    "mass spectrometry",
    "MassIVE",
    "Endotagging",
    "AP-MS",
    "Triple Negative Breast Cancer",
    "MDA-MB-468",
    "Paclitaxel",
    "CM4AI",
    "Bridge2AI"
   ],
   "contentSize": "441.2 GB",
   "version": "1.0",
   "datePublished": "2026-05-22T13:23:35.699025+00:00",
   "isPartOf": [
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-January-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release"
    }
   ],
   "hasPart": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 560,
    "id_families": {
     "dataset": 555,
     "schema": 2,
     "sample": 1,
     "instrument": 1,
     "experiment": 1
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/sample-msv000101915-1",
     "https://fairscape.net/api/ark:59853/instrument-msv000101915-1",
     "https://fairscape.net/api/ark:59853/experiment-msv000101915-1",
     "https://fairscape.net/api/ark:59853/dataset-msv000101915-7c441f304dddb5e2",
     "https://fairscape.net/api/ark:59853/dataset-msv000101915-bcfa9680459d3d30"
    ]
   },
   "author": "Richa Tiwari",
   "associatedPublication": "",
   "license": "https://creativecommons.org/publicdomain/zero/1.0/",
   "sameAs": "https://massive.ucsd.edu/ProteoSAFe/QueryMSV?id=MSV000101915",
   "principalInvestigator": "Trey Ideker",
   "copyrightNotice": "Copyright (c) 2026 The Regents of the University of California except where otherwise noted. Spatial proteomics raw image data is copyright (c) 2026 The Board of Trustees of the Leland Stanford Junior University.",
   "conditionsOfAccess": "Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived data products should cite the Related Publications below, as well as directly citing this data collection.",
   "contactEmail": "tideker@health.ucsd.edu",
   "confidentialityLevel": "Unrestricted",
   "citation": "Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Park, S; Polacco B; Pratt D; Qian G; Schaffer, LV; Sigaeva A; Thaker S; Zhang Y; Zhao, X; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, \"Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)\", https://doi.org/10.18130/V3/K7TGEM , https://dataverse.lib.virginia.edu/, V1",
   "usageInfo": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "ethicalReview": "Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.",
   "humanSubjectResearch": "None - data collected from commercially available cell lines",
   "humanSubjectExemption": "Exempt — research with commercially available de-identified human cell lines does not constitute human subjects research.",
   "dataGovernanceCommittee": "Jilian Parker",
   "completeness": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "prohibitedUses": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "rai:dataLimitations": "This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.",
   "rai:dataBiases": "Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.",
   "rai:dataUseCases": "AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable biologically-driven, interpretable ML applications, for example as proposed in Ma et al. 2018 (PMID: 29505029) and Kuenzi et al. 2020 (PMID: 33096023).",
   "rai:dataReleaseMaintenancePlan": "Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.",
   "rai:dataCollection": "Data collection processes are generally described in Clark T et al. (2024) \"Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines\" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized. ",
   "rai:dataCollectionType": [
    "Perturb-seq; IF imaging; SEC-MS"
   ],
   "rai:dataCollectionMissingData": "Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "rai:dataCollectionTimeframe": [
    "9/1/2022",
    "1/31/2026"
   ],
   "https://w3id.org/EVI#inputs": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 117,
    "id_families": {
     "sample": 117
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/sample-msv000101915-atm-dmso",
     "https://fairscape.net/api/ark:59853/sample-msv000101915-atm-paclitaxel",
     "https://fairscape.net/api/ark:59853/sample-msv000101915-aurka-dmso",
     "https://fairscape.net/api/ark:59853/sample-msv000101915-aurka-paclitaxel",
     "https://fairscape.net/api/ark:59853/sample-msv000101915-aurkb-dmso"
    ]
   },
   "https://w3id.org/EVI#outputs": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 555,
    "id_families": {
     "dataset": 555
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/dataset-msv000101915-7c441f304dddb5e2",
     "https://fairscape.net/api/ark:59853/dataset-msv000101915-bcfa9680459d3d30",
     "https://fairscape.net/api/ark:59853/dataset-msv000101915-a508263011888846",
     "https://fairscape.net/api/ark:59853/dataset-msv000101915-1271b8748c3e4374",
     "https://fairscape.net/api/ark:59853/dataset-msv000101915-8979852725104d65"
    ]
   },
   "localEvidenceGraph": {
    "@id": "AP-MS/apms-paclitaxel-rocrate/ro-crate-prov-graph.html"
   },
   "evi:processed": true,
   "ro-crate-metadata": "AP-MS/apms-paclitaxel-rocrate/ro-crate-metadata.json"
  },
  {
   "@id": "https://fairscape.net/api/ark:59853/rocrate-endotag-ap-ms-in-mda-mb-468-cells-vorinostat",
   "@type": [
    "Dataset",
    "https://w3id.org/EVI#ROCrate"
   ],
   "name": "EndoTag-AP-MS Profiling of Chromatin Modifier Interactome Rewiring in MDA-MB-468 Cells Upon Vorinostat Perturbation",
   "description": "This submission comprises affinity purification mass spectrometry (AP-MS) raw data generated from endogenously tagged MDA-MB-468 cells under Vorinostat perturbation. The dataset focuses on drug-induced interactome rewiring of chromatin modifiers relevant to the triple-negative breast cancer context. All AP-MS experiments were performed in four biological replicates. Each AP-MS batch includes an untagged parental control, 10 chromatin modifier tagged cell lines, and a positive control tagged line. DMSO-treated sample server as vehicle controls across all AP-MS experiments (Unpublished data).",
   "keywords": [
    "proteomics",
    "mass spectrometry",
    "MassIVE",
    "EndoTagging",
    "AP-MS",
    "Triple Negative Breast Cancer",
    "MDA-MB-468",
    "Vorinostat",
    "CM4AI",
    "Bridge2AI"
   ],
   "version": "1.0",
   "datePublished": "2026-05-22T13:58:58.942348+00:00",
   "isPartOf": [
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-January-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release"
    }
   ],
   "hasPart": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 723,
    "id_families": {
     "dataset": 718,
     "schema": 2,
     "sample": 1,
     "instrument": 1,
     "experiment": 1
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/sample-msv000101917-1",
     "https://fairscape.net/api/ark:59853/instrument-msv000101917-1",
     "https://fairscape.net/api/ark:59853/experiment-msv000101917-1",
     "https://fairscape.net/api/ark:59853/dataset-msv000101917-5309353eef290f2b",
     "https://fairscape.net/api/ark:59853/dataset-msv000101917-ce5c3cd7556aee4e"
    ]
   },
   "author": "Richa Tiwari",
   "associatedPublication": "",
   "contentSize": "532.5 GB",
   "license": "https://creativecommons.org/publicdomain/zero/1.0/",
   "sameAs": "https://massive.ucsd.edu/ProteoSAFe/QueryMSV?id=MSV000101917",
   "principalInvestigator": "Trey Ideker",
   "copyrightNotice": "Copyright (c) 2026 The Regents of the University of California except where otherwise noted. Spatial proteomics raw image data is copyright (c) 2026 The Board of Trustees of the Leland Stanford Junior University.",
   "conditionsOfAccess": "Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived data products should cite the Related Publications below, as well as directly citing this data collection.",
   "contactEmail": "tideker@health.ucsd.edu",
   "confidentialityLevel": "Unrestricted",
   "citation": "Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Park, S; Polacco B; Pratt D; Qian G; Schaffer, LV; Sigaeva A; Thaker S; Zhang Y; Zhao, X; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, \"Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)\", https://doi.org/10.18130/V3/K7TGEM , https://dataverse.lib.virginia.edu/, V1",
   "usageInfo": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "ethicalReview": "Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.",
   "humanSubjectResearch": "None - data collected from commercially available cell lines",
   "humanSubjectExemption": "Exempt — research with commercially available de-identified human cell lines does not constitute human subjects research.",
   "dataGovernanceCommittee": "Jilian Parker",
   "completeness": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "prohibitedUses": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "rai:dataLimitations": "This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.",
   "rai:dataBiases": "Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.",
   "rai:dataUseCases": "AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable biologically-driven, interpretable ML applications, for example as proposed in Ma et al. 2018 (PMID: 29505029) and Kuenzi et al. 2020 (PMID: 33096023).",
   "rai:dataReleaseMaintenancePlan": "Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.",
   "rai:dataCollection": "Data collection processes are generally described in Clark T et al. (2024) \"Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines\" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized. ",
   "rai:dataCollectionType": [
    "Perturb-seq; IF imaging; SEC-MS"
   ],
   "rai:dataCollectionMissingData": "Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "rai:dataCollectionTimeframe": [
    "9/1/2022",
    "1/31/2026"
   ],
   "https://w3id.org/EVI#inputs": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 138,
    "id_families": {
     "sample": 138
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/sample-msv000101917-atm-dmso",
     "https://fairscape.net/api/ark:59853/sample-msv000101917-atm-vorinostat",
     "https://fairscape.net/api/ark:59853/sample-msv000101917-aurka-dmso",
     "https://fairscape.net/api/ark:59853/sample-msv000101917-aurka-vorinostat",
     "https://fairscape.net/api/ark:59853/sample-msv000101917-aurkb-dmso"
    ]
   },
   "https://w3id.org/EVI#outputs": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 718,
    "id_families": {
     "dataset": 718
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/dataset-msv000101917-5309353eef290f2b",
     "https://fairscape.net/api/ark:59853/dataset-msv000101917-ce5c3cd7556aee4e",
     "https://fairscape.net/api/ark:59853/dataset-msv000101917-a17ec0f3791ced91",
     "https://fairscape.net/api/ark:59853/dataset-msv000101917-5fe39ba432ce8132",
     "https://fairscape.net/api/ark:59853/dataset-msv000101917-63a240d19a595d82"
    ]
   },
   "localEvidenceGraph": {
    "@id": "AP-MS/apms-vorinostat-rocrate/ro-crate-prov-graph.html"
   },
   "evi:processed": true,
   "ro-crate-metadata": "AP-MS/apms-vorinostat-rocrate/ro-crate-metadata.json"
  },
  {
   "@type": [
    "Dataset",
    "https://w3id.org/EVI#ROCrate"
   ],
   "isPartOf": [
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-october-2025-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-January-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release"
    }
   ],
   "version": "0.1.0",
   "datePublished": "02/28/2025",
   "license": "https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en",
   "associatedPublication": "",
   "author": [
    "Hansen JN",
    "Axelsson U",
    "Johannesson A",
    "Fall J",
    "Ballllosera Navarro F",
    "Lundberg E"
   ],
   "MD5": "9422486c80bc9e1d35b2fbbc72a5f043",
   "conditionsOfAccess": "Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived products should cite the related article as well as directly citing this data collection. ",
   "copyrightNotice": "Copyright (c) 2025 by The Board of Trustees of the Leland Stanford Junior University",
   "hasPart": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 18625,
    "id_families": {
     "experiment": 468,
     "b2ai": 465,
     "stain": 3,
     "dataset": 1,
     "cell": 1,
     "treatment": 1,
     "sample": 1,
     "B2AI_5_Paclitaxel_H12_R8_z01_blue": 1,
     "B2AI_5_Paclitaxel_H12_R8_z01_red": 1,
     "B2AI_5_Paclitaxel_H12_R8_z01_yellow": 1,
     "B2AI_5_Paclitaxel_H12_R8_z01_green": 1,
     "B2AI_4_Paclitaxel_H12_R8_z02_blue": 1
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/dataset-paclitaxel-manifest",
     "https://fairscape.net/api/ark:59853/cell-line-MDA-MB-468",
     "https://fairscape.net/api/ark:59853/treatment-paclitaxel",
     "https://fairscape.net/api/ark:59853/stain-dapi",
     "https://fairscape.net/api/ark:59853/stain-calreticulin-antibody"
    ]
   },
   "contactEmail": "emmalu@stanford.edu",
   "funder": "National Institutes of Health: 1OT2OD032742-01",
   "@id": "https://fairscape.net/api/ark:59853/rocrate-paclitaxel-if-data-release",
   "name": "Paclitaxel IF Images",
   "description": "This data set displays the spatial localization of 464 proteins of interest in cells of the breast cancer cell line MDA-MB-468 treated with paclitaxel as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).",
   "keywords": [
    "AI",
    "artificial intelligence",
    "breast cancer",
    "cell maps",
    "CM4AI",
    "machine learning",
    "MDA-MB-468",
    "paclitaxel",
    "protein localization"
   ],
   "contentSize": "2.6 GB",
   "https://w3id.org/EVI#inputs": [
    {
     "@id": "https://fairscape.net/api/ark:59853/sample-mda-mb-468-lNzmhj2iYbi"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-paclitaxel-manifest"
    }
   ],
   "https://w3id.org/EVI#outputs": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 17685,
    "id_families": {
     "dataset": 1,
     "B2AI_5_Paclitaxel_H12_R8_z01_blue": 1,
     "B2AI_5_Paclitaxel_H12_R8_z01_red": 1,
     "B2AI_5_Paclitaxel_H12_R8_z01_yellow": 1,
     "B2AI_5_Paclitaxel_H12_R8_z01_green": 1,
     "B2AI_4_Paclitaxel_H12_R8_z02_blue": 1,
     "B2AI_4_Paclitaxel_H12_R8_z02_red": 1,
     "B2AI_4_Paclitaxel_H12_R8_z02_yellow": 1,
     "B2AI_4_Paclitaxel_H12_R8_z02_green": 1,
     "B2AI_3_Paclitaxel_H12_R8_z02_blue": 1,
     "B2AI_3_Paclitaxel_H12_R8_z02_red": 1,
     "B2AI_3_Paclitaxel_H12_R8_z02_yellow": 1
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/dataset-paclitaxel-manifest",
     "https://fairscape.net/api/ark:59853/B2AI_5_Paclitaxel_H12_R8_z01_blue",
     "https://fairscape.net/api/ark:59853/B2AI_5_Paclitaxel_H12_R8_z01_red",
     "https://fairscape.net/api/ark:59853/B2AI_5_Paclitaxel_H12_R8_z01_yellow",
     "https://fairscape.net/api/ark:59853/B2AI_5_Paclitaxel_H12_R8_z01_green"
    ]
   },
   "principalInvestigator": "Trey Ideker",
   "confidentialityLevel": "Unrestricted",
   "citation": "Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Marquez C; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Polacco B; Pratt D; Qian G; Schaffer L; Sigaeva A; Thaker S; Zhang Y; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, \"Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)\", https://doi.org/10.18130/V3/B35XWX , https://dataverse.lib.virginia.edu/, V1",
   "usageInfo": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "additionalProperty": [
    {
     "@type": "PropertyValue",
     "name": "Completeness",
     "value": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release."
    },
    {
     "@type": "PropertyValue",
     "name": "Human Subject",
     "value": "No"
    },
    {
     "@type": "PropertyValue",
     "name": "Prohibited Uses",
     "value": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval."
    }
   ],
   "rai:dataLimitations": "This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.",
   "rai:dataBiases": "Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.",
   "rai:dataUseCases": "AI-ready datasets to support research in functional genomics, AI model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable visible machine learning applications, as proposed in Ma et al. (2018) Nature Methods.",
   "rai:dataReleaseMaintenancePlan": "Dataset will be regularly updated and augmented through the end of the project in November 2026, on a quarterly basis. Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.",
   "localEvidenceGraph": {
    "@id": "Images/paclitaxel/ro-crate-prov-graph.html"
   },
   "rai:dataCollection": "Data collection processes are generally described in Clark T et al. (2024) \"Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines\" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized. ",
   "rai:dataCollectionType": [
    "Perturb-seq; IF imaging; SEC-MS"
   ],
   "rai:dataCollectionTimeframe": [
    "9/1/2022",
    "10/13/25"
   ],
   "ethicalReview": "Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.",
   "humanSubjectResearch": "None - data collected from commercially available cell lines",
   "humanSubjectExemption": "Exempt — research with commercially available de-identified human cell lines does not constitute human subjects research.",
   "dataGovernanceCommittee": "Jilian Parker",
   "completeness": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "prohibitedUses": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "rai:dataCollectionMissingData": "Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "evi:processed": true,
   "ro-crate-metadata": "Images/paclitaxel/ro-crate-metadata.json"
  },
  {
   "@type": [
    "Dataset",
    "https://w3id.org/EVI#ROCrate"
   ],
   "isPartOf": [
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-october-2025-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-January-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release"
    }
   ],
   "version": "0.1.0",
   "MD5": "0b4d129f5fbc3bb7f7ea564cd032cef7",
   "datePublished": "02/28/2025",
   "license": "https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en",
   "associatedPublication": "",
   "author": [
    "Hansen JN",
    "Axelsson U",
    "Johannesson A",
    "Fall J",
    "Ballllosera Navarro F",
    "Lundberg E"
   ],
   "conditionsOfAccess": "Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived data products should cite the Related Publications below, as well as directly citing this data collection.",
   "copyrightNotice": "Copyright (c) 2025 by The Board of Trustees of the Leland Stanford Junior University",
   "hasPart": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 16845,
    "id_families": {
     "experiment": 468,
     "b2ai": 465,
     "stain": 3,
     "dataset": 1,
     "cell": 1,
     "treatment": 1,
     "sample": 1,
     "B2AI_5_untreated_H12_R8_z01_blue": 1,
     "B2AI_5_untreated_H12_R8_z01_red": 1,
     "B2AI_5_untreated_H12_R8_z01_yellow": 1,
     "B2AI_5_untreated_H12_R8_z01_green": 1,
     "B2AI_4_untreated_H12_R8_z02_blue": 1
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/dataset-untreated-manifest",
     "https://fairscape.net/api/ark:59853/cell-line-MDA-MB-468",
     "https://fairscape.net/api/ark:59853/treatment-control",
     "https://fairscape.net/api/ark:59853/stain-dapi",
     "https://fairscape.net/api/ark:59853/stain-calreticulin-antibody"
    ]
   },
   "@id": "https://fairscape.net/api/ark:59853/rocrate-untreated-if-data-release",
   "name": "Untreated IF Images",
   "description": "This data set displays the spatial localization of 464 proteins of interest in untreated cells of the breast cancer cell line MDA-MB-468 as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).",
   "keywords": [
    "artificial intelligence",
    "breast cancer",
    "cell maps",
    "CM4AI",
    "machine learning",
    "MDA-MB-468",
    "protein localization"
   ],
   "contentSize": "3.2 GB",
   "https://w3id.org/EVI#inputs": [
    {
     "@id": "https://fairscape.net/api/ark:59853/sample-mda-mb-468-UK59eH6L75"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-untreated-manifest"
    }
   ],
   "https://w3id.org/EVI#outputs": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 15905,
    "id_families": {
     "dataset": 1,
     "B2AI_5_untreated_H12_R8_z01_blue": 1,
     "B2AI_5_untreated_H12_R8_z01_red": 1,
     "B2AI_5_untreated_H12_R8_z01_yellow": 1,
     "B2AI_5_untreated_H12_R8_z01_green": 1,
     "B2AI_4_untreated_H12_R8_z02_blue": 1,
     "B2AI_4_untreated_H12_R8_z02_red": 1,
     "B2AI_4_untreated_H12_R8_z02_yellow": 1,
     "B2AI_4_untreated_H12_R8_z02_green": 1,
     "B2AI_3_untreated_H12_R8_z01_blue": 1,
     "B2AI_3_untreated_H12_R8_z01_red": 1,
     "B2AI_3_untreated_H12_R8_z01_yellow": 1
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/dataset-untreated-manifest",
     "https://fairscape.net/api/ark:59853/B2AI_5_untreated_H12_R8_z01_blue",
     "https://fairscape.net/api/ark:59853/B2AI_5_untreated_H12_R8_z01_red",
     "https://fairscape.net/api/ark:59853/B2AI_5_untreated_H12_R8_z01_yellow",
     "https://fairscape.net/api/ark:59853/B2AI_5_untreated_H12_R8_z01_green"
    ]
   },
   "principalInvestigator": "Trey Ideker",
   "contactEmail": "tideker@health.ucsd.edu",
   "confidentialityLevel": "Unrestricted",
   "citation": "Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Marquez C; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Polacco B; Pratt D; Qian G; Schaffer L; Sigaeva A; Thaker S; Zhang Y; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, \"Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)\", https://doi.org/10.18130/V3/B35XWX , https://dataverse.lib.virginia.edu/, V1",
   "usageInfo": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "additionalProperty": [
    {
     "@type": "PropertyValue",
     "name": "Completeness",
     "value": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release."
    },
    {
     "@type": "PropertyValue",
     "name": "Human Subject",
     "value": "No"
    },
    {
     "@type": "PropertyValue",
     "name": "Prohibited Uses",
     "value": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval."
    }
   ],
   "rai:dataLimitations": "This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.",
   "rai:dataBiases": "Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.",
   "rai:dataUseCases": "AI-ready datasets to support research in functional genomics, AI model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable visible machine learning applications, as proposed in Ma et al. (2018) Nature Methods.",
   "rai:dataReleaseMaintenancePlan": "Dataset will be regularly updated and augmented through the end of the project in November 2026, on a quarterly basis. Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.",
   "localEvidenceGraph": {
    "@id": "Images/untreated/ro-crate-prov-graph.html"
   },
   "rai:dataCollection": "Data collection processes are generally described in Clark T et al. (2024) \"Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines\" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized. ",
   "rai:dataCollectionType": [
    "Perturb-seq; IF imaging; SEC-MS"
   ],
   "rai:dataCollectionTimeframe": [
    "9/1/2022",
    "10/13/25"
   ],
   "ethicalReview": "Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.",
   "humanSubjectResearch": "None - data collected from commercially available cell lines",
   "humanSubjectExemption": "Exempt — research with commercially available de-identified human cell lines does not constitute human subjects research.",
   "dataGovernanceCommittee": "Jilian Parker",
   "completeness": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "prohibitedUses": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "rai:dataCollectionMissingData": "Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "evi:processed": true,
   "ro-crate-metadata": "Images/untreated/ro-crate-metadata.json"
  },
  {
   "@type": [
    "Dataset",
    "https://w3id.org/EVI#ROCrate"
   ],
   "isPartOf": [
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-october-2025-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-January-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release"
    }
   ],
   "version": "0.1.0",
   "datePublished": "02/28/2025",
   "MD5": "ac577109a41a9806978461157b777d52",
   "license": "https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en",
   "associatedPublication": "",
   "author": [
    "Hansen JN",
    "Axelsson U",
    "Johannesson A",
    "Fall J",
    "Ballllosera Navarro F",
    "Lundberg E"
   ],
   "conditionsOfAccess": "Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived products should cite the related article as well as directly citing this data collection. ",
   "copyrightNotice": "Copyright (c) 2025 by The Board of Trustees of the Leland Stanford Junior University",
   "hasPart": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 18789,
    "id_families": {
     "experiment": 468,
     "b2ai": 465,
     "stain": 3,
     "dataset": 1,
     "cell": 1,
     "treatment": 1,
     "sample": 1,
     "B2AI_5_Vorinostat_H12_R8_z01_blue": 1,
     "B2AI_5_Vorinostat_H12_R8_z01_red": 1,
     "B2AI_5_Vorinostat_H12_R8_z01_yellow": 1,
     "B2AI_5_Vorinostat_H12_R8_z01_green": 1,
     "B2AI_4_Vorinostat_H12_R8_z02_blue": 1
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/dataset-vorinostat-manifest",
     "https://fairscape.net/api/ark:59853/cell-line-MDA-MB-468",
     "https://fairscape.net/api/ark:59853/treatment-vorinostat",
     "https://fairscape.net/api/ark:59853/stain-dapi",
     "https://fairscape.net/api/ark:59853/stain-calreticulin-antibody"
    ]
   },
   "@id": "https://fairscape.net/api/ark:59853/rocrate-vorinostat-if-data-release",
   "name": "Vorinostat IF Images",
   "description": "This data set displays the spatial localization of 464 proteins of interest in cells of the breast cancer cell line MDA-MB-468 treated with vorinostat as imaged by immunofluorescence-based staining (ICC-IF) and confocal microscopy in the Lundberg Lab at Stanford University, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) project. Nuclei were stained with DAPI (blue channel); endoplasmic reticulum with a calreticulin antibody (yellow channel); microtubules with tubulin antibody (red channel); and antibody against protein of interest (green channel).",
   "keywords": [
    "AI",
    "artificial intelligence",
    "breast cancer",
    "cell maps",
    "CM4AI",
    "machine learning",
    "MDA-MB-468",
    "protein localization",
    "vorinostat"
   ],
   "contentSize": "2.8 GB",
   "contactEmail": "emmalu@stanford.edu",
   "funder": "National Institutes of Health: 1OT2OD032742-01",
   "https://w3id.org/EVI#inputs": [
    {
     "@id": "https://fairscape.net/api/ark:59853/sample-mda-mb-468-7Ja0Ix9dXRq"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-vorinostat-manifest"
    }
   ],
   "https://w3id.org/EVI#outputs": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 17849,
    "id_families": {
     "dataset": 1,
     "B2AI_5_Vorinostat_H12_R8_z01_blue": 1,
     "B2AI_5_Vorinostat_H12_R8_z01_red": 1,
     "B2AI_5_Vorinostat_H12_R8_z01_yellow": 1,
     "B2AI_5_Vorinostat_H12_R8_z01_green": 1,
     "B2AI_4_Vorinostat_H12_R8_z02_blue": 1,
     "B2AI_4_Vorinostat_H12_R8_z02_red": 1,
     "B2AI_4_Vorinostat_H12_R8_z02_yellow": 1,
     "B2AI_4_Vorinostat_H12_R8_z02_green": 1,
     "B2AI_3_Vorinostat_H12_R8_z02_blue": 1,
     "B2AI_3_Vorinostat_H12_R8_z02_red": 1,
     "B2AI_3_Vorinostat_H12_R8_z02_yellow": 1
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/dataset-vorinostat-manifest",
     "https://fairscape.net/api/ark:59853/B2AI_5_Vorinostat_H12_R8_z01_blue",
     "https://fairscape.net/api/ark:59853/B2AI_5_Vorinostat_H12_R8_z01_red",
     "https://fairscape.net/api/ark:59853/B2AI_5_Vorinostat_H12_R8_z01_yellow",
     "https://fairscape.net/api/ark:59853/B2AI_5_Vorinostat_H12_R8_z01_green"
    ]
   },
   "principalInvestigator": "Trey Ideker",
   "confidentialityLevel": "Unrestricted",
   "citation": "Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Marquez C; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Polacco B; Pratt D; Qian G; Schaffer L; Sigaeva A; Thaker S; Zhang Y; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, \"Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)\", https://doi.org/10.18130/V3/B35XWX , https://dataverse.lib.virginia.edu/, V1",
   "usageInfo": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "additionalProperty": [
    {
     "@type": "PropertyValue",
     "name": "Completeness",
     "value": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release."
    },
    {
     "@type": "PropertyValue",
     "name": "Human Subject",
     "value": "No"
    },
    {
     "@type": "PropertyValue",
     "name": "Prohibited Uses",
     "value": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval."
    }
   ],
   "rai:dataLimitations": "This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.",
   "rai:dataBiases": "Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.",
   "rai:dataUseCases": "AI-ready datasets to support research in functional genomics, AI model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable visible machine learning applications, as proposed in Ma et al. (2018) Nature Methods.",
   "rai:dataReleaseMaintenancePlan": "Dataset will be regularly updated and augmented through the end of the project in November 2026, on a quarterly basis. Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.",
   "localEvidenceGraph": {
    "@id": "Images/vorinostat/ro-crate-prov-graph.html"
   },
   "rai:dataCollection": "Data collection processes are generally described in Clark T et al. (2024) \"Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines\" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized. ",
   "rai:dataCollectionType": [
    "Perturb-seq; IF imaging; SEC-MS"
   ],
   "rai:dataCollectionTimeframe": [
    "9/1/2022",
    "10/13/25"
   ],
   "ethicalReview": "Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.",
   "humanSubjectResearch": "None - data collected from commercially available cell lines",
   "humanSubjectExemption": "Exempt — research with commercially available de-identified human cell lines does not constitute human subjects research.",
   "dataGovernanceCommittee": "Jilian Parker",
   "completeness": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "prohibitedUses": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "rai:dataCollectionMissingData": "Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "evi:processed": true,
   "ro-crate-metadata": "Images/vorinostat/ro-crate-metadata.json"
  },
  {
   "@id": "https://fairscape.net/api/ark:59853/rocrate-sec-ms-characterization-of-kolf2-neuronal-and-cardiomyocyte-differentiation-cm4ai-yjftt2oec6c",
   "@type": [
    "Dataset",
    "https://w3id.org/EVI#ROCrate"
   ],
   "name": "SEC-MS characterization of KOLF2 neuronal and cardiomyocyte differentiation",
   "description": "This dataset contains mass spectrometry-based proteomics data investigating the protein-protein interaction networks during the differentiation of KOLF2.1J human induced pluripotent stem cells (iPSCs). We employed Size Exclusion Chromatography coupled with Mass Spectrometry (SEC-MS) to map multiscale biological organization across three distinct stages:  Parental: Undifferentiated iPSCs.  NPC: Neural Progenitor Cells.  Neurons: Terminally differentiated neurons.  Cardio: Terminally differentiated Cardiomyocyte.  The study aims to identify how protein complexes and networks are reorganized during neuronal lineage commitment and cardiomycyte differentiation. Data was acquired on a Bruker timsTOF platform, and quantitative analysis was performed using Spectronaut.  Experimental Design:  Sample Groups: Parental (Reps 1-4), NPC (Reps 1-2), Neuron (Reps 1-3), and Cardio (Reps 1-2).  Modality: SEC-MS for protein correlation profiling.  Data Processing: Raw data was processed using Spectronaut; downstream analysis was conducted in R using MSstats-compatible formats.",
   "keywords": [
    "proteomics",
    "mass spectrometry",
    "MassIVE",
    "SEC-MS",
    "iPSC",
    "KOLF2.1J",
    "KOLF2",
    "IPSC",
    "Neuron",
    "Cardiomyocyte"
   ],
   "version": "1.0",
   "datePublished": "2026-05-27T20:36:55.634084+00:00",
   "isPartOf": [
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release"
    }
   ],
   "contentSize": "1.11 TB",
   "hasPart": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 867,
    "id_families": {
     "dataset": 862,
     "schema": 2,
     "sample": 1,
     "instrument": 1,
     "experiment": 1
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/sample-msv000100676-1",
     "https://fairscape.net/api/ark:59853/instrument-msv000100676-1",
     "https://fairscape.net/api/ark:59853/experiment-msv000100676-1",
     "https://fairscape.net/api/ark:59853/dataset-msv000100676-cc8e39efd0c8a96b",
     "https://fairscape.net/api/ark:59853/dataset-msv000100676-1ed62cf423da0185"
    ]
   },
   "author": "Antoine Forget",
   "associatedPublication": "",
   "license": "https://creativecommons.org/publicdomain/zero/1.0/",
   "sameAs": "https://massive.ucsd.edu/ProteoSAFe/QueryMSV?id=MSV000100676",
   "localEvidenceGraph": {
    "@id": "mass-spec/iPSCs/ro-crate-prov-graph.html"
   },
   "principalInvestigator": "Trey Ideker",
   "copyrightNotice": "Copyright (c) 2026 The Regents of the University of California except where otherwise noted. Spatial proteomics raw image data is copyright (c) 2026 The Board of Trustees of the Leland Stanford Junior University.",
   "conditionsOfAccess": "Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived data products should cite the Related Publications below, as well as directly citing this data collection.",
   "contactEmail": "tideker@health.ucsd.edu",
   "confidentialityLevel": "Unrestricted",
   "citation": "Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Park, S; Polacco B; Pratt D; Qian G; Schaffer, LV; Sigaeva A; Thaker S; Zhang Y; Zhao, X; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, \"Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)\", https://doi.org/10.18130/V3/K7TGEM , https://dataverse.lib.virginia.edu/, V1",
   "usageInfo": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "ethicalReview": "Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.",
   "humanSubjectResearch": "None - data collected from commercially available cell lines",
   "humanSubjectExemption": "Exempt — research with commercially available de-identified human cell lines does not constitute human subjects research.",
   "dataGovernanceCommittee": "Jilian Parker",
   "completeness": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "prohibitedUses": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "rai:dataLimitations": "This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.",
   "rai:dataBiases": "Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.",
   "rai:dataUseCases": "AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable biologically-driven, interpretable ML applications, for example as proposed in Ma et al. 2018 (PMID: 29505029) and Kuenzi et al. 2020 (PMID: 33096023).",
   "rai:dataReleaseMaintenancePlan": "Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.",
   "rai:dataCollection": "Data collection processes are generally described in Clark T et al. (2024) \"Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines\" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized. ",
   "rai:dataCollectionType": [
    "Perturb-seq; IF imaging; SEC-MS; AP-MS"
   ],
   "rai:dataCollectionMissingData": "Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "rai:dataCollectionTimeframe": [
    "9/1/2022",
    "6/1/2026"
   ],
   "publisher": "MassIVE",
   "https://w3id.org/EVI#inputs": [
    {
     "@id": "https://fairscape.net/api/ark:59853/sample-msv000100676-kolf2-1j"
    }
   ],
   "https://w3id.org/EVI#outputs": [
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-msv000100676-cc8e39efd0c8a96b"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-msv000100676-1ed62cf423da0185"
    }
   ],
   "evi:processed": true,
   "ro-crate-metadata": "mass-spec/iPSCs/ro-crate-metadata.json"
  },
  {
   "@id": "https://fairscape.net/api/ark:59853/rocrate-data-from-treated-human-cancer-cells-jan-26",
   "@type": [
    "Dataset",
    "https://w3id.org/EVI#ROCrate"
   ],
   "name": "SEC-MS of MDA-MB468 following treatment of vorinostat or paclitaxel.",
   "description": "This dataset was generated by size exclusion chromatography-mass spectroscopy (SEC-MS) following the treatment of vorinostat or paclitaxel on MDA-MB468 human breast cancer cells, in the Nevan Krogan laboratory at the University of California San Francisco, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) Functional Genomics Grand Challenge, a component of the U.S. National Institute of Health's (NIH) Bridge2AI program.",
   "keywords": [
    "AI",
    "Artificial intelligence",
    "Breast cancer",
    "Cell maps",
    "CM4AI",
    "IPSC",
    "KOLF2.1J",
    "Machine learning",
    "Mass spectroscopy",
    "Protein-protein interaction",
    "SEC-MS",
    "vorinostat",
    "paclitaxel"
   ],
   "identifier": "https://doi.org/doi:10.25345/C5348GV4S",
   "isPartOf": [
    {
     "@id": "https://fairscape.net/api/ark:59853/organization-university-of-california-san-diego"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/project-cm4ai"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-January-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release"
    }
   ],
   "version": "1.5",
   "hasPart": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 45,
    "id_families": {
     "dataset": 18,
     "experiment": 9,
     "computation": 9,
     "schema": 9
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/experiment-control-1-sec-ms-mda-mb468",
     "https://fairscape.net/api/ark:59853/experiment-control-2-sec-ms-mda-mb468",
     "https://fairscape.net/api/ark:59853/experiment-control-4-sec-ms-mda-mb468",
     "https://fairscape.net/api/ark:59853/experiment-paclitaxel-1-sec-ms-mda-mb468",
     "https://fairscape.net/api/ark:59853/experiment-paclitaxel-2-sec-ms-mda-mb468"
    ]
   },
   "author": "Forget A, Obernier K, Krogan N",
   "license": "https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en",
   "associatedPublication": "http://doi.org/10.1101/2024.05.21.589311",
   "conditionsOfAccess": "Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived products should cite the related article as well as directly citing this data collection.",
   "copyrightNotice": "Copyright (c) 2026 by The Regents of the University of California",
   "contactEmail": "nevan.krogan@ucsf.edu",
   "funder": "National Institutes of Health: 1OT2OD032742-01",
   "MD5": "cb67e7749b15ce87b9042a9feba9d032",
   "contentUrl": "ftp://massive-ftp.ucsd.edu/v10/MSV000098237/",
   "url": "https://massive.ucsd.edu/ProteoSAFe/dataset.jsp?task=ad8b8084f5b14af5bafac70fdd42a577",
   "contentSize": "910 GB",
   "publisher": "MassIVE",
   "https://w3id.org/EVI#inputs": [
    {
     "@id": "https://fairscape.net/api/ark:59853/sample-MDA-MB-468-cell-line-mass-spec-sample-oPzB4vOldoF"
    }
   ],
   "https://w3id.org/EVI#outputs": [
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-control-1-report"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-control-2-report"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-control-4-report"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-paclitaxel-1-report"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-paclitaxel-2-report"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-paclitaxel-3-report"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-vorinostat-1-report"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-vorinostat-2-report"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-vorinostat-3-report"
    }
   ],
   "principalInvestigator": "Trey Ideker",
   "confidentialityLevel": "Unrestricted",
   "citation": "Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Marquez C; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Polacco B; Pratt D; Qian G; Schaffer L; Sigaeva A; Thaker S; Zhang Y; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, \"Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)\", https://doi.org/10.18130/V3/B35XWX , https://dataverse.lib.virginia.edu/, V1",
   "usageInfo": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "additionalProperty": [
    {
     "@type": "PropertyValue",
     "name": "Completeness",
     "value": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release."
    },
    {
     "@type": "PropertyValue",
     "name": "Human Subject",
     "value": "None - data collected from commercially available cell lines"
    },
    {
     "@type": "PropertyValue",
     "name": "Prohibited Uses",
     "value": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval."
    }
   ],
   "rai:dataLimitations": "This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.",
   "rai:dataBiases": "Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.",
   "rai:dataUseCases": "AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable biologically-driven, interpretable ML applications, for example as proposed in Ma et al. 2018 (PMID: 29505029) and Kuenzi et al. 2020 (PMID: 33096023).",
   "rai:dataReleaseMaintenancePlan": "Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.",
   "rai:dataCollection": "Data collection processes are generally described in Clark T et al. (2024) \"Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines\" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized. ",
   "rai:dataCollectionType": [
    "Perturb-seq; IF imaging; SEC-MS"
   ],
   "rai:dataCollectionTimeframe": [
    "9/1/2022",
    "1/31/2026"
   ],
   "ethicalReview": "Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.",
   "humanSubjectResearch": "None - data collected from commercially available cell lines",
   "humanSubjectExemption": "Exempt — research with commercially available de-identified human cell lines does not constitute human subjects research.",
   "dataGovernanceCommittee": "Jilian Parker",
   "completeness": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "prohibitedUses": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "rai:dataCollectionMissingData": "Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "localEvidenceGraph": {
    "@id": "mass-spec/cancer-cells/ro-crate-prov-graph.html"
   },
   "evi:processed": true,
   "ro-crate-metadata": "mass-spec/cancer-cells/ro-crate-metadata.json"
  },
  {
   "@id": "https://fairscape.net/api/ark:59853/rocrate-sra-data-for-perturbation-cell-atlas-58RmCGMQBj",
   "@type": [
    "Dataset",
    "https://w3id.org/EVI#ROCrate"
   ],
   "name": "Perturbation Cell Atlas of Human Induced Pluripotent Stem Cells - Raw Sequence Data",
   "description": "This dataset represents raw sequence data from an expressed genome-scale CRISPRi Perturbation Cell Atlas in KOLF2.1J human induced pluripotent stem cells (hiPSCs) mapping transcriptional and fitness phenotypes associated with 11,739 targeted genes, as part of the Cell Maps for Artificial Intelligence (CM4AI; CM4AI.org) Functional Genomics Grand Challenge, a component of the U.S. National Institute of Health's (NIH) Bridge2AI program.",
   "keywords": [
    "AI",
    "Artificial intelligence",
    "Cell maps",
    "CM4AI",
    "CRISPR perturbation",
    "IPSC",
    "KOLF2.1J",
    "Machine learning",
    "Perturb-seq",
    "scRNAseq",
    "single-cell RNA sequencing"
   ],
   "isPartOf": [
    {
     "@id": "https://fairscape.net/api/ark:59853/organization-university-of-california-san-diego"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/project-cm4ai"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-october-2025-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release"
    }
   ],
   "version": "0.6",
   "license": "https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en",
   "associatedPublication": "Nourreddine S, Doctor Y, Dailamy A, et al. (2024) \"A Perturbation Cell Atlas of Human Induced Pluripotent Stem Cells\" bioRxiv 2024.11.03.621734; doi: https://doi.org/10.1101/2024.11.03.621734",
   "author": "Doctor Y; Dailamy A; Forget A; Lee YH; Chinn B; Khaliq H; Polacco B; Muralidharan, M; Pan E; Zhang Y; Sigaeva A; Hansen JN; Gao J; Parker JA; Obernier K; Clark T; Chen JY; Metallo C; Lundberg E; Idkeker T; Krogan N; Mali P",
   "conditionsOfAccess": "Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived products should cite the related article as well as directly citing this data collection.",
   "copyrightNotice": "Copyright (c) 2024 by The Regents of the University of California",
   "hasPart": [],
   "contactEmail": "pmali@ucsd.edu",
   "funder": "National Institutes of Health: 1OT2OD032742-01, R01HG012351, R01NS131560, U54CA274502, #S10 OD026929. Department of Defense: W81XWH-22-1-0401. CIRM training: EDUC4-12804. Dutch Research Council: NWO, 019.231EN.013. National Cancer Institute: P30CA023100",
   "MD5": "cbdb263b1c099396d75e16f00a79a818",
   "url": "Embargoed",
   "contentSize": "16.7TB",
   "https://w3id.org/EVI#inputs": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 192,
    "id_families": {
     "sample": 192
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/sample-sra-sample-53424698-GkKwucuUUW8",
     "https://fairscape.net/api/ark:59853/sample-sra-sample-53424707-Ks8UBhB5Y3c",
     "https://fairscape.net/api/ark:59853/sample-sra-sample-53424708-B9F2ilij312",
     "https://fairscape.net/api/ark:59853/sample-sra-sample-53424709-Oyr1MZMJGM",
     "https://fairscape.net/api/ark:59853/sample-sra-sample-53424710-Oyr1MZMJcNy"
    ]
   },
   "https://w3id.org/EVI#outputs": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 103,
    "id_families": {
     "dataset": 103
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/dataset-sra-data-kolf2-alpha-crispr-chip-2-channel-14-n0TB2HDqaSh",
     "https://fairscape.net/api/ark:59853/dataset-sra-data-kolf2-alpha-crispr-chip-2-channel-3-Ewf13NZhZlZ",
     "https://fairscape.net/api/ark:59853/dataset-sra-data-kolf2-alpha-gex-chip-2-channel-14-QRexSUcAae",
     "https://fairscape.net/api/ark:59853/dataset-sra-data-kolf2-alpha-gex-chip-2-channel-3-6XbNlQwgbw7",
     "https://fairscape.net/api/ark:59853/dataset-sra-data-kolf2-beta-crispr-chip-1-channel-1-uhBWMiS21VX"
    ]
   },
   "principalInvestigator": "Trey Ideker",
   "confidentialityLevel": "Unrestricted",
   "citation": "Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Marquez C; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Polacco B; Pratt D; Qian G; Schaffer L; Sigaeva A; Thaker S; Zhang Y; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, \"Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)\", https://doi.org/10.18130/V3/B35XWX , https://dataverse.lib.virginia.edu/, V1",
   "usageInfo": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "additionalProperty": [
    {
     "@type": "PropertyValue",
     "name": "Completeness",
     "value": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release."
    },
    {
     "@type": "PropertyValue",
     "name": "Human Subject",
     "value": "None - data collected from commercially available cell lines"
    },
    {
     "@type": "PropertyValue",
     "name": "Prohibited Uses",
     "value": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval."
    }
   ],
   "rai:dataLimitations": "This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.",
   "rai:dataBiases": "Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.",
   "rai:dataUseCases": "AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable biologically-driven, interpretable ML applications, for example as proposed in Ma et al. 2018 (PMID: 29505029) and Kuenzi et al. 2020 (PMID: 33096023).",
   "rai:dataReleaseMaintenancePlan": "Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.",
   "rai:dataCollection": "Data collection processes are generally described in Clark T et al. (2024) \"Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines\" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized. ",
   "rai:dataCollectionType": [
    "Perturb-seq; IF imaging; SEC-MS"
   ],
   "rai:dataCollectionTimeframe": [
    "9/1/2022",
    "10/13/25"
   ],
   "ethicalReview": "Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.",
   "humanSubjectResearch": "None - data collected from commercially available cell lines",
   "humanSubjectExemption": "Exempt — research with commercially available de-identified human cell lines does not constitute human subjects research.",
   "dataGovernanceCommittee": "Jilian Parker",
   "completeness": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "prohibitedUses": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "rai:dataCollectionMissingData": "Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "evi:processed": true,
   "localEvidenceGraph": {
    "@id": "Perturb-Seq/sra/ro-crate-prov-graph.html"
   },
   "ro-crate-metadata": "Perturb-Seq/sra/ro-crate-metadata.json"
  },
  {
   "@id": "https://fairscape.net/api/ark:59853/rocrate-a-perturbation-cell-atlas-of-human-induced-pluripotent-stem-cells",
   "@type": [
    "Dataset",
    "https://w3id.org/EVI#ROCrate"
   ],
   "name": "Perturbation Cell Atlas of Human Induced Pluripotent Stem Cells - Perturb Seq",
   "description": "This dataset represents an expressed genome-scale CRISPRi Perturbation Cell Atlas in KOLF2.1J human induced pluripotent stem cells (hiPSCs) mapping transcriptional and fitness phenotypes associated with 11,739 targeted genes, as part of the Cell Maps for Artificial Intelligence (CM4AI) Functional Genomics Grand Challenge, a component of the U.S. National Institute of Health's (NIH) Bridge2AI program. We validated these findings via phenotypic, protein-interaction, and metabolic tracing assays.",
   "publisher": "FigShare",
   "keywords": [
    "AI",
    "Artificial intelligence",
    "Cell maps",
    "CM4AI",
    "CRISPR perturbation",
    "IPSC",
    "KOLF2.1J",
    "Machine learning",
    "Perturb-seq",
    "scRNAseq",
    "single-cell RNA sequencing"
   ],
   "isPartOf": [
    {
     "@id": "https://fairscape.net/api/ark:59853/organization-university-of-california-san-diego"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/project-cm4ai"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-January-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release,"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/rocrate-cell-maps-for-artificial-intelligence-June-2026-data-release"
    }
   ],
   "version": "1.0",
   "license": "https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en",
   "associatedPublication": "Nourreddine S, Doctor Y, Dailamy A, et al. (2024) \"A Perturbation Cell Atlas of Human Induced Pluripotent Stem Cells\" bioRxiv 2024.11.03.621734; doi: https://doi.org/10.1101/2024.11.03.621734",
   "author": "Doctor Y; Dailamy A; Forget A; Lee YH; Chinn B; Khaliq H; Polacco B; Muralidharan, M; Pan E; Zhang Y; Sigaeva A; Hansen JN; Gao J; Parker JA; Obernier K; Clark T; Chen JY; Metallo C; Lundberg E; Ideker T; Krogan N; Mali P",
   "conditionsOfAccess": "Attribution is required to the copyright holders and the authors. Any publications referencing this data or derived products should cite the related article as well as directly citing this data collection.",
   "copyrightNotice": "Copyright (c) 2026 by The Regents of the University of California",
   "hasPart": [
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-kolf-pan-genome-aggregated-data"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/schema-kolf-pan-genome-aggregate-20250529-143000"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/computation-perturbation-cell-atlas-data-processing-Cmx85JtSguG"
    }
   ],
   "contentSize": "177.35 GB",
   "contactEmail": "pmali@ucsd.edu",
   "funder": "National Institutes of Health: 1OT2OD032742-01, R01HG012351, R01NS131560, U54CA274502, #S10 OD026929. Department of Defense: W81XWH-22-1-0401. CIRM training: EDUC4-12804. Dutch Research Council: NWO, 019.231EN.013. National Cancer Institute: P30CA023100",
   "MD5": "1cafefa32a897998e3e2ba0a29a3ef5c",
   "https://w3id.org/EVI#outputs": [
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-kolf-pan-genome-aggregated-data"
    },
    {
     "@id": "https://fairscape.net/api/ark:59853/dataset-kolf-pan-genome-aggregated-data-small"
    }
   ],
   "https://w3id.org/EVI#inputs": {
    "_summarized_by": "d4d rocrate normalize",
    "count": 90,
    "id_families": {
     "dataset": 90
    },
    "sample_ids": [
     "https://fairscape.net/api/ark:59853/dataset-molecule-data-gamma-chip-1-channel-1",
     "https://fairscape.net/api/ark:59853/dataset-molecule-data-alpha-chip-2-channel-4",
     "https://fairscape.net/api/ark:59853/dataset-molecule-data-gamma-chip-1-channel-8",
     "https://fairscape.net/api/ark:59853/dataset-molecule-data-beta-chip-2-channel-7",
     "https://fairscape.net/api/ark:59853/dataset-molecule-data-beta-chip-2-channel-1"
    ]
   },
   "principalInvestigator": "Trey Ideker",
   "confidentialityLevel": "Unrestricted",
   "citation": "Clark T; Parker J; Al Manir S; Axelsson U; Ballllosero Navarro F; Chinn B; Churas CP; Dailamy A; Doctor Y; Fall J; Forget A; Gao J; Hansen JN; Hu M; Johannesson A; Khaliq H; Lee YH; Lenkiewicz J; Levinson MA; Marquez C; Metallo C; Muralidharan M; Nourreddine S; Niestroy J; Obernier K; Pan E; Polacco B; Pratt D; Qian G; Schaffer L; Sigaeva A; Thaker S; Zhang Y; Bélisle-Pipon JC; Brandt C; Chen JY; Ding Y; Fodeh S; Krogan N; Lundberg E; Mali P; Payne-Foster P; Ratcliffe S; Ravitsky V; Sali A; Schulz W; Ideker T, 2025, \"Cell Maps for Artificial Intelligence - March 2025 Data Release (Beta)\", https://doi.org/10.18130/V3/B35XWX , https://dataverse.lib.virginia.edu/, V1",
   "usageInfo": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "additionalProperty": [
    {
     "@type": "PropertyValue",
     "name": "Completeness",
     "value": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release."
    },
    {
     "@type": "PropertyValue",
     "name": "Human Subject",
     "value": "None - data collected from commercially available cell lines"
    },
    {
     "@type": "PropertyValue",
     "name": "Prohibited Uses",
     "value": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval."
    }
   ],
   "rai:dataLimitations": "This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis.",
   "rai:dataBiases": "Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large.",
   "rai:dataUseCases": "AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable biologically-driven, interpretable ML applications, for example as proposed in Ma et al. 2018 (PMID: 29505029) and Kuenzi et al. 2020 (PMID: 33096023).",
   "rai:dataReleaseMaintenancePlan": "Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds.",
   "rai:dataCollection": "Data collection processes are generally described in Clark T et al. (2024) \"Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines\" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized. ",
   "rai:dataCollectionType": [
    "Perturb-seq; IF imaging; SEC-MS"
   ],
   "url": "https://figshare.com/s/ee85bb1880921326249b",
   "rai:dataCollectionTimeframe": [
    "9/1/2022",
    "1/31/2026"
   ],
   "ethicalReview": "Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca.",
   "humanSubjectResearch": "None - data collected from commercially available cell lines",
   "humanSubjectExemption": "Exempt — research with commercially available de-identified human cell lines does not constitute human subjects research.",
   "dataGovernanceCommittee": "Jilian Parker",
   "completeness": "These data are not yet in completed final form, and some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "prohibitedUses": "These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval.",
   "rai:dataCollectionMissingData": "Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incompletely overlap. Computed cell maps not included in this release.",
   "localEvidenceGraph": {
    "@id": "Perturb-Seq/cell-atlas/ro-crate-prov-graph.html"
   },
   "evi:processed": true,
   "ro-crate-metadata": "Perturb-Seq/cell-atlas/ro-crate-metadata.json"
  }
 ]
}

================================================================================
FILE: ai_ready_score.json
ROLE: AI-readiness self-assessment
SIZE: 6,003 characters
--------------------------------------------------------------------------------
{
  "name": "AI-Ready Score for Cell Maps for Artificial Intelligence - June 2026 Data Release (Beta)",
  "fairness": {
    "findable": {
      "has_content": true,
      "details": "Dataset has DOI: https://doi.org/10.18130/V3/HIGT4C"
    },
    "accessible": {
      "has_content": true,
      "details": "The RO-Crate's JSON-LD metadata is machine-readable and publicly accessible by design."
    },
    "interoperable": {
      "has_content": true,
      "details": "The dataset uses the schema.org vocabulary within the RO-Crate framework and conforms to the Croissant RAI specification for interoperability."
    },
    "reusable": {
      "has_content": true,
      "details": "License: https://creativecommons.org/licenses/by-nc-sa/4.0/"
    }
  },
  "provenance": {
    "transparent": {
      "has_content": true,
      "details": "53877 dataset(s) documented"
    },
    "traceable": {
      "has_content": true,
      "details": "1976 computation/experiment steps documented"
    },
    "interpretable": {
      "has_content": true,
      "details": "6 software instances documented"
    },
    "key_actors_identified": {
      "has_content": true,
      "details": "47 authors, Publisher: https://dataverse.lib.virginia.edu/, PI: Trey Ideker"
    }
  },
  "characterization": {
    "semantics": {
      "has_content": true,
      "details": "Data is semantically described using the schema.org vocabulary within a machine-readable RO-Crate."
    },
    "statistics": {
      "has_content": true,
      "details": "Total size: 19.9 TB"
    },
    "standards": {
      "has_content": true,
      "details": "20 schema(s) documented"
    },
    "potential_sources_of_bias": {
      "has_content": true,
      "details": "Data in this release was derived from commercially available de-identified human cell lines, and does not represent all biological variants which may be seen in the population at large."
    },
    "data_quality": {
      "has_content": true,
      "details": "Some datasets are under temporary pre-publication embargo. Protein-protein interaction (SEC-MS), protein localization (IF imaging), and CRISPRi perturbSeq data interrogate sets of proteins which incom..."
    }
  },
  "pre_model_explainability": {
    "data_documentation_template": {
      "has_content": true,
      "details": "Documentation is provided via the RO-Crate's structured JSON-LD metadata, this HTML Datasheet, and Croissant RAI properties."
    },
    "fit_for_purpose": {
      "has_content": true,
      "details": "Use cases: AI-ready datasets to support research in functional genomics, AI/machine learning model training, cellular process analysis, cell architectural changes, and interactions in presence of specific disease processes, treatment conditions, or genetic perturbations. A major goal is to enable biologically-driven, interpretable ML applications, for example as proposed in Ma et al. 2018 (PMID: 29505029) and Kuenzi et al. 2020 (PMID: 33096023)., Limitations: This is an interim release. It does not contain predicted cell maps, which will be added in future releases. The current release is most suitable for bioinformatics analysis of the individual datasets. Requires domain expertise for meaningful analysis."
    },
    "verifiable": {
      "has_content": true,
      "details": "0% of files have checksums (8/55859)"
    }
  },
  "ethics": {
    "ethically_acquired": {
      "has_content": true,
      "details": "Data collection: Data collection processes are generally described in Clark T et al. (2024) \"Cell Maps for Artificial Intelligence: AI-Ready Maps of Human Cell Architecture from Disease-Relevant Cell Lines\" bioRxiv 2024.05.21.589311; doi: https://doi.org/10.1101/2024.05.21.589311. Additional data collection details will be subsequently published once finalized. , Human subject info: None - data collected from commercially available cell lines"
    },
    "ethically_managed": {
      "has_content": true,
      "details": "Ethical review: Vardit Ravistky ravitskyv@thehastingscenter.org and Jean-Christophe Belisle-Pipon jean-christophe_belisle-pipon@sfu.ca., Governance: Jilian Parker"
    },
    "ethically_disseminated": {
      "has_content": true,
      "details": "License: https://creativecommons.org/licenses/by-nc-sa/4.0/, Prohibited uses: These laboratory data are not to be used in clinical decision-making or in any context involving patient care without appropriate regulatory oversight and approval."
    },
    "secure": {
      "has_content": true,
      "details": "Confidentiality level: Unrestricted"
    }
  },
  "sustainability": {
    "persistent": {
      "has_content": true,
      "details": "Dataset has DOI: https://doi.org/10.18130/V3/HIGT4C"
    },
    "domain_appropriate": {
      "has_content": true,
      "details": "Maintenance plan: Dataset will be regularly updated and augmented on a quarterly basis through the end of the project (November, 2026). Long term preservation in the https://dataverse.lib.virginia.edu/, supported by committed institutional funds."
    },
    "well_governed": {
      "has_content": true,
      "details": "Governance committee: Jilian Parker"
    },
    "associated": {
      "has_content": true,
      "details": "All data, software, and computations are explicitly linked within the RO-Crate's provenance graph."
    }
  },
  "computability": {
    "standardized": {
      "has_content": true,
      "details": "Formats: .d, .d directory group, .tsv, .xml, TSV..."
    },
    "computationally_accessible": {
      "has_content": true,
      "details": "Publisher: https://dataverse.lib.virginia.edu/"
    },
    "portable": {
      "has_content": true,
      "details": "The dataset is packaged as a self-contained RO-Crate, a standard designed for portability across systems."
    },
    "contextualized": {
      "has_content": true,
      "details": "Context is provided by the RO-Crate's graph structure and detailed in properties such as rai:dataLimitations."
    }
  }
}

