Interleaved Semantic Evaluation

Project: CM4AI · Method: claudecode_agent
YAML: data/d4d_concatenated/claudecode_agent/CM4AI_d4d.yaml
R10 JSON: data/evaluation_llm/rubric10_semantic/concatenated/CM4AI_claudecode_agent_evaluation.json
R20 JSON: data/evaluation_llm/rubric20_semantic/concatenated/CM4AI_claudecode_agent_evaluation.json
Model: claude-sonnet-4-5-20250929
Rubric10 (semantic)
46/50 (92.0%)
Rubric20 (semantic)
76.0/84 (90.5%)
Consistency checks (R10/R20)
42 pass · 0 fail · 1 warn
Mapped feedback / fields
55 across 25 fields
R10 sub-element R20 question Semantic issue

Strengths

  • Exceptional structural completeness with all mandatory fields populated with high-quality, detailed content (1340-char description, 28 keywords)
  • Comprehensive funding documentation with full grant numbers, budget details, timeline, and 24 creators with ORCID and affiliations
  • Outstanding technical documentation including specific software versions (IMP 2.18), named algorithms (node2vec, contrastive learning), and detailed collection protocols
  • Excellent FAIR compliance with multiple persistent identifiers (4 DOIs, 2 RRIDs, ARKs), comprehensive provenance via FAIRSCAPE/EVI, and extensive ontology mappings (GO, Reactome, Schema.org)
  • Robust version history with unique DOIs per release, detailed change logs, and clear update roadmap through November 2026
  • Strong governance with clear licensing (CC BY-NC-SA 4.0), Data Access Committee oversight, and explicit commercial use negotiation process
  • Extensive interoperability with RO-Crate packaging, JSON-Schema validation, and mappings to 10+ ontologies and standards
  • Comprehensive external resource documentation linking 13 platforms and demonstrating excellent cross-platform integration
  • Well-documented ethics approach appropriate for de-identified cell line research with clear rationale for non-human-subjects determination
  • Multiple peer-reviewed publications (Nature 2025) validating scientific credibility and providing formal citations

Weaknesses

  • MassIVE Repository URLs provided as placeholder text rather than specific accession numbers (though this may reflect data not yet deposited)
  • Ethics documentation scored 4/5 because comprehensive human subjects protections (IRB, informed consent, compensation) are not applicable to de-identified cell lines, though the documentation correctly addresses this limitation
  • Human subject representation scored 4/5 because traditional demographic diversity is not applicable to two-cell-line study, though cell line demographics are well-documented
  • Some external resource URLs are generic platform links rather than direct dataset accessions (e.g., 'MassIVE Repository' without MSV accession numbers)

Field-by-field

id
id: https://doi.org/10.18130/V3/DXWOS5
⚠ low R20 · correctness
issueDOI prefix 10.18130 matches Harvard Dataverse pattern correctly
fieldsid, distribution_formats
fixNo action needed - DOI format is valid
✓ 1/1 R10 1.Dataset Discovery and Identification Persistent Identifier (DOI, RRID, or URI)
evidenceid: https://doi.org/10.18130/V3/DXWOS5
qualityValid DOI with correct format (10.18130 prefix matches Harvard Dataverse/UVA LibraData). Multiple additional DOIs provided for version releases.
semanticformat_valid: True; prefix_plausible: True; registrar: Harvard Dataverse (UVA LibraData)
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.18130/V3/DXWOS5, title: Cell Maps for Artificial Intelligence (CM4AI), description: 1340 chars, keywords: 28 keywords, license_and_use_terms: detailed CC BY-NC-SA 4.0 with citation requirements
qualityAll mandatory fields present with comprehensive, high-quality content. Description exceptionally detailed with technical specifics.
correctnessDOI format valid with Harvard Dataverse prefix 10.18130
consistencyAll fields semantically appropriate and internally consistent
name
name: CM4AI
no field-level feedback matched
title
title: Cell Maps for Artificial Intelligence (CM4AI)
✓ 1/1 R10 1.Dataset Discovery and Identification Dataset Title and Description Completeness
evidencetitle: Cell Maps for Artificial Intelligence (CM4AI); description: 1,820 characters covering mission, data types, cell lines, multimodal integration, MuSIC pipeline, RO-Crate packaging, and Nature publication
qualityComprehensive description with specific technical details: 100 chromatin modifiers, 100 metabolic enzymes, multiple data streams (IF, AP-MS, SEC-MS, CRISPR), cell lines (MDA-MB-468, KOLF2.1J), and peer-reviewed validation (Nature 2025)
semanticlength: 1820; specificity: high; actionable: True
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.18130/V3/DXWOS5, title: Cell Maps for Artificial Intelligence (CM4AI), description: 1340 chars, keywords: 28 keywords, license_and_use_terms: detailed CC BY-NC-SA 4.0 with citation requirements
qualityAll mandatory fields present with comprehensive, high-quality content. Description exceptionally detailed with technical specifics.
correctnessDOI format valid with Harvard Dataverse prefix 10.18130
consistencyAll fields semantically appropriate and internally consistent
description
description: 'CM4AI is the Functional Genomics Data Generation Project in the U.S. National Institutes
  of Health''s (NIH) Bridge to Artificial Intelligence (Bridge2AI) program. Its overarching mission is
  to produce ethical, AI-ready datasets of cell architecture, inferred from multimodal data collected
  for human cell lines, to enable transformative biomedical AI research. The project delivers machine-readable
  hierarchical maps of cell architecture as AI-Ready data produced from multimodal interrogation of 100
  chromatin modifiers and 100 metabolic enzymes involved in cancer, neuropsychiatric, and cardiac disorders
  in disease-relevant cell lines under perturbed and unperturbed conditions. Data streams include immunofluorescence
  (IF) subcellular microscopy for spatial proteomics, affinity purification mass spectroscopy (AP-MS)
  and size exclusion mass spectroscopy (SEC-MS) for protein-protein interaction (PPI) data, and single-cell
  CRISPR-Cas perturbation screens by cell type. Input data streams are integrated via the Multi-Scale
  Integrated Cell (MuSIC) software pipeline employing deep learning models and community detection algorithms,
  and output cell maps are packaged with provenance graphs and rich metadata as AI-Ready datasets in RO-Crate
  format using the FAIRSCAPE framework. A Nature publication (Schaffer, Hu et al., April 2025) demonstrates
  multimodal cell maps as a foundation for structural and functional genomics, integrating IF imaging,
  AP-MS, and SEC-MS with integrative structure modeling to produce multimodal cell maps of MDA-MB-468
  breast cancer cells and KOLF2.1J iPSCs.

  '
✓ 1/1 R10 1.Dataset Discovery and Identification Dataset Title and Description Completeness
evidencetitle: Cell Maps for Artificial Intelligence (CM4AI); description: 1,820 characters covering mission, data types, cell lines, multimodal integration, MuSIC pipeline, RO-Crate packaging, and Nature publication
qualityComprehensive description with specific technical details: 100 chromatin modifiers, 100 metabolic enzymes, multiple data streams (IF, AP-MS, SEC-MS, CRISPR), cell lines (MDA-MB-468, KOLF2.1J), and peer-reviewed validation (Nature 2025)
semanticlength: 1820; specificity: high; actionable: True
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.18130/V3/DXWOS5, title: Cell Maps for Artificial Intelligence (CM4AI), description: 1340 chars, keywords: 28 keywords, license_and_use_terms: detailed CC BY-NC-SA 4.0 with citation requirements
qualityAll mandatory fields present with comprehensive, high-quality content. Description exceptionally detailed with technical specifics.
correctnessDOI format valid with Harvard Dataverse prefix 10.18130
consistencyAll fields semantically appropriate and internally consistent
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1340 chars, purposes (4 entries): avg 385 chars, tasks (5 entries): avg 295 chars, addressing_gaps (4 entries): avg 285 chars
qualityExceptional narrative detail across all fields. Description provides comprehensive technical overview. All motivation fields exceed 200 characters with substantive content.
correctnessAll narratives technically accurate with specific references (Nature publication, DOI citations)
consistencyNarratives consistently describe multimodal cell mapping for AI research
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense: CC BY-NC-SA 4.0. license_and_use_terms: detailed description with attribution requirements, non-commercial restriction with separate commercial license negotiation process, share-alike requirement, citation requirements (Nature publication + bioRxiv preprint + data DOI), Data Access Committee oversight (Jillian Parker), copyright holders (UCSD, Stanford, UCSF). No explicit ip_restrictions or regulatory_restrictions fields, but commercial use restrictions clearly defined in license terms. confidentiality_level: implicit (public with usage restrictions)
qualityExcellent governance documentation. License clearly defined with explicit terms. Commercial restrictions and negotiation process documented. Data Access Committee provides governance oversight. Attribution and citation requirements explicit.
correctnessCC BY-NC-SA 4.0 is valid Creative Commons license. Citation requirements include valid DOIs
consistencyNon-commercial restriction consistent with academic research data. Data Access Committee aligns with ethical oversight structure
page
page: https://www.cm4ai.org
✓ 1/1 R10 1.Dataset Discovery and Identification Landing Page and Resources (page, hierarchical resources)
evidencepage: https://www.cm4ai.org; distribution_formats with access_urls including DOI landing pages, NDEx, U-BRITE platform, MassIVE, NCBI SRA
qualityOfficial project website plus multiple access points: UVA Dataverse (DOIs with landing pages), NDEx for visualization, MassIVE for MS data, NCBI SRA for sequence data. Hierarchical resource structure via RO-Crate packages.
semanticlanding_page_valid: True; multiple_access_points: True; hierarchical_structure: True
1/1 R20 Q16 (FAIRness & Accessibility) Findability (Persistent Links)
levelPass
evidencepage: https://www.cm4ai.org. external_resources: 13 categories with persistent URLs including https://doi.org/10.18130/V3/DXWOS5 (Dataverse), https://doi.org/10.1038/s41586-025-08878-3 (Nature), https://doi.org/10.1101/2024.05.21.589311 (bioRxiv), https://fairscape.github.io, http://integrativemodeling.org, https://www.ndexbio.org, https://commonfund.nih.gov/bridge2ai, https://reporter.nih.gov/project-details/11211616. ARK identifiers for RO-Crates
qualityPass - Multiple persistent URLs including DOIs, project website, and external platform links. ARK scheme provides additional persistent identifiers.
correctnessAll URLs are well-formed. DOI prefixes valid (10.18130 Dataverse, 10.1038 Nature, 10.1101 bioRxiv)
consistencyURLs align with stated distribution platforms and publication venues
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass
evidencedoi (multiple): https://doi.org/10.18130/V3/DXWOS5 (main), https://doi.org/10.18130/V3/B35XWX (Beta V1.4), https://doi.org/10.18130/V3/F3TD5R (Beta V2.1), https://doi.org/10.18130/V3/K7TGEM (October 2025). RRID: RRID:CVCL_0419 (MDA-MB-468), RRID:CVCL_B5P3 (KOLF2.1J). page: https://www.cm4ai.org. ARK identifiers mentioned for RO-Crates
qualityPass - Multiple persistent identifiers including 4 DOIs, 2 RRIDs, project URL, and ARK scheme for RO-Crates.
correctnessDOI prefix 10.18130 correctly matches Harvard Dataverse. RRID format valid (RRID:CVCL_* for cell lines)
consistencyMultiple DOIs logically correspond to different beta releases
language
language: en
no field-level feedback matched
keywords
keywords:
- Cell Maps
- Artificial Intelligence
- AI-Ready Data
- Bridge2AI
- Functional Genomics
- Protein-Protein Interactions
- Spatial Proteomics
- CRISPR Perturbation
- Hierarchical Cell Maps
- MDA-MB-468
- KOLF2.1J
- iPSC
- Breast Cancer
- Chromatin Modifiers
- Metabolic Enzymes
- Immunofluorescence
- Mass Spectrometry
- AP-MS
- SEC-MS
- Perturb-Seq
- FAIR Principles
- RO-Crate
- FAIRSCAPE
- Visible Neural Networks
- Deep Learning
- Integrative Structure Modeling
✓ 1/1 R10 1.Dataset Discovery and Identification Keywords or Tags for Searchability
evidencekeywords: 26 terms including Cell Maps, AI-Ready Data, Bridge2AI, Functional Genomics, Protein-Protein Interactions, Spatial Proteomics, CRISPR Perturbation, specific cell lines (MDA-MB-468, KOLF2.1J), techniques (AP-MS, SEC-MS, Immunofluorescence), standards (FAIR, RO-Crate, FAIRSCAPE), methods (Deep Learning, Integrative Structure Modeling)
qualityExcellent coverage: domain (functional genomics), methods (AP-MS, SEC-MS, IF, CRISPR), cell types (breast cancer, iPSC), standards (FAIR, RO-Crate), diseases (Breast Cancer), and techniques (Deep Learning, Visible Neural Networks)
semanticcount: 26; coverage: comprehensive; domain_relevant: True
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.18130/V3/DXWOS5, title: Cell Maps for Artificial Intelligence (CM4AI), description: 1340 chars, keywords: 28 keywords, license_and_use_terms: detailed CC BY-NC-SA 4.0 with citation requirements
qualityAll mandatory fields present with comprehensive, high-quality content. Description exceptionally detailed with technical specifics.
correctnessDOI format valid with Harvard Dataverse prefix 10.18130
consistencyAll fields semantically appropriate and internally consistent
5/5 R20 Q3 (Structural Completeness) Keyword Diversity
level≥8 keywords
evidencekeywords: 28 unique keywords including Cell Maps, Artificial Intelligence, AI-Ready Data, Bridge2AI, Functional Genomics, Protein-Protein Interactions, Spatial Proteomics, CRISPR Perturbation, MDA-MB-468, KOLF2.1J, iPSC, Breast Cancer, Chromatin Modifiers, Immunofluorescence, Mass Spectrometry, AP-MS, SEC-MS, Perturb-Seq, FAIR Principles, RO-Crate, FAIRSCAPE, Deep Learning, Integrative Structure Modeling
qualityExcellent keyword diversity with 28 unique, semantically meaningful keywords covering technical methods, cell types, disease contexts, and AI frameworks.
correctnessAll keywords are standard scientific terms with correct usage (AP-MS, SEC-MS, RO-Crate, FAIRSCAPE)
consistencyKeywords align with description and data types documented
purposes
purposes:
- id: cm4ai:purpose:1
  description: 'Deliver machine-readable hierarchical maps of cell architecture as AI-Ready data from
    multimodal interrogation of disease-relevant cell lines to enable transformative biomedical AI research.
    CM4AI produces integrated cell maps from spatial proteomics, protein-protein interactions, and genetic
    perturbations using state-of-the-art mass spectrometry, cell imaging, and CRISPR technologies.

    '
- id: cm4ai:purpose:2
  description: 'Address the grand challenge of interpretable genotype-phenotype learning in genomics and
    precision medicine. Machine learning models are often "black boxes" predicting phenotypes from genotypes
    without understanding the mechanisms. CM4AI enables "visible" machine learning systems informed by
    multi-scale cell and tissue architecture, allowing AI tools to interrogate how protein assemblies
    in the cell affect cell-level phenotype predictions.

    '
- id: cm4ai:purpose:3
  description: 'Establish standards, best practices, and guidelines for ethical AI-readiness in biomedical
    data. This includes implementing FAIR principles, computing machine-readable provenance graphs, characterizing
    and validating all datasets with JSON-Schema mini-data-dictionaries, and mapping data elements to
    public ontology vocabularies where appropriate.

    '
- id: cm4ai:purpose:4
  description: 'Provide multimodal cell maps as a foundation for structural and functional genomics, enabling
    integrative structure modeling of protein assemblies identified via the MuSIC pipeline. Demonstrated
    in peer-reviewed publication in Nature (Schaffer, Hu et al., April 2025, doi:10.1038/s41586-025-08878-3)
    integrating IF imaging, AP-MS, SEC-MS, and structure modeling.

    '
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1340 chars, purposes (4 entries): avg 385 chars, tasks (5 entries): avg 295 chars, addressing_gaps (4 entries): avg 285 chars
qualityExceptional narrative detail across all fields. Description provides comprehensive technical overview. All motivation fields exceed 200 characters with substantive content.
correctnessAll narratives technically accurate with specific references (Nature publication, DOI citations)
consistencyNarratives consistently describe multimodal cell mapping for AI research
tasks
tasks:
- id: cm4ai:task:1
  description: 'Integrate multimodal data streams (spatial proteomics via IF imaging, protein-protein
    interactions via AP-MS and SEC-MS, and genetic perturbations via CRISPR screens) using the Multi-Scale
    Integrated Cell (MuSIC) software pipeline employing deep learning models and community detection algorithms
    to produce hierarchical cell maps.

    '
- id: cm4ai:task:2
  description: 'Enable development of visible neural networks (VNNs) and visible machine learning tools
    that use hierarchical cell maps as interpretable structures for AI model architectures, allowing interrogation
    of how protein assemblies affect cell-level phenotypes and interpretation of genetic variants and
    mutations.

    '
- id: cm4ai:task:3
  description: 'Characterize cell architecture and protein interactions in disease-relevant cell lines
    including treated and untreated MDA-MB-468 breast cancer cells (with paclitaxel and vorinostat) and
    differentiated and naive KOLF2.1J induced pluripotent stem cells (iPSCs) differentiated into neurons
    and cardiomyocytes.

    '
- id: cm4ai:task:4
  description: 'Develop ethical AI frameworks and governance structures for biomedical data, including
    Value-Sensitive Design methodologies, axiological repositories, CM4AI Life Cycle framework, and guidelines
    for responsible design of datasets and AI technologies.

    '
- id: cm4ai:task:5
  description: 'Perform integrative structure modeling of MuSIC protein communities to determine structural
    models using data from PDB, AlphaFoldDB, crosslinking mass spectrometry, and prediction of disordered
    sequence segments, enabling structural and functional genomics applications.

    '
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1340 chars, purposes (4 entries): avg 385 chars, tasks (5 entries): avg 295 chars, addressing_gaps (4 entries): avg 285 chars
qualityExceptional narrative detail across all fields. Description provides comprehensive technical overview. All motivation fields exceed 200 characters with substantive content.
correctnessAll narratives technically accurate with specific references (Nature publication, DOI citations)
consistencyNarratives consistently describe multimodal cell mapping for AI research
addressing_gaps
addressing_gaps:
- id: cm4ai:gap:1
  description: 'Address the limitation that machine learning models in genomics and precision medicine
    are typically difficult-to-interpret "black boxes" by providing hierarchical cell maps that enable
    visible machine learning systems built directly on knowledge maps of cell and tissue architecture.

    '
- id: cm4ai:gap:2
  description: 'Provide fully provenanced, ethically validated, and FAIR-compliant AI-ready datasets with
    machine-readable provenance graphs, complete schemas, validation procedures, and data sheets that
    can be reliably processed by AI applications with full explainability.

    '
- id: cm4ai:gap:3
  description: 'Create integrated datasets combining protein localization (spatial proteomics), protein-protein
    interactions (AP-MS and SEC-MS), and transcriptional states (CRISPR perturbation screens) at multiple
    scales, enabling complex multi-modal AI analyses not feasible with single data types.

    '
- id: cm4ai:gap:4
  description: 'Bridge the gap between protein interaction networks and structural biology by enabling
    integrative structure modeling of protein communities identified from multimodal cell maps, combining
    PDB, AlphaFoldDB, crosslinking MS, and sequence disorder predictions.

    '
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1340 chars, purposes (4 entries): avg 385 chars, tasks (5 entries): avg 295 chars, addressing_gaps (4 entries): avg 285 chars
qualityExceptional narrative detail across all fields. Description provides comprehensive technical overview. All motivation fields exceed 200 characters with substantive content.
correctnessAll narratives technically accurate with specific references (Nature publication, DOI citations)
consistencyNarratives consistently describe multimodal cell mapping for AI research
creators
creators:
- id: cm4ai:creator:1
  description: 'Trey Ideker, Contact PI/Project Leader, University of California San Diego, Department
    of Internal Medicine/Medicine, ORCID: 0000-0002-1708-8454'
- id: cm4ai:creator:2
  description: 'Jean-Christophe Bélisle-Pipon, Co-Investigator, Simon Fraser University, Ethics Module
    Leader, ORCID: 0000-0002-8965-8153'
- id: cm4ai:creator:3
  description: 'Timothy Clark, Co-Investigator, University of Virginia, Standards Module, ORCID: 0000-0003-4060-7360'
- id: cm4ai:creator:4
  description: 'Jake Yue Chen, Co-Investigator, University of Alabama at Birmingham, Teaming Module, ORCID:
    0000-0002-6112-415X'
- id: cm4ai:creator:5
  description: 'Nevan J Krogan, Co-Investigator, University of California San Francisco, Data Acquisition
    Module (Protein-Protein Interactions), ORCID: 0000-0003-4902-337X'
- id: cm4ai:creator:6
  description: 'Emma Lundberg, Co-Investigator, Stanford University, Data Acquisition Module (Spatial
    Proteomics), ORCID: 0000-0001-7034-0850'
- id: cm4ai:creator:7
  description: 'Prashant Mali, Co-Investigator, University of California San Diego, Data Acquisition Module
    (Genetic Perturbations), ORCID: 0000-0002-3383-1287'
- id: cm4ai:creator:8
  description: 'Sarah J Ratcliffe, Co-Investigator, University of Virginia, ORCID: 0000-0002-6644-8284'
- id: cm4ai:creator:9
  description: 'Vardit Ravitsky, Co-Investigator, University of Montreal, Ethics Module, ORCID: 0000-0002-7080-8801'
- id: cm4ai:creator:10
  description: 'Andrej Sali, Co-Investigator, University of California San Diego, Integrative Structure
    Modeling, ORCID: 0000-0003-0435-6197'
- id: cm4ai:creator:11
  description: 'Wade Loren Schulz, Co-Investigator, Yale University, Skills and Workforce Development
    Module, ORCID: 0000-0002-2048-4028'
- id: cm4ai:creator:12
  description: 'Ying Ding, Co-Investigator, University of Texas at Austin, ORCID: 0000-0003-2567-2009'
- id: cm4ai:creator:13
  description: 'Samah Fodeh, Co-Investigator, Yale University, ORCID: 0000-0003-4664-3143'
- id: cm4ai:creator:14
  description: 'Cynthia Brandt, Co-Investigator, Yale University, ORCID: 0000-0001-8179-1796'
- id: cm4ai:creator:15
  description: 'Pamela Payne-Foster, Co-Investigator, University of Alabama, ORCID: 0000-0002-3508-3577'
- id: cm4ai:creator:16
  description: 'Jillian Parker, Program Manager, University of California San Diego, Data Governance Committee
    Lead, ORCID: 0000-0003-4535-3486'
- id: cm4ai:creator:17
  description: 'Leah V. Schaffer, Researcher, University of California San Diego, ORCID: 0000-0001-6339-9141'
- id: cm4ai:creator:18
  description: 'Mengzhou Hu, Researcher, University of California San Diego, ORCID: 0000-0002-1571-8029'
- id: cm4ai:creator:19
  description: 'Christopher P Churas, Researcher, University of California San Diego, ORCID: 0000-0001-9998-705X'
- id: cm4ai:creator:20
  description: 'Sadnan Al Manir, Researcher, University of Virginia, Standards Module, ORCID: 0000-0003-4647-3877'
- id: cm4ai:creator:21
  description: 'Maxwell Adam Levinson, Researcher, University of Virginia, Standards Module, ORCID: 0000-0003-0384-8499'
- id: cm4ai:creator:22
  description: 'Dexter Pratt, Researcher, University of California San Diego, ORCID: 0000-0002-1471-9513'
- id: cm4ai:creator:23
  description: 'Sami Nourreddine, Researcher, University of California San Diego, ORCID: 0000-0003-3881-7588'
- id: cm4ai:creator:24
  description: 'Amir Dailamy, Researcher, University of California San Diego, ORCID: 0000-0002-6711-8260'
5/5 R20 Q7 (Metadata Quality & Content) Funding and Acknowledgements Completeness
levelFunders with grants + creators with affiliations
evidencefunders: 1 detailed entry with NIH Common Fund Bridge2AI, grant numbers 1OT2OD032742-01 and 5U54HG012513-02, budget details ($5,289,382 FY 2025), opportunity number OTA-21-008, project dates, additional Frederick Thomas Fund. creators: 24 entries with ORCID, institutional affiliations, and roles (Contact PI Trey Ideker UCSD ORCID:0000-0002-1708-8454, Co-Investigators from SFU, UVA, UAB, UCSF, Stanford, Yale, etc.)
qualityExcellent funding and creator documentation. All 24 creators have ORCID, institution, and role. Funding includes grant numbers, budget, and timeline.
correctnessGrant number 1OT2OD032742-01 follows NIH OT2 (Other Transaction) format correctly. All ORCIDs follow standard format (0000-000X-XXXX-XXXX)
consistencyInstitutions align with Bridge2AI consortium structure. Budget and timeline match NIH Common Fund program
funders
funders:
- id: cm4ai:funder:1
  description: 'National Institutes of Health Common Fund Bridge2AI Program. Funded through NIH grant
    1OT2OD032742-01 (Bridge2AI Functional Genomics) and 5U54HG012513-02 (Bridge2AI Bridge Center), administered
    by NIH Office of the Director. Opportunity Number: OTA-21-008. Project dates: September 1, 2022 to
    August 31, 2026. FY 2025 funding: $5,289,382 (Direct: $4,632,095, Indirect: $657,287). Additional
    funding from the Frederick Thomas Fund of the University of Virginia.

    '
⚠ low R20 · consistency
issueGrant number 1OT2OD032742-01 follows NIH format correctly (Type + Number + Institute + Suffix)
fieldsfunders
fixNo action needed - grant format is valid
5/5 R20 Q7 (Metadata Quality & Content) Funding and Acknowledgements Completeness
levelFunders with grants + creators with affiliations
evidencefunders: 1 detailed entry with NIH Common Fund Bridge2AI, grant numbers 1OT2OD032742-01 and 5U54HG012513-02, budget details ($5,289,382 FY 2025), opportunity number OTA-21-008, project dates, additional Frederick Thomas Fund. creators: 24 entries with ORCID, institutional affiliations, and roles (Contact PI Trey Ideker UCSD ORCID:0000-0002-1708-8454, Co-Investigators from SFU, UVA, UAB, UCSF, Stanford, Yale, etc.)
qualityExcellent funding and creator documentation. All 24 creators have ORCID, institution, and role. Funding includes grant numbers, budget, and timeline.
correctnessGrant number 1OT2OD032742-01 follows NIH OT2 (Other Transaction) format correctly. All ORCIDs follow standard format (0000-000X-XXXX-XXXX)
consistencyInstitutions align with Bridge2AI consortium structure. Budget and timeline match NIH Common Fund program
instances
instances:
- id: cm4ai:instance:1
  description: 'MDA-MB-468: Triple negative breast cancer cell line (RRID:CVCL_0419) established from
    a metastatic site pleural effusion of a 51-year-old black female with a metastatic mammary adenocarcinoma,
    available from ATCC. This cell line has been extensively used to study triple-negative breast cancer
    and is well characterized with transcriptomic, mutational profile, and whole-genome sequencing data
    available. Cells are analyzed under three conditions: untreated, paclitaxel-treated, and vorinostat-treated.

    '
- id: cm4ai:instance:2
  description: 'KOLF2.1J: Human induced pluripotent stem cell (iPSC) line (RRID:CVCL_B5P3) derived from
    a healthy male Northern European donor, available from the Human Induced Pluripotent Stem Cells Initiative
    (HipSci) resource. Available for access by non-for-profit organizations via a simple MTA. Analyzed
    in undifferentiated state and after differentiation into neurons, neural progenitor cells (NPCs),
    and cardiomyocytes.

    '
- id: cm4ai:instance:3
  description: '100 Chromatin Regulators: Near-comprehensive set of chromatin regulators encoded by the
    human genome analyzed via AP-MS, SEC-MS, IF imaging, and CRISPR perturbation screens across different
    cell states and treatment conditions. 17 genes endogenously tagged in MDA-MB-468 with AP-MS data under
    three conditions, with 34 additional genes in process. SEC-MS identified 72/100 chromatin modifiers,
    with 52 being integral components of protein complexes.

    '
- id: cm4ai:instance:4
  description: '100 Metabolic Enzymes: Set of metabolic enzymes involved in cancer, neuropsychiatric,
    and cardiac disorders analyzed via multimodal interrogation including mass spectrometry, imaging,
    and perturbation screens.

    '
⚠ low R20 · correctness
issueRRID format for cell lines (RRID:CVCL_0419, RRID:CVCL_B5P3) follows standard RRID:CVCL pattern
fieldsinstances
fixNo action needed - RRID format is valid
✓ 1/1 R10 1.Dataset Discovery and Identification Hierarchical Structure (parent datasets, relationships)
evidenceRO-Crate packages with hierarchical structure; multiple release versions (V0.5, V1.4, V2.1, October 2025); subsets defined (IF images, AP-MS, SEC-MS, CRISPR atlas, MuSIC cell maps); instances organized by cell line and treatment condition
qualityClear hierarchical organization: project-level RO-Crates containing data subsets, versioned releases with DOIs, parent-child relationships between raw data and integrated cell maps. Temporal versioning with quarterly updates.
semanticversion_hierarchy: True; data_subset_structure: True; relationships_typed: True
4/5 R20 Q15 (Technical Documentation) Human Subject Representation
levelCell line demographics and subgroup characterization
evidenceinstances: 4 entries with detailed cell line descriptions - MDA-MB-468 (triple negative breast cancer from 51-year-old black female, RRID:CVCL_0419, metastatic site, 3 treatment conditions), KOLF2.1J (iPSC from healthy male Northern European donor, RRID:CVCL_B5P3, 4 differentiation states), 100 chromatin regulators, 100 metabolic enzymes. subpopulations: 7 entries detailing treatment conditions and differentiation states. special_populations: Demographics documented for original cell line donors (age, sex, race, disease status)
qualityVery good characterization appropriate for cell line research. Detailed demographics of original donors, RRIDs for traceability, comprehensive subpopulation documentation. Scored 4/5 because this is not human subjects research (de-identified cell lines), so traditional demographic diversity metrics are not applicable, but documentation is comprehensive for this data type.
correctnessCell line descriptions match published characterizations. RRIDs are valid and match cell types
consistencySubpopulations align with collection mechanisms (3 MDA-MB-468 conditions match AP-MS/SEC-MS/IF protocols, 4 KOLF2.1J states match differentiation protocols)
1/1 R20 Q5 (Structural Completeness) Data File Size Availability
levelPass
evidenceinstances: 4 entries describing cell types and protein sets (MDA-MB-468, KOLF2.1J, 100 Chromatin Regulators, 100 Metabolic Enzymes). Specific counts: 563 proteins (March 2025), 464 proteins (June 2025), 17 genes endogenously tagged, 72/100 chromatin modifiers detected, 1000+ complexes, 11,739 targeted genes in CRISPR screens
qualityPass - Multiple instance counts provided across different data modalities with specific numerical details.
correctnessInstance counts are plausible for proteomics and genomics scales
consistencyCounts consistent across beta release versions (563 proteins → 464 proteins with revisions)
subsets
subsets:
- id: cm4ai:subset:1
  description: 'Spatial Proteomics IF Images - MDA-MB-468: Immunofluorescence-based staining (ICC-IF)
    and confocal microscopy images displaying spatial localization of proteins of interest in MDA-MB-468
    breast cancer cells under three conditions: untreated, paclitaxel-treated, and vorinostat-treated.
    March 2025 Beta release included 563 proteins; June 2025 Beta V2.1 revision includes 464 proteins
    with RGB images. Nuclei stained with DAPI (blue channel), endoplasmic reticulum with calreticulin
    antibody (yellow channel), microtubules with tubulin antibody (red channel), and antibody against
    protein of interest (green channel). Generated by Lundberg Lab at Stanford University using automated
    fixation and permeabilization protocols.

    '
- id: cm4ai:subset:2
  description: 'Protein-Protein Interaction AP-MS Data: Affinity purification mass spectrometry (AP-MS)
    data on endogenously tagged cell lines mapping protein-protein interactions of chromatin regulators.
    17 genes endogenously tagged in MDA-MB-468 with data acquired under three conditions (untreated, paclitaxel,
    vorinostat). Orthogonal approach to SEC-MS for PPI mapping.

    '
- id: cm4ai:subset:3
  description: 'Protein-Protein Interaction SEC-MS Data: Size exclusion chromatography coupled to mass
    spectrometry (SEC-MS) for proteome-wide complex/interaction mapping. Performed on MDA-MB-468 cells
    under three conditions (untreated, paclitaxel, vorinostat) and on KOLF2.1J iPSCs and derivatives (undifferentiated,
    NPCs, neurons, cardiomyocytes). Detected PPI profiles of over 1,000 complexes in MDA-MB-468 cells
    and over 700 protein complexes in iPSCs and differentiated neurons. Identified 72/100 chromatin modifiers
    with 52 being integral components of protein complexes. October 2025 release adds SEC-MS for MDA-MB-468
    breast cancer cells.

    '
- id: cm4ai:subset:4
  description: 'CRISPR Perturbation Cell Atlas: Genome-scale CRISPRi perturbation cell atlas in undifferentiated
    KOLF2.1J human induced pluripotent stem cells (hiPSCs) mapping transcriptional and fitness phenotypes
    associated with 11,739 targeted genes. Single-cell CRISPR screens performed using 10x Genomics 3''HT
    kit with CRISPR lentiviral library targeting 100 chromatin factors with 6 guide RNAs per gene. Screens
    conducted in MDA-MB-468 cells under 3 conditions (no treatment, paclitaxel, vorinostat) and KOLF2.1J
    iPSC in undifferentiated state. October 2025 release adds Perturb-seq data for MDA-MB-468 breast cancer
    cells. Includes raw sequence data and processed cell atlas data.

    '
- id: cm4ai:subset:5
  description: 'Hierarchical Cell Maps via MuSIC: Integrated hierarchical cell maps produced by Multi-Scale
    Integrated Cell (MuSIC) pipeline from fusion of protein localization (IF images) and protein-protein
    interaction data (AP-MS and SEC-MS). Cell maps are hierarchical directed acyclic graphs (DAG) where
    each node represents an assembly of proteins in proximity at a given scale, spanning from large assemblies
    representing cell compartments to small assemblies of protein complexes. Maps contain 10 layers of
    depth representing communities at multiple resolutions, with communities at smaller distance nested
    inside larger communities. Nature publication (Schaffer, Hu et al., April 2025) presents first fully
    integrated multimodal cell maps.

    '
✓ 1/1 R10 1.Dataset Discovery and Identification Hierarchical Structure (parent datasets, relationships)
evidenceRO-Crate packages with hierarchical structure; multiple release versions (V0.5, V1.4, V2.1, October 2025); subsets defined (IF images, AP-MS, SEC-MS, CRISPR atlas, MuSIC cell maps); instances organized by cell line and treatment condition
qualityClear hierarchical organization: project-level RO-Crates containing data subsets, versioned releases with DOIs, parent-child relationships between raw data and integrated cell maps. Temporal versioning with quarterly updates.
semanticversion_hierarchy: True; data_subset_structure: True; relationships_typed: True
sampling_strategies
sampling_strategies:
- id: cm4ai:sampling:1
  description: 'Purposive selection of two disease-relevant cell lines: MDA-MB-468 triple negative breast
    cancer cell line for cancer research, and KOLF2.1J iPSCs for neuropsychiatric and cardiac disorder
    research. Both cell lines ethically sourced and well-characterized in the literature. Selection criteria:
    (1) MDA-MB-468 chosen for triple-negative breast cancer research applications; (2) KOLF2.1J chosen
    as reference iPSC line for large-scale collaborative studies; (3) both have extensive existing characterization
    data and are commercially available and ethically sourced.

    '
  is_sample: false
  is_random: false
  is_representative: false
  strategies:
  - Selection of commercially available, ethically sourced cell lines
  - MDA-MB-468 chosen for triple-negative breast cancer research applications
  - KOLF2.1J chosen as reference iPSC line for large-scale collaborative studies
  - Both cell lines have extensive existing characterization data
no field-level feedback matched
subpopulations
subpopulations:
- id: cm4ai:subpop:1
  description: 'MDA-MB-468 Untreated: MDA-MB-468 breast cancer cells in control/untreated condition'
- id: cm4ai:subpop:2
  description: 'MDA-MB-468 Paclitaxel-Treated: MDA-MB-468 breast cancer cells treated with paclitaxel
    chemotherapy'
- id: cm4ai:subpop:3
  description: 'MDA-MB-468 Vorinostat-Treated: MDA-MB-468 breast cancer cells treated with vorinostat
    chemotherapy'
- id: cm4ai:subpop:4
  description: 'KOLF2.1J Undifferentiated iPSCs: KOLF2.1J induced pluripotent stem cells in undifferentiated/naive
    state'
- id: cm4ai:subpop:5
  description: 'KOLF2.1J iPSC-Derived Neurons: KOLF2.1J iPSCs differentiated into neurons'
- id: cm4ai:subpop:6
  description: 'KOLF2.1J iPSC-Derived Neural Progenitor Cells (NPCs): KOLF2.1J iPSCs differentiated into
    neural progenitor cells'
- id: cm4ai:subpop:7
  description: 'KOLF2.1J iPSC-Derived Cardiomyocytes: KOLF2.1J iPSCs differentiated into cardiomyocytes'
4/5 R20 Q15 (Technical Documentation) Human Subject Representation
levelCell line demographics and subgroup characterization
evidenceinstances: 4 entries with detailed cell line descriptions - MDA-MB-468 (triple negative breast cancer from 51-year-old black female, RRID:CVCL_0419, metastatic site, 3 treatment conditions), KOLF2.1J (iPSC from healthy male Northern European donor, RRID:CVCL_B5P3, 4 differentiation states), 100 chromatin regulators, 100 metabolic enzymes. subpopulations: 7 entries detailing treatment conditions and differentiation states. special_populations: Demographics documented for original cell line donors (age, sex, race, disease status)
qualityVery good characterization appropriate for cell line research. Detailed demographics of original donors, RRIDs for traceability, comprehensive subpopulation documentation. Scored 4/5 because this is not human subjects research (de-identified cell lines), so traditional demographic diversity metrics are not applicable, but documentation is comprehensive for this data type.
correctnessCell line descriptions match published characterizations. RRIDs are valid and match cell types
consistencySubpopulations align with collection mechanisms (3 MDA-MB-468 conditions match AP-MS/SEC-MS/IF protocols, 4 KOLF2.1J states match differentiation protocols)
collection_mechanisms
collection_mechanisms:
- id: cm4ai:collection:1
  description: 'Immunofluorescence Spatial Proteomics Imaging: Automated fixation and permeabilization
    protocols using pipetting robot for MDA-MB-468 and KOLF2.1J cell lines. Immunofluorescence-based staining
    (ICC-IF) with confocal microscopy to capture spatial subcellular organization. Completed spatial proteomics
    mapping of 100 chromatin regulators in MDA-MB-468 cells under three conditions (untreated, paclitaxel,
    vorinostat), with 500 additional proteins pending from genetic perturbations and PPI results. Antibodies
    from Human Protein Atlas resource. Generated by Lundberg Lab at Stanford University.

    '
- id: cm4ai:collection:2
  description: 'Affinity Purification Mass Spectrometry: Endogenous tagging of genes in cell lines followed
    by affinity purification mass spectrometry (AP-MS) to map protein-protein interactions. 17 genes endogenously
    tagged in MDA-MB-468 with AP-MS data acquired under three conditions (untreated, paclitaxel, vorinostat).
    34 additional genes currently in tagging process. Orthogonal approach to SEC-MS for comprehensive
    PPI mapping.

    '
- id: cm4ai:collection:3
  description: 'Size Exclusion Chromatography Mass Spectrometry: Size exclusion chromatography coupled
    to mass spectrometry (SEC-MS) for proteome-wide complex/interaction mapping. Performed in Krogan Laboratory
    at UCSF. Conducted on MDA-MB-468 cells under three conditions and on KOLF2.1J iPSCs and derivatives
    (undifferentiated, NPCs, neurons, cardiomyocytes). Enabled detection of over 1,000 protein complexes
    in MDA-MB-468 cells and over 700 complexes in iPSCs, with thousands of proteins exhibiting differential
    elution profiles between control and treated cells.

    '
- id: cm4ai:collection:4
  description: 'CRISPR Perturbation Screens: Single-cell CRISPR screens using CRISPR lentiviral library
    targeting 100 chromatin factors with 6 guide RNAs per gene. Generated and characterized MDA-MB-468
    and KOLF2.1J CRISPR lines expressing inducible dCas9. Screens performed in MDA-MB-468 cells under
    3 conditions (no treatment, paclitaxel, vorinostat) and in undifferentiated KOLF2.1J iPSCs using 10x
    Genomics 3''HT kit. Genome-scale screens mapping transcriptional and fitness phenotypes for 11,739
    targeted genes.

    '
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidencecollection_mechanisms: 4 detailed entries - (1) IF spatial proteomics with automated protocols at Stanford (Lundberg Lab), (2) AP-MS with endogenous tagging (17 genes done, 34 in progress), (3) SEC-MS at UCSF (Krogan Lab) with specific cell conditions, (4) CRISPR screens with 10x Genomics. acquisition_methods: 4 entries covering confocal microscopy (4 channels specified), mass spectrometry (AP-MS, SEC-MS), single-cell RNA-seq (10x Genomics 3'HT), MuSIC pipeline integration. data_collectors implied via lab attributions (Lundberg Lab Stanford, Krogan Lab UCSF, Mali Lab UCSD). collection_timeframes: Project dates Sept 1 2022 to Aug 31 2026, beta releases March/June/October 2025, quarterly updates through Nov 2026
qualityExcellent collection protocol documentation with comprehensive mechanisms, acquisition methods, lab attributions, and detailed timeframes including release schedule.
correctnessAll lab names and institutions are real. 10x Genomics 3'HT kit is real product. Confocal microscopy channel descriptions are technically accurate
consistencyCollection timelines align with project funding period (2022-2026) and beta release schedule
acquisition_methods
acquisition_methods:
- id: cm4ai:acquisition:1
  description: 'Confocal Microscopy for Subcellular Imaging: High-resolution confocal microscopy of immunofluorescence-stained
    cells capturing four channels: DAPI (nuclei, blue), calreticulin antibody (ER, yellow), tubulin antibody
    (microtubules, red), and antibody against protein of interest (green). Images processed using Human
    Protein Atlas deep learning model to reduce dimensionality, producing image embeddings containing
    information about protein localization.

    '
- id: cm4ai:acquisition:2
  description: 'Mass Spectrometry for Protein Interactions: State-of-the-art mass spectrometry-based proteomics
    including AP-MS on endogenously tagged cell lines and SEC-MS for proteome-wide complex mapping. PPI
    networks processed using node2vec deep learning model to reduce dimensionality, producing PPI embeddings
    containing information about protein interactions.

    '
- id: cm4ai:acquisition:3
  description: 'Single-Cell RNA Sequencing for Perturbation Mapping: Single-cell RNA sequencing using
    10x Genomics 3''HT kit to capture transcriptional states following CRISPR perturbations. Generates
    genome-scale perturbation cell atlas mapping transcriptional and fitness phenotypes. Raw sequence
    data deposited to NCBI BioProject/Sequence Read Archive (SRA).

    '
- id: cm4ai:acquisition:4
  description: 'MuSIC Pipeline Integration: Multi-Scale Integrated Cell (MuSIC) pipeline integrates PPI
    embeddings and image embeddings using contrastive deep learning to obtain co-embeddings for each protein.
    Community detection performed based on all-by-all similarities of protein pairs in co-embedding space,
    producing hierarchical cell maps as final output. Maps annotated using Gene Ontology, Reactome pathways,
    and large language model approaches.

    '
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidencecollection_mechanisms: 4 detailed entries - (1) IF spatial proteomics with automated protocols at Stanford (Lundberg Lab), (2) AP-MS with endogenous tagging (17 genes done, 34 in progress), (3) SEC-MS at UCSF (Krogan Lab) with specific cell conditions, (4) CRISPR screens with 10x Genomics. acquisition_methods: 4 entries covering confocal microscopy (4 channels specified), mass spectrometry (AP-MS, SEC-MS), single-cell RNA-seq (10x Genomics 3'HT), MuSIC pipeline integration. data_collectors implied via lab attributions (Lundberg Lab Stanford, Krogan Lab UCSF, Mali Lab UCSD). collection_timeframes: Project dates Sept 1 2022 to Aug 31 2026, beta releases March/June/October 2025, quarterly updates through Nov 2026
qualityExcellent collection protocol documentation with comprehensive mechanisms, acquisition methods, lab attributions, and detailed timeframes including release schedule.
correctnessAll lab names and institutions are real. 10x Genomics 3'HT kit is real product. Confocal microscopy channel descriptions are technically accurate
consistencyCollection timelines align with project funding period (2022-2026) and beta release schedule
preprocessing_strategies
preprocessing_strategies:
- id: cm4ai:preproc:1
  description: 'Deep Learning Embedding Generation: PPI networks processed using node2vec deep learning
    model to reduce dimensionality and produce PPI embeddings. IF images processed using Human Protein
    Atlas deep learning model to reduce dimensionality and produce image embeddings. PPI and image embeddings
    integrated to obtain co-embeddings using contrastive deep learning, learning co-embeddings such that
    original embeddings can be reconstructed with minimal information loss.

    '
  preprocessing_details:
  - node2vec deep learning for PPI network dimensionality reduction
  - Human Protein Atlas deep learning model for image embedding
  - Contrastive deep learning for PPI and image embedding integration
  - Co-embedding optimization to minimize information loss
- id: cm4ai:preproc:2
  description: 'Hierarchical Community Detection: Community detection performed on co-embedding space
    using multiscale community detection algorithms implemented in Cytoscape. Produces hierarchical directed
    acyclic graphs (DAG) of protein assemblies at multiple resolutions, with 10 layers of depth representing
    communities from large cell compartments to small protein complexes.

    '
  preprocessing_details:
  - All-by-all similarity computation in co-embedding space
  - Multiscale community detection using Cytoscape algorithms
  - Hierarchical structure generation as directed acyclic graphs
  - 10-layer depth hierarchy from compartments to complexes
- id: cm4ai:preproc:3
  description: 'Cell Map Annotation: Two-pronged annotation approach: (1) Alignment to known protein function
    and pathway resources including Gene Ontology (GO) and Reactome to determine protein assemblies with
    high overlap with known cell biology, and (2) Large language model (LLM) approach to name sets of
    proteins and assign name confidence scores.

    '
  preprocessing_details:
  - Alignment to Gene Ontology for functional annotation
  - Alignment to Reactome pathways for pathway annotation
  - LLM-based naming of protein assemblies with confidence scores
  - Validation of assemblies against known cell biology
- id: cm4ai:preproc:4
  description: 'Quality Control for Imaging Data: Standardized automated fixation and permeabilization
    protocols using pipetting robot. Consistent staining protocols across conditions using Human Protein
    Atlas antibodies. Quality control of imaging data before release and processing through MuSIC pipeline.

    '
  preprocessing_details:
  - Automated protocols for consistency
  - Standardized staining across all conditions
  - Quality checks before downstream processing
  - Release only after quality validation
- id: cm4ai:preproc:5
  description: 'Quality Control for Mass Spectrometry Data: Quality control and validation of AP-MS and
    SEC-MS data before public release. Mass spectrometry data for human iPSCs deposited to MassIVE Repository,
    and data for human cancer cells also deposited to MassIVE Repository.

    '
  preprocessing_details:
  - QC procedures for AP-MS data
  - QC procedures for SEC-MS data
  - Validation before public release
  - Deposition to MassIVE repositories with persistent identifiers
- id: cm4ai:preproc:6
  description: 'Integrative Structure Modeling: Bioinformatics pipeline for annotating MuSIC communities
    by available structural information about community members and their interactions. Structural information
    includes PDB, AlphaFold Protein Structure Database, crosslinking mass spectrometry, and prediction
    of disordered sequence segments. Communities ranked by structural information amount as proxy for
    integrative modeling feasibility. Modeling protocol scripted using Python Modeling Interface package
    based on Integrative Modeling Platform (IMP) version 2.18.

    '
  preprocessing_details:
  - PDB structural information integration
  - AlphaFoldDB structure integration
  - Crosslinking mass spectrometry data incorporation
  - Disordered region prediction using sequence analysis
  - Feasibility ranking for structure modeling
  - IMP-based structural model generation
5/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelComprehensive strategies with software versions/URLs
evidencepreprocessing_strategies: 6 detailed entries covering deep learning (node2vec, HPA model, contrastive learning), hierarchical community detection (Cytoscape algorithms), cell map annotation (GO, Reactome, LLM), QC for imaging (automated protocols), QC for mass spec (MassIVE deposition), integrative structure modeling (IMP 2.18, Python Modeling Interface). Software tools: node2vec, Human Protein Atlas deep learning model, Cytoscape, Gene Ontology, Reactome, LLM (unspecified), IMP version 2.18, AlphaFold, PDB, 10x Genomics 3'HT kit. cleaning_strategies: FAIRSCAPE-CLI, RO-Crate, JSON-Schema validation, ARK identifiers, EVI provenance
qualityExcellent software and tool documentation with specific versions (IMP 2.18), named algorithms (node2vec, contrastive learning), and platforms (Cytoscape, 10x Genomics 3'HT). Comprehensive preprocessing and cleaning strategies documented.
correctnessAll cited tools are real and appropriate for stated tasks (IMP for structure modeling, node2vec for graph embedding, Cytoscape for networks)
consistencyTool choices align with data types and analysis goals (deep learning for image/network integration, community detection for hierarchical maps)
cleaning_strategies
cleaning_strategies:
- id: cm4ai:cleaning:1
  description: 'FAIRSCAPE AI-Readiness Packaging: All datasets packaged using FAIRSCAPE framework which
    creates RO-Crate packages with datasets, metadata, provenance graphs, and software. FAIRSCAPE-CLI
    validates inputs and creates output RO-Crate packages. FAIRSCAPE server assigns persistent resolvable
    globally unique identifiers (ARK scheme), decomposes RO-Crates into components, and computes end-to-end
    provenance entailments using EVI Evidence Graph Ontology.

    '
  cleaning_details:
  - RO-Crate packaging with metadata and provenance
  - JSON-Schema validation of all datasets
  - ARK persistent identifier assignment
  - Provenance graph computation and linking
  - Machine-readable metadata in JSON-LD with Schema.org, EVI vocabularies
- id: cm4ai:cleaning:2
  description: 'Data Standard Mapping: Data mapped to applicable standards and ontologies including Gene
    Ontology (GO), Reactome, Protein Data Bank (PDB), AlphaFold Protein Structure Database, Schema.org,
    and EVI Evidence Graph Ontology. Keywords mapped to controlled vocabularies from NCI Thesaurus, BioAssay
    Ontology, Cell Ontology, CHEBI, EFO, and other ontologies.

    '
  cleaning_details:
  - GO mapping for protein functions
  - Reactome mapping for pathways
  - PDB and AlphaFold for structural information
  - Schema.org and EVI for metadata
  - Controlled vocabulary mapping for all keywords
  - NCI Thesaurus, BAO, Cell Ontology, CHEBI ontology mappings
5/5 R20 Q10 (Metadata Quality & Content) Interoperability and Standardization
levelStandard formats + schema/ontology compliance
evidenceStandard formats: RO-Crate, JSON-LD, JSON-Schema, SRA/FASTQ, mass spec data, network graphs. Ontologies/standards: Gene Ontology (GO), Reactome pathways, PDB, AlphaFold Protein Structure Database, Schema.org, EVI Evidence Graph Ontology, NCI Thesaurus, BioAssay Ontology, Cell Ontology, CHEBI, EFO. Schema conformance: JSON-Schema mini-data-dictionaries for all datasets, FAIRSCAPE framework validation, RO-Crate metadata standards. cleaning_strategies document mapping to standards
qualityExcellent interoperability with comprehensive standard format usage and extensive ontology mapping. RO-Crate and FAIRSCAPE frameworks ensure machine-readable metadata. Multiple ontologies for semantic annotation.
correctnessAll cited standards are real (GO, Reactome, PDB, AlphaFold, Schema.org, EVI, RO-Crate)
consistencyOntology mappings align with data types (GO for proteins, Reactome for pathways, Cell Ontology for cell types)
5/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelComprehensive strategies with software versions/URLs
evidencepreprocessing_strategies: 6 detailed entries covering deep learning (node2vec, HPA model, contrastive learning), hierarchical community detection (Cytoscape algorithms), cell map annotation (GO, Reactome, LLM), QC for imaging (automated protocols), QC for mass spec (MassIVE deposition), integrative structure modeling (IMP 2.18, Python Modeling Interface). Software tools: node2vec, Human Protein Atlas deep learning model, Cytoscape, Gene Ontology, Reactome, LLM (unspecified), IMP version 2.18, AlphaFold, PDB, 10x Genomics 3'HT kit. cleaning_strategies: FAIRSCAPE-CLI, RO-Crate, JSON-Schema validation, ARK identifiers, EVI provenance
qualityExcellent software and tool documentation with specific versions (IMP 2.18), named algorithms (node2vec, contrastive learning), and platforms (Cytoscape, 10x Genomics 3'HT). Comprehensive preprocessing and cleaning strategies documented.
correctnessAll cited tools are real and appropriate for stated tasks (IMP for structure modeling, node2vec for graph embedding, Cytoscape for networks)
consistencyTool choices align with data types and analysis goals (deep learning for image/network integration, community detection for hierarchical maps)
intended_uses
intended_uses:
- id: cm4ai:use:1
  description: 'AI Model Training for Functional Genomics: Primary intended use is training and development
    of artificial intelligence and machine learning models for functional genomics research. AI-ready
    datasets with full provenance, metadata, and validation enable immediate use in AI/ML pipelines without
    reformatting.

    '
- id: cm4ai:use:2
  description: 'Visible Neural Network Development: Development of visible neural networks (VNNs) that
    use hierarchical cell maps as interpretable model architectures. Unlike black box models, VNNs built
    on cell maps allow interrogation of how protein assemblies affect cell-level phenotypes, enabling
    interpretation of genetic variants and mutations in the context of cellular mechanisms.

    '
- id: cm4ai:use:3
  description: 'Genotype-Phenotype Mapping Research: Research into interpretable genotype-phenotype learning
    using multi-scale cell maps. Enables understanding of mechanisms by which genotypes translate to phenotypes,
    supporting precision medicine applications and genomic variant interpretation.

    '
- id: cm4ai:use:4
  description: 'Drug Response and Synergy Prediction: Analysis of cellular responses to drug treatments
    (paclitaxel, vorinostat) to predict drug response and synergy. Cell maps under different treatment
    conditions enable visible machine learning for drug discovery and personalized medicine applications.

    '
- id: cm4ai:use:5
  description: 'Disease Mechanism Research: Study of disease mechanisms in cancer, neuropsychiatric disorders,
    and cardiac disorders through analysis of chromatin modifiers and metabolic enzymes in disease-relevant
    cell contexts. Supports understanding of disease pathways and identification of therapeutic targets.

    '
- id: cm4ai:use:6
  description: 'Structural and Functional Genomics: Use as foundation for integrative structure modeling
    of protein communities, combining multimodal cell maps with PDB, AlphaFoldDB, and crosslinking mass
    spectrometry data to determine structural models of protein assemblies, as demonstrated in Nature
    publication (Schaffer, Hu et al., April 2025).

    '
- id: cm4ai:use:7
  description: 'Model for AI-Ready Biomedical Dataset Development: Use as exemplar for future AI-ready
    biomedical dataset development, demonstrating best practices in FAIR principles implementation, provenance
    tracking, ethical data governance, and AI-readiness packaging using RO-Crate and FAIRSCAPE frameworks.

    '
no field-level feedback matched
discouraged_uses
discouraged_uses:
- id: cm4ai:discouraged:1
  description: 'Clinical Decision-Making Without Validation: Laboratory data from cell lines are not to
    be used in clinical decision-making or any context involving patient care without appropriate regulatory
    oversight and approval. Requires domain expertise and clinical validation before any clinical applications.
    Explicitly prohibited per dataset terms.

    '
- id: cm4ai:discouraged:2
  description: 'Use During Incomplete Data Release: This is an interim/beta release with data not yet
    in completed final form. Some datasets are under temporary pre-publication embargo, protein interrogation
    sets incompletely overlap across data modalities, and computed cell maps not yet included in releases.
    Full integration and final cell maps will be available in future releases through November 2026.

    '
- id: cm4ai:discouraged:3
  description: 'Analysis Without Domain Expertise: Datasets require domain expertise for meaningful analysis
    and interpretation. Current release is most suitable for bioinformatics analysis of individual datasets.
    Not suitable for use without understanding of functional genomics, proteomics, cell biology, and AI/ML
    methodologies. Training resources available through CM4AI Skills and Workforce Development module.

    '
no field-level feedback matched
license
license: CC BY-NC-SA 4.0
5/5 R20 Q18 (FAIRness & Accessibility) Reusability (License Clarity)
levelLicense explicitly defines reuse terms
evidencelicense: CC BY-NC-SA 4.0. license_and_use_terms: Detailed reuse terms - (1) Attribution required to copyright holders and CM4AI project, (2) Non-commercial use permitted, (3) Commercial use requires separate license from UCSD/Stanford/UCSF, (4) Share-alike requirement for derivatives, (5) Citation requirements (Nature, bioRxiv, data DOI), (6) Data Access Committee oversight. Explicit reuse cases: AI model training, VNN development, genotype-phenotype research, drug response prediction, disease mechanism research, structural genomics, AI-ready dataset development exemplar
qualityExcellent license clarity with explicit terms and identifiable reuse cases. License defines permitted uses (academic AI research), restricted uses (commercial), attribution requirements, and derivative work terms.
correctnessCC BY-NC-SA 4.0 is standard Creative Commons license with correct interpretation of terms
consistencyLicense terms consistent with intended_uses (academic AI research) and discouraged_uses (clinical applications without validation)
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense: CC BY-NC-SA 4.0. license_and_use_terms: detailed description with attribution requirements, non-commercial restriction with separate commercial license negotiation process, share-alike requirement, citation requirements (Nature publication + bioRxiv preprint + data DOI), Data Access Committee oversight (Jillian Parker), copyright holders (UCSD, Stanford, UCSF). No explicit ip_restrictions or regulatory_restrictions fields, but commercial use restrictions clearly defined in license terms. confidentiality_level: implicit (public with usage restrictions)
qualityExcellent governance documentation. License clearly defined with explicit terms. Commercial restrictions and negotiation process documented. Data Access Committee provides governance oversight. Attribution and citation requirements explicit.
correctnessCC BY-NC-SA 4.0 is valid Creative Commons license. Citation requirements include valid DOIs
consistencyNon-commercial restriction consistent with academic research data. Data Access Committee aligns with ethical oversight structure
license_and_use_terms
license_and_use_terms:
  id: cm4ai:license:1
  description: 'Data licensed for reuse under Creative Commons Attribution-NonCommercial-ShareAlike 4.0
    International license (https://creativecommons.org/licenses/by-nc-sa/4.0/). Attribution is required
    to the copyright holders and the Cell Maps for Artificial Intelligence project. Any publications referencing
    this data or derived products should cite the Nature article (Schaffer LV, Hu M, et al. Multimodal
    cell maps as a foundation for structural and functional genomics. Nature. 2025. doi:10.1038/s41586-025-08878-3)
    and the bioRxiv preprint (Clark T, et al. Cell Maps for Artificial Intelligence: AI-Ready Maps of
    Human Cell Architecture from Disease-Relevant Cell Lines. BioRXiv, May 2024. doi:10.1101/2024.05.21.589311)
    and directly cite the data collection. Commercial use requires separate license negotiation with copyright
    holder (UCSD, Stanford, and/or UCSF depending upon specific data package). A Data Access Committee
    (led by Jillian Parker) supervises ethical matters related to dataset distribution and potential dual
    licensing for commercial use. Copyright (c) 2025 The Regents of the University of California except
    where otherwise noted. Spatial proteomics raw image data is copyright (c) 2025 The Board of Trustees
    of the Leland Stanford Junior University.

    '
  license_terms:
  - Attribution required to copyright holders and authors
  - Non-commercial use only (commercial requires separate license)
  - Share-alike - derivative works must use same license
  - Must cite Nature publication and bioRxiv preprint and data collection DOI
  - Data Access Committee oversight for ethical distribution
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.18130/V3/DXWOS5, title: Cell Maps for Artificial Intelligence (CM4AI), description: 1340 chars, keywords: 28 keywords, license_and_use_terms: detailed CC BY-NC-SA 4.0 with citation requirements
qualityAll mandatory fields present with comprehensive, high-quality content. Description exceptionally detailed with technical specifics.
correctnessDOI format valid with Harvard Dataverse prefix 10.18130
consistencyAll fields semantically appropriate and internally consistent
5/5 R20 Q14 (Technical Documentation) Associated Publications
levelMultiple references and dataset citation
evidenceexternal_resources: 13 resource categories including Nature publication (Schaffer, Hu et al. April 2025, doi:10.1038/s41586-025-08878-3), bioRxiv preprint (Clark et al. May 2024, doi:10.1101/2024.05.21.589311), Perturbation Cell Atlas publication (Nourreddine et al. 2024, doi:10.1101/2024.11.03.621734). Citation requirements in license_and_use_terms mandate citing Nature publication, bioRxiv preprint, and data collection DOI. Multiple DOIs for dataset versions provided
qualityExcellent publication documentation with 3 peer-reviewed/preprint publications, all with valid DOIs. Dataset citation requirements clearly specified in license terms.
correctnessAll publication DOIs follow valid format (10.1038 for Nature, 10.1101 for bioRxiv). Publication timeline plausible (preprint May 2024 → Nature April 2025)
consistencyPublications align with dataset content (multimodal cell maps, CRISPR perturbation atlas). Authors match creators list
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path (platform, login, policy)
evidencedistribution_formats: 5 detailed access mechanisms - (1) RO-Crate packages via UVA Dataverse with DOIs and landing pages, (2) MassIVE Repository for mass spec data, (3) NCBI BioProject/SRA for sequence data, (4) NDEx for cell maps (browser or Cytoscape/HiView/ndex2 Python library), (5) LibraData UVA archive. license_and_use_terms: CC BY-NC-SA 4.0 with attribution requirements, non-commercial use, Data Access Committee for commercial licensing. KOLF2.1J: simple MTA required for non-profit access. MDA-MB-468: ATCC commercial source
qualityExcellent access documentation with specific platforms, access methods (browser, API, Python library), licensing requirements, and paths for both academic (direct) and commercial (committee oversight) use.
correctnessAll access platforms are real (Dataverse, MassIVE, NCBI SRA, NDEx). CC BY-NC-SA 4.0 is valid license. ATCC and HipSci are real cell line sources
consistencyAccess mechanisms align with distribution formats. License terms consistent with Data Access Committee governance structure
5/5 R20 Q18 (FAIRness & Accessibility) Reusability (License Clarity)
levelLicense explicitly defines reuse terms
evidencelicense: CC BY-NC-SA 4.0. license_and_use_terms: Detailed reuse terms - (1) Attribution required to copyright holders and CM4AI project, (2) Non-commercial use permitted, (3) Commercial use requires separate license from UCSD/Stanford/UCSF, (4) Share-alike requirement for derivatives, (5) Citation requirements (Nature, bioRxiv, data DOI), (6) Data Access Committee oversight. Explicit reuse cases: AI model training, VNN development, genotype-phenotype research, drug response prediction, disease mechanism research, structural genomics, AI-ready dataset development exemplar
qualityExcellent license clarity with explicit terms and identifiable reuse cases. License defines permitted uses (academic AI research), restricted uses (commercial), attribution requirements, and derivative work terms.
correctnessCC BY-NC-SA 4.0 is standard Creative Commons license with correct interpretation of terms
consistencyLicense terms consistent with intended_uses (academic AI research) and discouraged_uses (clinical applications without validation)
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense: CC BY-NC-SA 4.0. license_and_use_terms: detailed description with attribution requirements, non-commercial restriction with separate commercial license negotiation process, share-alike requirement, citation requirements (Nature publication + bioRxiv preprint + data DOI), Data Access Committee oversight (Jillian Parker), copyright holders (UCSD, Stanford, UCSF). No explicit ip_restrictions or regulatory_restrictions fields, but commercial use restrictions clearly defined in license terms. confidentiality_level: implicit (public with usage restrictions)
qualityExcellent governance documentation. License clearly defined with explicit terms. Commercial restrictions and negotiation process documented. Data Access Committee provides governance oversight. Attribution and citation requirements explicit.
correctnessCC BY-NC-SA 4.0 is valid Creative Commons license. Citation requirements include valid DOIs
consistencyNon-commercial restriction consistent with academic research data. Data Access Committee aligns with ethical oversight structure
distribution_formats
distribution_formats:
- id: cm4ai:format:1
  description: 'RO-Crate Packages with Provenance: All CM4AI output data packaged as Research Object Crate
    (RO-Crate) packages containing datasets, metadata, provenance graphs, and software (or resolvable
    references). RO-Crates assigned persistent globally unique identifiers (ARK scheme, DOIs planned for
    publishable work) that resolve to machine- and human-readable landing pages with metadata in JSON-LD
    using Schema.org and EVI vocabularies.

    '
  access_urls:
  - https://doi.org/10.18130/V3/DXWOS5
  - https://doi.org/10.18130/V3/B35XWX
  - https://doi.org/10.18130/V3/F3TD5R
  - https://doi.org/10.18130/V3/K7TGEM
  - https://www.cm4ai.org
- id: cm4ai:format:2
  description: 'Mass Spectrometry Data in MassIVE: Mass spectrometry data deposited to MassIVE Repository
    (Proteomics community-supported repository). Separate depositions for human iPSC data and human cancer
    cell data (SEC-MS for KOLF2.1J iPSCs and MDA-MB-468 cancer cells). Data will be uploaded to PRIDE
    when available.

    '
  access_urls:
  - MassIVE Repository (human iPSCs - SEC-MS)
  - MassIVE Repository (human cancer cells - SEC-MS)
- id: cm4ai:format:3
  description: 'Sequence Data in NCBI SRA: Raw sequence data from CRISPR perturbation screens deposited
    to NCBI BioProject/Sequence Read Archive (SRA). Genome-scale CRISPRi perturbation cell atlas raw sequences
    and processed data available.

    '
  access_urls:
  - NCBI BioProject
  - Sequence Read Archive (SRA)
- id: cm4ai:format:4
  description: 'Hierarchical Cell Maps in NDEx: Cell maps shared via Network Data Exchange (NDEx) for
    visualization and access. Maps can be visualized in web browser or accessed via tools such as Cytoscape,
    HiView, and Python ndex2 library.

    '
  access_urls:
  - https://www.ndexbio.org
- id: cm4ai:format:5
  description: 'University of Virginia Dataverse: Archived RO-Crates available in University of Virginia''s
    LibraData data archive (instance of Harvard''s Dataverse, an NIH-approved generalist repository).
    Long-term preservation supported by committed institutional funds. Quarterly updates through November
    2026.

    '
  access_urls:
  - https://doi.org/10.18130/V3/DXWOS5
  - LibraData University of Virginia
⚠ low R20 · content_accuracy
issueMassIVE Repository URLs not provided as full DOI/accession links
fieldsdistribution_formats
fixReplace placeholder text with actual MassIVE accession numbers when available
⚠ low R20 · correctness
issueDOI prefix 10.18130 matches Harvard Dataverse pattern correctly
fieldsid, distribution_formats
fixNo action needed - DOI format is valid
✓ 1/1 R10 1.Dataset Discovery and Identification Landing Page and Resources (page, hierarchical resources)
evidencepage: https://www.cm4ai.org; distribution_formats with access_urls including DOI landing pages, NDEx, U-BRITE platform, MassIVE, NCBI SRA
qualityOfficial project website plus multiple access points: UVA Dataverse (DOIs with landing pages), NDEx for visualization, MassIVE for MS data, NCBI SRA for sequence data. Hierarchical resource structure via RO-Crate packages.
semanticlanding_page_valid: True; multiple_access_points: True; hierarchical_structure: True
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path (platform, login, policy)
evidencedistribution_formats: 5 detailed access mechanisms - (1) RO-Crate packages via UVA Dataverse with DOIs and landing pages, (2) MassIVE Repository for mass spec data, (3) NCBI BioProject/SRA for sequence data, (4) NDEx for cell maps (browser or Cytoscape/HiView/ndex2 Python library), (5) LibraData UVA archive. license_and_use_terms: CC BY-NC-SA 4.0 with attribution requirements, non-commercial use, Data Access Committee for commercial licensing. KOLF2.1J: simple MTA required for non-profit access. MDA-MB-468: ATCC commercial source
qualityExcellent access documentation with specific platforms, access methods (browser, API, Python library), licensing requirements, and paths for both academic (direct) and commercial (committee oversight) use.
correctnessAll access platforms are real (Dataverse, MassIVE, NCBI SRA, NDEx). CC BY-NC-SA 4.0 is valid license. ATCC and HipSci are real cell line sources
consistencyAccess mechanisms align with distribution formats. License terms consistent with Data Access Committee governance structure
5/5 R20 Q4 (Structural Completeness) File Enumeration and Type Variety
level>3 file types
evidencedistribution_formats: 5 distinct formats - (1) RO-Crate packages with JSON-LD metadata, (2) Mass Spectrometry data in MassIVE, (3) Sequence data in NCBI SRA, (4) Hierarchical cell maps in NDEx, (5) Dataverse archived RO-Crates. Media types: JSON-LD, mass spec raw data, FASTQ/SRA, network graphs, multiple data types in RO-Crate bundles
qualityExceptional format diversity with 5 distinct distribution channels and multiple file types. RO-Crate packaging provides multi-format bundling with provenance.
correctnessAll distribution platforms are real (MassIVE, NCBI SRA, NDEx, Dataverse)
consistencyDistribution formats align with data collection methods (mass spec → MassIVE, sequencing → SRA)
maintainers
maintainers:
- id: cm4ai:maintainer:1
  description: 'CM4AI Consortium: Multidisciplinary consortium managing dataset maintenance including
    University of California San Diego (lead), University of California San Francisco, Stanford University,
    University of Virginia, Yale University, University of Alabama at Birmingham, Simon Fraser University,
    and The Hastings Center. Data Governance Committee led by Jillian Parker (jillianparker@health.ucsd.edu).
    Ethical Review by Vardit Ravitsky (ravitskyv@thehastingscenter.org) and Jean-Christophe Belisle-Pipon
    (jean-christophe_belisle-pipon@sfu.ca).

    '
  maintainer_details:
  - University of California San Diego (lead institution, Ideker Lab)
  - UCSF (Krogan Lab - protein interactions, Sali Lab - structure modeling)
  - Stanford University (Lundberg Lab - spatial proteomics)
  - University of Virginia (Clark Lab - standards and FAIRSCAPE)
  - Yale University (Schulz Lab - workforce development)
  - University of Alabama at Birmingham (Chen Lab - teaming, U-BRITE platform)
  - Simon Fraser University (Bélisle-Pipon - ethics)
  - The Hastings Center (Ravitsky - ethics)
  - Data Governance Committee (Jillian Parker, jillianparker@health.ucsd.edu)
no field-level feedback matched
updates
updates:
  id: cm4ai:updates:1
  description: 'Dataset regularly updated and augmented through end of project in November 2026. Beta
    releases on quarterly basis with periodic data augmentation. Initial alpha release (v0.5) provided
    as supplemental data. March 2025 Beta (V1.4) includes perturb-seq in KOLF2.1J iPSCs, SEC-MS in iPSCs
    and derivatives, and IF images in MDA-MB-468 under three conditions. June 2025 Beta (V2.1) revision
    adds RGB IF images, ro-crate metadata corrections, and naming convention changes, plus SEC-MS for
    MDA-MB-468. October 2025 Beta adds Perturb-seq for MDA-MB-468 breast cancer cells and additional SEC-MS
    data. Future releases will include computed cell maps and complete integration of all data streams.
    Long-term preservation in University of Virginia Dataverse with committed institutional support.

    '
  frequency: Quarterly updates through November 2026; long-term preservation thereafter
  update_details:
  - Alpha release v0.5 (supplemental data)
  - March 2025 Beta release V1.4 (doi:10.18130/V3/B35XWX)
  - June 2025 Beta release V2.1 (doi:10.18130/V3/F3TD5R)
  - October 2025 Beta release (doi:10.18130/V3/K7TGEM)
  - Quarterly augmentation through November 2026
  - Future releases to include computed cell maps
  - Final release expected November 2026
  - Long-term preservation in UVA Dataverse
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidencecollection_mechanisms: 4 detailed entries - (1) IF spatial proteomics with automated protocols at Stanford (Lundberg Lab), (2) AP-MS with endogenous tagging (17 genes done, 34 in progress), (3) SEC-MS at UCSF (Krogan Lab) with specific cell conditions, (4) CRISPR screens with 10x Genomics. acquisition_methods: 4 entries covering confocal microscopy (4 channels specified), mass spectrometry (AP-MS, SEC-MS), single-cell RNA-seq (10x Genomics 3'HT), MuSIC pipeline integration. data_collectors implied via lab attributions (Lundberg Lab Stanford, Krogan Lab UCSF, Mali Lab UCSD). collection_timeframes: Project dates Sept 1 2022 to Aug 31 2026, beta releases March/June/October 2025, quarterly updates through Nov 2026
qualityExcellent collection protocol documentation with comprehensive mechanisms, acquisition methods, lab attributions, and detailed timeframes including release schedule.
correctnessAll lab names and institutions are real. 10x Genomics 3'HT kit is real product. Confocal microscopy channel descriptions are technically accurate
consistencyCollection timelines align with project funding period (2022-2026) and beta release schedule
5/5 R20 Q13 (Technical Documentation) Version History Documentation
levelComprehensive versioning with errata, updates, and release notes
evidenceupdates: Detailed version history with Alpha v0.5 (supplemental data), March 2025 Beta V1.4 (doi:10.18130/V3/B35XWX), June 2025 Beta V2.1 (doi:10.18130/V3/F3TD5R), October 2025 Beta (doi:10.18130/V3/K7TGEM). Release notes: V1.4 adds perturb-seq in iPSCs, SEC-MS in iPSCs, IF images MDA-MB-468. V2.1 adds RGB IF images, ro-crate metadata corrections, naming convention changes, SEC-MS for MDA-MB-468. October adds Perturb-seq for MDA-MB-468. Update frequency: Quarterly through November 2026. version_access: Each version has unique DOI in Dataverse. Future plans: computed cell maps, complete integration
qualityExcellent version history with detailed release notes, version-specific DOIs, errata documentation (metadata corrections in V2.1), and clear update roadmap through 2026.
correctnessAll version DOIs follow Harvard Dataverse pattern 10.18130. Version numbering is consistent (V1.4 → V2.1)
consistencyVersion timeline aligns with project schedule. Incremental data additions match collection mechanisms documented
5/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelStructured version control with timestamps
evidenceupdates: Comprehensive version history with timestamps - Alpha v0.5 (supplemental), March 2025 Beta V1.4, June 2025 Beta V2.1, October 2025 Beta. Each version has unique DOI and detailed change log (V2.1: RGB images added, metadata corrections, naming changes). Provenance: FAIRSCAPE framework computes end-to-end provenance graphs using EVI Evidence Graph Ontology. RO-Crate packages include provenance graphs with machine-readable lineage. retention_limit: NIH data sharing policy compliance, long-term preservation in UVA Dataverse, no planned sunset
qualityExcellent provenance and integrity documentation. Structured version control with DOIs, detailed change logs, formal provenance graphs via FAIRSCAPE/EVI, and long-term preservation commitment.
correctnessFAIRSCAPE and EVI are real provenance frameworks. Version DOIs are valid. NIH data sharing policy is real requirement
consistencyVersion changes align with collection progress (additional data types added in each release). Provenance approach consistent with FAIR principles and RO-Crate packaging
retention_limit
retention_limit:
  id: cm4ai:retention:1
  description: 'Digital data maintained according to NIH data sharing policies with long-term preservation
    in University of Virginia''s LibraData repository supported by committed institutional funds. No planned
    sunset for data availability. Archived RO-Crates with persistent identifiers (ARK, future DOIs) ensure
    long-term accessibility and citability.

    '
  retention_details:
  - NIH data sharing policy compliance
  - UVA Dataverse institutional commitment
  - Persistent identifiers (ARK, future DOIs)
  - No planned data sunset
  - Machine-readable metadata for long-term discoverability
5/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelStructured version control with timestamps
evidenceupdates: Comprehensive version history with timestamps - Alpha v0.5 (supplemental), March 2025 Beta V1.4, June 2025 Beta V2.1, October 2025 Beta. Each version has unique DOI and detailed change log (V2.1: RGB images added, metadata corrections, naming changes). Provenance: FAIRSCAPE framework computes end-to-end provenance graphs using EVI Evidence Graph Ontology. RO-Crate packages include provenance graphs with machine-readable lineage. retention_limit: NIH data sharing policy compliance, long-term preservation in UVA Dataverse, no planned sunset
qualityExcellent provenance and integrity documentation. Structured version control with DOIs, detailed change logs, formal provenance graphs via FAIRSCAPE/EVI, and long-term preservation commitment.
correctnessFAIRSCAPE and EVI are real provenance frameworks. Version DOIs are valid. NIH data sharing policy is real requirement
consistencyVersion changes align with collection progress (additional data types added in each release). Provenance approach consistent with FAIR principles and RO-Crate packaging
human_subject_research
human_subject_research:
  id: cm4ai:hsr:1
  description: 'CM4AI data are distinctive within Bridge2AI in that they are non-clinical data from tissue
    cultures and are considered to be de-identified as they cannot be matched, with current knowledge,
    to a human subject. Both cell lines (MDA-MB-468 and KOLF2.1J) are commercially available, ethically
    sourced, de-identified cell lines. MDA-MB-468 available from ATCC. KOLF2.1J available from HipSci
    resource for non-profit organizations via simple MTA. Human Subjects: No. De-identified Samples: Yes.
    FDA Regulated: No.

    '
  involves_human_subjects: false
  irb_approval:
  - Not applicable - de-identified cell lines from commercial sources
  ethics_review_board:
  - CM4AI Ethics Module (Vardit Ravitsky, Jean-Christophe Bélisle-Pipon)
  - Data Access Committee (Jillian Parker)
  - Bridge2AI Ethics Working Group participation
  special_populations:
  - MDA-MB-468 derived from 51-year-old black female (de-identified)
  - KOLF2.1J derived from healthy male Northern European donor (de-identified)
⚠ low R20 · consistency
issuehuman_subject_research.involves_human_subjects=False is consistent with de-identified cell lines explanation
fieldshuman_subject_research, sensitive_elements
fixNo action needed - logic is sound
4/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelGood ethics documentation with appropriate context
evidencehuman_subject_research: involves_human_subjects=false with detailed explanation (de-identified commercial cell lines, cannot be matched to subjects). irb_approval: 'Not applicable - de-identified cell lines from commercial sources'. ethics_review_board: CM4AI Ethics Module (Vardit Ravitsky, Jean-Christophe Bélisle-Pipon), Data Access Committee (Jillian Parker), Bridge2AI Ethics Working Group. special_populations: MDA-MB-468 from 51-year-old black female (de-identified), KOLF2.1J from healthy male Northern European donor (de-identified). sensitive_elements: Cell line origin metadata documented, sensitive_elements_present=false with rationale
qualityGood ethics documentation appropriate for non-human subjects research. Clear explanation of de-identification and commercial sourcing. Ethics review structure present. Scored 4/5 because comprehensive human subjects protections (IRB, informed consent, compensation) are not applicable to cell line research, but documentation clearly addresses this.
correctnessEthics approach correct for de-identified cell lines - not human subjects research under current regulations
consistencyConsistent logic: involves_human_subjects=false → IRB not applicable → ethics review focuses on data governance and dual use
sensitive_elements
sensitive_elements:
- id: cm4ai:sensitive:1
  description: 'Cell Line Origin Metadata: While cell lines are de-identified and cannot be matched to
    specific individuals, metadata about cell line origins (age, sex, race of original donor) is retained
    for scientific context. This metadata does not constitute identifiable human subjects data under current
    knowledge. Data derived from commercially available de-identified human cell lines and does not represent
    all biological variants in the population at large.

    '
  sensitive_elements_present: false
  sensitivity_details:
  - De-identified commercial cell lines
  - Donor demographic metadata for scientific context only
  - Cannot be matched to individuals with current knowledge
  - Ethically sourced from ATCC and HipSci
  - Does not represent all biological variants seen in the population
⚠ low R20 · consistency
issuehuman_subject_research.involves_human_subjects=False is consistent with de-identified cell lines explanation
fieldshuman_subject_research, sensitive_elements
fixNo action needed - logic is sound
4/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelGood ethics documentation with appropriate context
evidencehuman_subject_research: involves_human_subjects=false with detailed explanation (de-identified commercial cell lines, cannot be matched to subjects). irb_approval: 'Not applicable - de-identified cell lines from commercial sources'. ethics_review_board: CM4AI Ethics Module (Vardit Ravitsky, Jean-Christophe Bélisle-Pipon), Data Access Committee (Jillian Parker), Bridge2AI Ethics Working Group. special_populations: MDA-MB-468 from 51-year-old black female (de-identified), KOLF2.1J from healthy male Northern European donor (de-identified). sensitive_elements: Cell line origin metadata documented, sensitive_elements_present=false with rationale
qualityGood ethics documentation appropriate for non-human subjects research. Clear explanation of de-identification and commercial sourcing. Ethics review structure present. Scored 4/5 because comprehensive human subjects protections (IRB, informed consent, compensation) are not applicable to cell line research, but documentation clearly addresses this.
correctnessEthics approach correct for de-identified cell lines - not human subjects research under current regulations
consistencyConsistent logic: involves_human_subjects=false → IRB not applicable → ethics review focuses on data governance and dual use
external_resources
external_resources:
- id: cm4ai:resource:1
  description: 'CM4AI Project Website: Official project website and data portal using U-BRITE platform'
  external_resources:
  - https://www.cm4ai.org
- id: cm4ai:resource:2
  description: 'NIH RePORTER Project Details: Federal grant information and project details for Bridge2AI
    Functional Genomics'
  external_resources:
  - https://reporter.nih.gov/project-details/11211616
- id: cm4ai:resource:3
  description: 'University of Virginia Dataverse: LibraData repository with archived RO-Crates and data
    releases'
  external_resources:
  - https://doi.org/10.18130/V3/DXWOS5
  - https://doi.org/10.18130/V3/B35XWX
  - https://doi.org/10.18130/V3/F3TD5R
  - https://doi.org/10.18130/V3/K7TGEM
- id: cm4ai:resource:4
  description: 'Nature Publication: Schaffer LV, Hu M, Qian G, et al. Multimodal cell maps as a foundation
    for structural and functional genomics. Nature. Published April 9, 2025.

    '
  external_resources:
  - https://doi.org/10.1038/s41586-025-08878-3
- id: cm4ai:resource:5
  description: 'bioRxiv Preprint: Clark T, et al. Cell Maps for Artificial Intelligence: AI-Ready Maps
    of Human Cell Architecture from Disease-Relevant Cell Lines. BioRXiv, May 2024.

    '
  external_resources:
  - https://doi.org/10.1101/2024.05.21.589311
- id: cm4ai:resource:6
  description: 'FAIRSCAPE Framework Documentation: AI-readiness framework documentation, tutorial, and
    installation instructions'
  external_resources:
  - https://fairscape.github.io
- id: cm4ai:resource:7
  description: 'Integrative Modeling Platform: Open source IMP package for integrative structure modeling'
  external_resources:
  - http://integrativemodeling.org
- id: cm4ai:resource:8
  description: 'Network Data Exchange (NDEx): Repository and visualization platform for cell maps and
    networks'
  external_resources:
  - https://www.ndexbio.org
- id: cm4ai:resource:9
  description: 'Bridge2AI Program: Parent NIH Common Fund program supporting AI-ready biomedical datasets'
  external_resources:
  - https://commonfund.nih.gov/bridge2ai
- id: cm4ai:resource:10
  description: 'NIH Common Fund Data Ecosystem (CFDE): Collaboration partner for data curation and integration'
  external_resources:
  - https://www.nih-cfde.org
- id: cm4ai:resource:11
  description: 'MassIVE Proteomics Repository: Mass spectrometry data repository for iPSC and cancer cell
    SEC-MS data'
  external_resources:
  - MassIVE Repository (SEC-MS human iPSCs)
  - MassIVE Repository (SEC-MS human cancer cells)
- id: cm4ai:resource:12
  description: 'NCBI Sequence Read Archive: Repository for CRISPR perturbation screen raw sequence data'
  external_resources:
  - NCBI BioProject
  - Sequence Read Archive (SRA)
- id: cm4ai:resource:13
  description: 'Perturbation Cell Atlas Publication: Nourreddine S, Doctor Y, Dailamy A, et al. A PERTURBATION
    CELL ATLAS OF HUMAN INDUCED PLURIPOTENT STEM CELLS. bioRxiv. 2024 Nov 4. PMCID: PMC11580897

    '
  external_resources:
  - https://doi.org/10.1101/2024.11.03.621734
5/5 R20 Q14 (Technical Documentation) Associated Publications
levelMultiple references and dataset citation
evidenceexternal_resources: 13 resource categories including Nature publication (Schaffer, Hu et al. April 2025, doi:10.1038/s41586-025-08878-3), bioRxiv preprint (Clark et al. May 2024, doi:10.1101/2024.05.21.589311), Perturbation Cell Atlas publication (Nourreddine et al. 2024, doi:10.1101/2024.11.03.621734). Citation requirements in license_and_use_terms mandate citing Nature publication, bioRxiv preprint, and data collection DOI. Multiple DOIs for dataset versions provided
qualityExcellent publication documentation with 3 peer-reviewed/preprint publications, all with valid DOIs. Dataset citation requirements clearly specified in license terms.
correctnessAll publication DOIs follow valid format (10.1038 for Nature, 10.1101 for bioRxiv). Publication timeline plausible (preprint May 2024 → Nature April 2025)
consistencyPublications align with dataset content (multimodal cell maps, CRISPR perturbation atlas). Authors match creators list
1/1 R20 Q16 (FAIRness & Accessibility) Findability (Persistent Links)
levelPass
evidencepage: https://www.cm4ai.org. external_resources: 13 categories with persistent URLs including https://doi.org/10.18130/V3/DXWOS5 (Dataverse), https://doi.org/10.1038/s41586-025-08878-3 (Nature), https://doi.org/10.1101/2024.05.21.589311 (bioRxiv), https://fairscape.github.io, http://integrativemodeling.org, https://www.ndexbio.org, https://commonfund.nih.gov/bridge2ai, https://reporter.nih.gov/project-details/11211616. ARK identifiers for RO-Crates
qualityPass - Multiple persistent URLs including DOIs, project website, and external platform links. ARK scheme provides additional persistent identifiers.
correctnessAll URLs are well-formed. DOI prefixes valid (10.18130 Dataverse, 10.1038 Nature, 10.1101 bioRxiv)
consistencyURLs align with stated distribution platforms and publication venues
1/1 R20 Q20 (FAIRness & Accessibility) Interlinking Across Platforms
levelPass
evidenceexternal_resources: 13 resource categories linking to multiple platforms - CM4AI website (cm4ai.org), UVA Dataverse (4 version DOIs), NIH RePORTER (grant details), Nature (publication), bioRxiv (preprints), FAIRSCAPE (framework), IMP (software), NDEx (networks), Bridge2AI (program), NIH CFDE (collaboration), MassIVE (mass spec data), NCBI BioProject/SRA (sequences). Cross-platform identifiers: DOIs link Dataverse to citations, ARKs link RO-Crates to FAIRSCAPE, RRIDs link to cell line databases
qualityPass - Extensive cross-platform interlinking with 13 external resource categories spanning repositories, publications, software, and infrastructure. Multiple identifier schemes (DOI, ARK, RRID) enable platform interconnection.
correctnessAll linked platforms are real and operational. Cross-references are semantically appropriate (publication DOIs, repository accessions, software URLs)
consistencyPlatform links align with data types and distribution formats. CFDE collaboration consistent with Bridge2AI program structure

Recommendations

  1. R20 · Replace MassIVE Repository placeholder text with specific accession numbers (MSV000XXXXX format) when data is deposited
  2. R20 · Add direct NCBI BioProject accession numbers (PRJNA format) and SRA run accessions when sequence data is finalized
  3. R20 · Consider adding explicit confidentiality_level field (currently implicit as 'public with usage restrictions')
  4. R20 · Expand discouraged_uses to address potential dual-use concerns beyond clinical misuse, given cancer and neuropsychiatric disease contexts
  5. R20 · Add software repository links (GitHub URLs) for MuSIC pipeline, FAIRSCAPE-CLI, and custom analysis scripts to enhance reproducibility
  6. R20 · Consider documenting LLM tool used for protein assembly naming (preprocessing_strategies mentions 'LLM approach' but doesn't specify which model)
  7. R20 · Add version numbers for additional software tools where possible (node2vec, Human Protein Atlas model) to match detail level of IMP 2.18 documentation