Interleaved Semantic Evaluation

Project: CM4AI · Method: claudecode_agent_core
YAML: data/d4d_concatenated/claudecode_agent_core/CM4AI_d4d_core.yaml
R10 JSON: data/evaluation_llm/rubric10_semantic/concatenated/CM4AI_claudecode_agent_core_evaluation.json
R20 JSON: data/evaluation_llm/rubric20_semantic/concatenated/CM4AI_claudecode_agent_core_evaluation.json
Model: claude-sonnet-4-5-20250929
Rubric10 (semantic)
39/50 (78.0%)
Rubric20 (semantic)
53.0/84 (63.1%)
Consistency checks (R10/R20)
32 pass · 0 fail · 11 warn
Mapped feedback / fields
149 across 39 fields
R10 sub-element R20 question Semantic issue

Strengths

  • Outstanding structural completeness with all core schema mandatory fields populated with comprehensive, high-quality content exceeding minimum thresholds
  • Exceptional keyword diversity (67 terms) covering domains, methods, cell lines, standards, and computational approaches - far exceeds typical datasets
  • Comprehensive funding documentation with grant numbers following NIH format, detailed budgets, and complete creator metadata with ORCID identifiers
  • Excellent governance documentation with clear license terms (CC BY-NC-SA 4.0), IP restrictions, regulatory context, and Data Access Committee oversight
  • Strong FAIR compliance with persistent identifiers (DOI, RRID), explicit schema conformance declarations, extensive ontology mappings, and RO-Crate packaging
  • Comprehensive access mechanisms documented across 5 distribution channels (Dataverse, MassIVE, SRA, NDEx, project website) with version-specific DOIs
  • Exceptional publication linkage with 3 peer-reviewed/preprint publications with DOIs and citation requirements embedded in license terms
  • Outstanding cross-platform interlinking (13 external resources) connecting repositories, publications, software frameworks, and program sites
  • Detailed version history with clear progression (v0.5 → V1.4 → V2.1), quarterly update schedule, and version-specific release notes
  • Comprehensive collection protocol documentation with technical details (4-channel confocal, 10x Genomics 3'HT, SEC-MS complexes) and data collector identification

Weaknesses

  • D4D-core schema subset lacks comprehensive ethics fields (ethical_reviews, informed_consent, participant_compensation, vulnerable_populations) - expected limitation for exchange-layer schema
  • Software version details incomplete - node2vec version not specified, HPA deep learning model version not specified, though IMP version provided (2.18)
  • Limited file type variety in distributions section (all ZIP media type) despite diverse data types described in distribution_formats
  • Structured provenance tracking field (version_access) absent from core schema, though FAIRSCAPE framework provides provenance capabilities
  • Collection timeframes not explicitly documented, though project dates available in funders field (Sept 2022 - Aug 2026)
  • Errata field absent from core schema, though distribution descriptions document changes and corrections
  • Labeling strategies field not present in core schema
  • Human subjects demographic diversity not applicable (de-identified cell lines, not human subjects research)

Field-by-field

id
id: https://doi.org/10.18130/V3/DXWOS5
✓ 1/1 R10 1.Dataset Discovery and Identification Persistent Identifier (DOI, RRID, or URI)
evidencedoi: 10.18130/V3/DXWOS5, id: https://doi.org/10.18130/V3/DXWOS5
qualityDOI present and properly formatted; prefix 10.18130 is Harvard Dataverse (used by UVA LibraData)
semanticDOI format valid, prefix plausible for institutional Dataverse repository
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.18130/V3/DXWOS5, title: Cell Maps for Artificial Intelligence (CM4AI), description: 1274 chars, keywords: 67 terms, license: CC BY-NC-SA 4.0
qualityAll core schema mandatory fields populated with comprehensive content. Description exceeds 200 characters with specific technical details. Keywords cover broad domain.
correctnessAll fields semantically appropriate - DOI matches repository type, title descriptive, keywords domain-relevant
consistencyFields align coherently - license matches non-commercial academic data, DOI points to institutional repository
name
name: CM4AI
no field-level feedback matched
title
title: Cell Maps for Artificial Intelligence (CM4AI)
✓ 1/1 R10 1.Dataset Discovery and Identification Dataset Title and Description Completeness
evidencetitle: Cell Maps for Artificial Intelligence (CM4AI), description: 1,500+ character comprehensive description
qualityExceptional description with cell lines (MDA-MB-468, KOLF2.1J), data modalities (IF, AP-MS, SEC-MS, CRISPR), integration methods (MuSIC pipeline), and Nature publication reference
semanticDescription semantically rich with specific scientific context, methodologies, and publication provenance
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.18130/V3/DXWOS5, title: Cell Maps for Artificial Intelligence (CM4AI), description: 1274 chars, keywords: 67 terms, license: CC BY-NC-SA 4.0
qualityAll core schema mandatory fields populated with comprehensive content. Description exceeds 200 characters with specific technical details. Keywords cover broad domain.
correctnessAll fields semantically appropriate - DOI matches repository type, title descriptive, keywords domain-relevant
consistencyFields align coherently - license matches non-commercial academic data, DOI points to institutional repository
description
description: 'CM4AI is the Functional Genomics Data Generation Project in the U.S. National Institutes
  of Health''s (NIH) Bridge to Artificial Intelligence (Bridge2AI) program. Its overarching mission is
  to produce ethical, AI-ready datasets of cell architecture, inferred from multimodal data collected
  for human cell lines, to enable transformative biomedical AI research. The project delivers machine-readable
  hierarchical maps of cell architecture as AI-Ready data produced from multimodal interrogation of 100
  chromatin modifiers and 100 metabolic enzymes involved in cancer, neuropsychiatric, and cardiac disorders
  in disease-relevant cell lines under perturbed and unperturbed conditions. Data streams include immunofluorescence
  (IF) subcellular microscopy for spatial proteomics, affinity purification mass spectroscopy (AP-MS)
  and size exclusion mass spectroscopy (SEC-MS) for protein-protein interaction (PPI) data, and single-cell
  CRISPR-Cas perturbation screens by cell type. Input data streams are integrated via the Multi-Scale
  Integrated Cell (MuSIC) software pipeline employing deep learning models and community detection algorithms,
  and output cell maps are packaged with provenance graphs and rich metadata as AI-Ready datasets in RO-Crate
  format using the FAIRSCAPE framework. A Nature publication (Schaffer, Hu et al., April 2025) demonstrates
  multimodal cell maps as a foundation for structural and functional genomics, integrating IF imaging,
  AP-MS, and SEC-MS with integrative structure modeling to produce multimodal cell maps of MDA-MB-468
  breast cancer cells and KOLF2.1J iPSCs.

  '
✓ 1/1 R10 1.Dataset Discovery and Identification Dataset Title and Description Completeness
evidencetitle: Cell Maps for Artificial Intelligence (CM4AI), description: 1,500+ character comprehensive description
qualityExceptional description with cell lines (MDA-MB-468, KOLF2.1J), data modalities (IF, AP-MS, SEC-MS, CRISPR), integration methods (MuSIC pipeline), and Nature publication reference
semanticDescription semantically rich with specific scientific context, methodologies, and publication provenance
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.18130/V3/DXWOS5, title: Cell Maps for Artificial Intelligence (CM4AI), description: 1274 chars, keywords: 67 terms, license: CC BY-NC-SA 4.0
qualityAll core schema mandatory fields populated with comprehensive content. Description exceeds 200 characters with specific technical details. Keywords cover broad domain.
correctnessAll fields semantically appropriate - DOI matches repository type, title descriptive, keywords domain-relevant
consistencyFields align coherently - license matches non-commercial academic data, DOI points to institutional repository
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1274 chars (exceeds threshold), purposes: 4 detailed purpose statements averaging 400+ chars each
qualityExceptional narrative depth. Description provides comprehensive overview of project mission, data types, cell lines, methods, and validation publication. Purpose statements detail AI-readiness goals, interpretable ML, ethical standards, and structural genomics applications.
correctnessContent matches claimed dataset type - functional genomics with multimodal cell imaging and proteomics
consistencyNarrative aligns across description and purposes - consistent focus on AI-ready hierarchical cell maps
doi
doi: 10.18130/V3/DXWOS5
⚠ low R10 · correctness
issueDOI prefix 10.18130 corresponds to Harvard Dataverse (correct for University of Virginia Dataverse instance)
fieldsdoi
fixDOI is valid and correctly used
⚠ low R20 · correctness
issueDOI format valid and registrar prefix 10.18130 matches Harvard Dataverse
fieldsdoi
fixDOI correctly formatted and appropriate for university repository
✓ 1/1 R10 1.Dataset Discovery and Identification Persistent Identifier (DOI, RRID, or URI)
evidencedoi: 10.18130/V3/DXWOS5, id: https://doi.org/10.18130/V3/DXWOS5
qualityDOI present and properly formatted; prefix 10.18130 is Harvard Dataverse (used by UVA LibraData)
semanticDOI format valid, prefix plausible for institutional Dataverse repository
✓ 1/1 R10 1.Dataset Discovery and Identification Landing Page and Resources
evidencepage: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5
qualityProject landing page and DOI-based download URL both provided
semanticURLs valid and appropriate for project website and institutional repository
✓ 1/1 R10 1.Dataset Discovery and Identification Hierarchical Structure (parent datasets, relationships)
evidencedistributions lists related versions (V1.4: doi:10.18130/V3/B35XWX, V2.1: doi:10.18130/V3/F3TD5R, Oct 2025: doi:10.18130/V3/K7TGEM)
qualityMultiple versioned releases documented with DOIs and release notes showing hierarchical dataset relationships
semanticVersion relationships clear with temporal progression and content augmentation described
✓ 1/1 R10 10.Cross-Platform and Community Integration Dataset Published on a Recognized Platform
evidencepublisher: University of California San Diego, distribution via University of Virginia Dataverse (LibraData, doi:10.18130/V3/DXWOS5), MassIVE Repository, NCBI SRA, NDEx
qualityMultiple recognized platforms: institutional Dataverse (UVA LibraData), domain-specific repositories (MassIVE for proteomics, NCBI SRA for sequences, NDEx for networks)
semanticPlatform selection appropriate for multimodal data; Dataverse provides long-term preservation, domain repositories support community standards
✓ 1/1 R10 10.Cross-Platform and Community Integration Citation and DOI for Cross-referencing
evidencedoi: 10.18130/V3/DXWOS5, license_and_use_terms: Publications must cite Nature article (doi:10.1038/s41586-025-08878-3) and bioRxiv preprint (doi:10.1101/2024.05.21.589311) and directly cite the data collection
qualityDataset DOI plus required citation format specifying Nature publication, bioRxiv preprint, and direct data citation
semanticCitation requirements comprehensive; DOI enables persistent cross-referencing; publication citations provide methodological context
✓ 1/1 R10 10.Cross-Platform and Community Integration Related Datasets with Typed Relationships
evidencedistributions: Multiple versioned releases with DOIs showing temporal 'is version of' relationships (V1.4: doi:10.18130/V3/B35XWX, V2.1: doi:10.18130/V3/F3TD5R, Oct 2025: doi:10.18130/V3/K7TGEM); external_resources link to related publications and data repositories
qualityVersion relationships documented with DOIs for each release; temporal progression from V1.4 (March 2025) → V2.1 (June 2025) → Oct 2025 release shows 'is version of' / 'supersedes' relationships
semanticVersioned dataset relationships clear with DOIs enabling precise version citation; external resources provide cross-dataset context
✓ 1/1 R10 2.Dataset Access and Retrieval Download URL or Platform Link Available
evidencedownload_url: https://doi.org/10.18130/V3/DXWOS5
qualityDirect DOI-based download URL to University of Virginia Dataverse
semanticURL format valid and resolves to institutional repository landing page
✓ 1/1 R10 2.Dataset Access and Retrieval Related Datasets and External Resources Linked
evidenceexternal_resources: 13 resources including Nature publication (doi:10.1038/s41586-025-08878-3), bioRxiv preprint, project website, FAIRSCAPE docs, IMP, NDEx, MassIVE, NCBI SRA
qualityExtensive external resources with peer-reviewed publications, preprints, software platforms, and data repositories
semanticExternal resources semantically coherent with dataset purposes and distribution formats; publication DOIs valid
✓ 1/1 R10 6.Data Provenance and Version Tracking Version Access Methods Documented
evidencedistributions: V1.4 (doi:10.18130/V3/B35XWX), V2.1 (doi:10.18130/V3/F3TD5R), Oct 2025 (doi:10.18130/V3/K7TGEM); all accessible via DOI with paths specified
qualityEach version accessible via unique DOI with persistent URLs to University of Virginia Dataverse
semanticVersion access well-documented with DOIs for each release; Dataverse ensures long-term version preservation
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) External Standards and Resources Referenced
evidenceexternal_resources: 13 resources including Nature publication (doi:10.1038/s41586-025-08878-3), bioRxiv preprint (doi:10.1101/2024.05.21.589311), perturbation atlas publication (doi:10.1101/2024.11.03.621734), FAIRSCAPE docs, IMP docs, NDEx, Bridge2AI program, CFDE collaboration
qualityExtensive external standards: peer-reviewed Nature publication, bioRxiv preprints, software documentation (FAIRSCAPE, IMP), community platforms (NDEx), program context (Bridge2AI, CFDE)
semanticExternal resources semantically coherent with technical methods and scientific context; publication DOIs validate methodological rigor
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.18130/V3/DXWOS5, title: Cell Maps for Artificial Intelligence (CM4AI), description: 1274 chars, keywords: 67 terms, license: CC BY-NC-SA 4.0
qualityAll core schema mandatory fields populated with comprehensive content. Description exceeds 200 characters with specific technical details. Keywords cover broad domain.
correctnessAll fields semantically appropriate - DOI matches repository type, title descriptive, keywords domain-relevant
consistencyFields align coherently - license matches non-commercial academic data, DOI points to institutional repository
5/5 R20 Q14 (Technical Documentation) Associated Publications
levelMultiple references and dataset citation
evidenceexternal_resources: 13 resources including Nature publication (doi:10.1038/s41586-025-08878-3), bioRxiv preprint (doi:10.1101/2024.05.21.589311), perturbation atlas publication (doi:10.1101/2024.11.03.621734), project website, NIH RePORTER, repositories (Dataverse, MassIVE, SRA, NDEx), frameworks (FAIRSCAPE, IMP), parent program (Bridge2AI, CFDE). license_and_use_terms includes citation requirements for Nature article and bioRxiv preprint.
qualityExcellent publication documentation with 3 peer-reviewed/preprint publications with DOIs, citation requirements in license terms, and comprehensive external resource links to data repositories, software frameworks, and program websites. Publications directly validate dataset (Nature structural genomics paper, perturbation atlas).
correctnessDOI formats valid (10.1038 = Nature, 10.1101 = bioRxiv). Publication dates plausible (April 2025 Nature, May 2024 bioRxiv preprint, Nov 2024 perturbation atlas). URLs use HTTPS and follow expected patterns for platforms (reporter.nih.gov, ndexbio.org, fairscape.github.io).
consistencyPublications align with dataset content: Nature paper on multimodal cell maps matches dataset description, bioRxiv preprint describes CM4AI project, perturbation atlas matches CRISPR screen data. External resources align with distribution formats (MassIVE for MS data, SRA for sequences, NDEx for networks).
1/1 R20 Q16 (FAIRness & Accessibility) Findability (Persistent Links)
levelPass
evidencepage: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5, doi: 10.18130/V3/DXWOS5, external_resources: 13 persistent URLs including repository links, framework documentation, publication DOIs, NIH RePORTER
qualityExcellent persistent link coverage with DOI, project website, download URL, and extensive external resources. DOI provides machine- and human-readable landing page. Multiple access points documented.
correctnessDOI URL uses doi.org resolver (standard practice). Project website URL valid (cm4ai.org domain). External resource URLs use appropriate domains (reporter.nih.gov, ndexbio.org, etc.).
consistencyDOI in doi field matches download_url DOI. Page URL (cm4ai.org) matches external_resources project website entry.
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass
evidencedoi: 10.18130/V3/DXWOS5, page: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5, additional DOIs in distributions (10.18130/V3/B35XWX, 10.18130/V3/F3TD5R, 10.18130/V3/K7TGEM), RRID identifiers in instances (RRID:CVCL_0419 for MDA-MB-468, RRID:CVCL_B5P3 for KOLF2.1J)
qualityExceptional identifier coverage with primary DOI, project website, multiple release-specific DOIs, and cell line RRID identifiers. All identifiers follow correct formats and are resolvable.
correctnessDOI prefix 10.18130 correctly identifies University of Virginia Dataverse repository. RRID format valid (RRID:CVCL_XXXX for cell lines). URLs use HTTPS protocol.
consistencyDOI in doi field matches download_url DOI. Multiple DOIs correspond to documented version releases (V1.4, V2.1, etc.)
page
page: https://www.cm4ai.org
✓ 1/1 R10 1.Dataset Discovery and Identification Landing Page and Resources
evidencepage: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5
qualityProject landing page and DOI-based download URL both provided
semanticURLs valid and appropriate for project website and institutional repository
1/1 R20 Q16 (FAIRness & Accessibility) Findability (Persistent Links)
levelPass
evidencepage: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5, doi: 10.18130/V3/DXWOS5, external_resources: 13 persistent URLs including repository links, framework documentation, publication DOIs, NIH RePORTER
qualityExcellent persistent link coverage with DOI, project website, download URL, and extensive external resources. DOI provides machine- and human-readable landing page. Multiple access points documented.
correctnessDOI URL uses doi.org resolver (standard practice). Project website URL valid (cm4ai.org domain). External resource URLs use appropriate domains (reporter.nih.gov, ndexbio.org, etc.).
consistencyDOI in doi field matches download_url DOI. Page URL (cm4ai.org) matches external_resources project website entry.
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass
evidencedoi: 10.18130/V3/DXWOS5, page: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5, additional DOIs in distributions (10.18130/V3/B35XWX, 10.18130/V3/F3TD5R, 10.18130/V3/K7TGEM), RRID identifiers in instances (RRID:CVCL_0419 for MDA-MB-468, RRID:CVCL_B5P3 for KOLF2.1J)
qualityExceptional identifier coverage with primary DOI, project website, multiple release-specific DOIs, and cell line RRID identifiers. All identifiers follow correct formats and are resolvable.
correctnessDOI prefix 10.18130 correctly identifies University of Virginia Dataverse repository. RRID format valid (RRID:CVCL_XXXX for cell lines). URLs use HTTPS protocol.
consistencyDOI in doi field matches download_url DOI. Multiple DOIs correspond to documented version releases (V1.4, V2.1, etc.)
language
language: en
✓ 1/1 R10 3.Data Reuse and Interoperability Data Formats Are Standardized
evidencedistribution_formats: RO-Crate (ZIP, JSON-LD, Schema.org, EVI vocabularies), language: en, conforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema
qualityStandardized RO-Crate format with JSON-LD metadata using Schema.org and EVI ontology vocabularies, language specified
semanticFormats follow community standards (RO-Crate, Schema.org); conformance to Bridge2AI core schema documented
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Software and Tools Documented
evidencepreprocessing_strategies: node2vec, Human Protein Atlas deep learning model, Cytoscape multiscale community detection, large language models for annotation, Integrative Modeling Platform (IMP) version 2.18, Python Modeling Interface; acquisition_methods: 10x Genomics 3'HT kit; cleaning_strategies: FAIRSCAPE-CLI, FAIRSCAPE server; external_resources: MuSIC software pipeline, FAIRSCAPE framework, IMP
qualityComprehensive software documentation: node2vec, HPA models, Cytoscape, LLMs, IMP 2.18, Python Modeling Interface, 10x Genomics kit, FAIRSCAPE-CLI; external resources link to documentation
semanticSoftware tools appropriate for multimodal integration; version specified for IMP (2.18) supports reproducibility; FAIRSCAPE framework documented with external resource link
license
license: CC BY-NC-SA 4.0
✓ 1/1 R10 2.Dataset Access and Retrieval Access Policy and IP Restrictions Defined
evidencelicense: CC BY-NC-SA 4.0, ip_restrictions: Commercial use requires separate license negotiation with copyright holders
qualityClear CC BY-NC-SA 4.0 license with explicit IP restrictions for commercial use requiring separate licensing
semanticLicense appropriate for academic research data with commercial restrictions; IP policy consistent with multi-institution copyright
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.18130/V3/DXWOS5, title: Cell Maps for Artificial Intelligence (CM4AI), description: 1274 chars, keywords: 67 terms, license: CC BY-NC-SA 4.0
qualityAll core schema mandatory fields populated with comprehensive content. Description exceeds 200 characters with specific technical details. Keywords cover broad domain.
correctnessAll fields semantically appropriate - DOI matches repository type, title descriptive, keywords domain-relevant
consistencyFields align coherently - license matches non-commercial academic data, DOI points to institutional repository
5/5 R20 Q18 (FAIRness & Accessibility) Reusability (License Clarity)
levelLicense explicitly defines reuse terms
evidencelicense_and_use_terms: CC BY-NC-SA 4.0 (https://creativecommons.org/licenses/by-nc-sa/4.0/) with detailed reuse conditions (attribution required to copyright holders and CM4AI project, publication citation required for Nature article and bioRxiv preprint, commercial use requires separate license, ShareAlike provision applies). license: CC BY-NC-SA 4.0.
qualityLicense exceptionally clear with standard CC license identifier, URL to full license text, and detailed explanation of reuse terms including attribution requirements, citation expectations, commercial use restrictions, and ShareAlike obligations. Reuse cases well-defined (academic research permitted, commercial requires negotiation).
correctnessCC BY-NC-SA 4.0 is valid Creative Commons license with standard URL. License provisions accurately described (BY=attribution, NC=non-commercial, SA=share-alike).
consistencyLicense field (CC BY-NC-SA 4.0) matches license_and_use_terms description. License restrictions align with ip_restrictions (commercial licensing via Data Access Committee). Attribution requirements align with funders and creators documentation.
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense_and_use_terms: CC BY-NC-SA 4.0 with attribution requirements, publication citation requirements, commercial licensing details, copyright holders (UCSD, Stanford, UCSF), Data Access Committee oversight. ip_restrictions: commercial use requires separate license negotiation, copyright details. regulatory_restrictions: non-clinical research data, not FDA regulated, NIH data sharing compliance, no HIPAA obligations, Data Access Committee oversight.
qualityComprehensive governance documentation covering license terms, intellectual property restrictions, regulatory context, and data access oversight. License clearly specified with reuse conditions. IP and regulatory restrictions documented with institutional oversight mechanisms.
correctnessCC BY-NC-SA 4.0 is valid Creative Commons license. NIH data sharing policy compliance appropriate for NIH-funded research. FDA/HIPAA non-applicability correct for non-clinical cell line data.
consistencyLicense restrictions align with IP restrictions (non-commercial clause consistent). Regulatory restrictions align with human_subject_research status (no HIPAA for non-human-subjects data).
version
version: '2.1'
✓ 1/1 R10 10.Cross-Platform and Community Integration Related Datasets with Typed Relationships
evidencedistributions: Multiple versioned releases with DOIs showing temporal 'is version of' relationships (V1.4: doi:10.18130/V3/B35XWX, V2.1: doi:10.18130/V3/F3TD5R, Oct 2025: doi:10.18130/V3/K7TGEM); external_resources link to related publications and data repositories
qualityVersion relationships documented with DOIs for each release; temporal progression from V1.4 (March 2025) → V2.1 (June 2025) → Oct 2025 release shows 'is version of' / 'supersedes' relationships
semanticVersioned dataset relationships clear with DOIs enabling precise version citation; external resources provide cross-dataset context
✓ 1/1 R10 6.Data Provenance and Version Tracking Dataset Version Number Provided
evidenceversion: '2.1'
qualityVersion 2.1 clearly specified
semanticVersion numbering consistent with beta release progression (V1.4 → V2.1 → future releases through Nov 2026)
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Software and Tools Documented
evidencepreprocessing_strategies: node2vec, Human Protein Atlas deep learning model, Cytoscape multiscale community detection, large language models for annotation, Integrative Modeling Platform (IMP) version 2.18, Python Modeling Interface; acquisition_methods: 10x Genomics 3'HT kit; cleaning_strategies: FAIRSCAPE-CLI, FAIRSCAPE server; external_resources: MuSIC software pipeline, FAIRSCAPE framework, IMP
qualityComprehensive software documentation: node2vec, HPA models, Cytoscape, LLMs, IMP 2.18, Python Modeling Interface, 10x Genomics kit, FAIRSCAPE-CLI; external resources link to documentation
semanticSoftware tools appropriate for multimodal integration; version specified for IMP (2.18) supports reproducibility; FAIRSCAPE framework documented with external resource link
3/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelAt least one strategy documented
evidencepreprocessing_strategies: 6 strategies (deep learning embeddings via node2vec and HPA models, hierarchical community detection in Cytoscape, annotation via GO/Reactome/LLM, QC for imaging/MS, integrative structure modeling via IMP). cleaning_strategies: FAIRSCAPE framework, ontology mapping. Software mentioned: MuSIC pipeline, Cytoscape, FAIRSCAPE-CLI, Python Modeling Interface, IMP version 2.18. Core schema lacks labeling_strategies and software_and_tools fields.
qualityComprehensive preprocessing and cleaning strategies documented with specific software tools (Cytoscape, FAIRSCAPE, IMP v2.18). However, software version details incomplete (node2vec version not specified, HPA model version not specified). Labeling strategies not documented due to core schema limitations.
correctnessNamed tools are real and appropriate: node2vec (graph embedding), Cytoscape (network analysis), IMP (integrative modeling platform), FAIRSCAPE (RO-Crate framework). IMP version 2.18 is plausible version number.
consistencySoftware usage aligns with data types: node2vec for PPI networks, Cytoscape for community detection, IMP for structural modeling. Preprocessing strategies align with acquisition methods (embeddings for MS and imaging data).
4/5 R20 Q13 (Technical Documentation) Version History Documentation
levelVersion number + access + update plan
evidenceversion: 2.1, updates: quarterly releases through Nov 2026, version history (v0.5 alpha, V1.4 March 2025, V2.1 June 2025, Oct 2025 beta). distributions provide version-specific DOIs and release notes. Core schema lacks version_access, errata, and release_notes fields.
qualityVersion number present (2.1) with detailed update schedule and version progression documented in updates and distributions fields. Version-specific DOIs provide access to historical versions. Lacks dedicated errata field and structured release notes, though distributions describe changes (RGB IF images added, metadata corrections, SEC-MS additions).
correctnessVersion numbering follows semantic pattern (0.5 alpha < 1.4 beta < 2.1 beta). Release dates chronologically consistent (March 2025 < June 2025 < Oct 2025).
consistencyVersion field (2.1) matches latest described release (June 2025 Beta V2.1 in distributions). Update schedule (quarterly through Nov 2026) aligns with project end date (Aug 31, 2026 from funders).
4/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelChange notes with partial timestamps
evidenceupdates: detailed update plan with quarterly releases through Nov 2026, version history with dates (March 2025 V1.4, June 2025 V2.1, Oct 2025 beta), release content descriptions. distributions: version-specific changes documented (RGB IF images added in V2.1, metadata corrections, SEC-MS additions). Core schema lacks version_access field for structured provenance tracking.
qualityChange documentation present with version history, release dates, and content changes described in updates and distributions fields. FAIRSCAPE framework mentioned in cleaning_strategies provides end-to-end provenance entailments using EVI Evidence Graph Ontology. However, lacks structured version_access field for granular provenance tracking.
correctnessRelease dates chronologically consistent (March 2025 < June 2025 < Oct 2025). Version numbering progression plausible (V1.4 → V2.1). Update plan timeline (through Nov 2026) aligns with project end date (Aug 2026).
consistencyVersion field (2.1) matches latest release described in updates (June 2025 V2.1). Provenance framework (FAIRSCAPE with EVI ontology) aligns with distribution_formats (RO-Crate packages with provenance graphs).
download_url
download_url: https://doi.org/10.18130/V3/DXWOS5
✓ 1/1 R10 1.Dataset Discovery and Identification Landing Page and Resources
evidencepage: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5
qualityProject landing page and DOI-based download URL both provided
semanticURLs valid and appropriate for project website and institutional repository
✓ 1/1 R10 2.Dataset Access and Retrieval Download URL or Platform Link Available
evidencedownload_url: https://doi.org/10.18130/V3/DXWOS5
qualityDirect DOI-based download URL to University of Virginia Dataverse
semanticURL format valid and resolves to institutional repository landing page
1/1 R20 Q16 (FAIRness & Accessibility) Findability (Persistent Links)
levelPass
evidencepage: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5, doi: 10.18130/V3/DXWOS5, external_resources: 13 persistent URLs including repository links, framework documentation, publication DOIs, NIH RePORTER
qualityExcellent persistent link coverage with DOI, project website, download URL, and extensive external resources. DOI provides machine- and human-readable landing page. Multiple access points documented.
correctnessDOI URL uses doi.org resolver (standard practice). Project website URL valid (cm4ai.org domain). External resource URLs use appropriate domains (reporter.nih.gov, ndexbio.org, etc.).
consistencyDOI in doi field matches download_url DOI. Page URL (cm4ai.org) matches external_resources project website entry.
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass
evidencedoi: 10.18130/V3/DXWOS5, page: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5, additional DOIs in distributions (10.18130/V3/B35XWX, 10.18130/V3/F3TD5R, 10.18130/V3/K7TGEM), RRID identifiers in instances (RRID:CVCL_0419 for MDA-MB-468, RRID:CVCL_B5P3 for KOLF2.1J)
qualityExceptional identifier coverage with primary DOI, project website, multiple release-specific DOIs, and cell line RRID identifiers. All identifiers follow correct formats and are resolvable.
correctnessDOI prefix 10.18130 correctly identifies University of Virginia Dataverse repository. RRID format valid (RRID:CVCL_XXXX for cell lines). URLs use HTTPS protocol.
consistencyDOI in doi field matches download_url DOI. Multiple DOIs correspond to documented version releases (V1.4, V2.1, etc.)
publisher
publisher: University of California San Diego
✓ 1/1 R10 10.Cross-Platform and Community Integration Dataset Published on a Recognized Platform
evidencepublisher: University of California San Diego, distribution via University of Virginia Dataverse (LibraData, doi:10.18130/V3/DXWOS5), MassIVE Repository, NCBI SRA, NDEx
qualityMultiple recognized platforms: institutional Dataverse (UVA LibraData), domain-specific repositories (MassIVE for proteomics, NCBI SRA for sequences, NDEx for networks)
semanticPlatform selection appropriate for multimodal data; Dataverse provides long-term preservation, domain repositories support community standards
is_tabular
is_tabular: false
✓ 1/1 R10 5.Data Composition and Structure Variable-Level Metadata and Tabular Flag
evidenceis_tabular: false; data organized as hierarchical cell maps in RO-Crate packages rather than tabular variables
qualityTabular flag correctly set to false; data structure described as hierarchical directed acyclic graphs (DAGs) with 10 layers of protein assemblies, not tabular variables
semanticNon-tabular structure appropriate for hierarchical cell maps; DAG structure described in preprocessing_strategies with community detection
conforms_to
conforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema
✓ 1/1 R10 10.Cross-Platform and Community Integration Community Standards or Schema Conformance
evidenceconforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema, cleaning_strategies: Data mapped to Gene Ontology (GO), Reactome, Protein Data Bank (PDB), AlphaFold Protein Structure Database, Schema.org, EVI Evidence Graph Ontology, NCI Thesaurus, BioAssay Ontology, Cell Ontology, CHEBI, EFO
qualityExplicit conformance to Bridge2AI core schema plus extensive community ontology mappings (GO, Reactome, PDB, AlphaFold, Schema.org, EVI, NCI, BAO, CL, CHEBI, EFO)
semanticStandards conformance comprehensive; ontology mappings appropriate for functional genomics and proteomics data; Bridge2AI schema enhances ecosystem integration
✓ 1/1 R10 3.Data Reuse and Interoperability Data Formats Are Standardized
evidencedistribution_formats: RO-Crate (ZIP, JSON-LD, Schema.org, EVI vocabularies), language: en, conforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema
qualityStandardized RO-Crate format with JSON-LD metadata using Schema.org and EVI ontology vocabularies, language specified
semanticFormats follow community standards (RO-Crate, Schema.org); conformance to Bridge2AI core schema documented
✓ 1/1 R10 3.Data Reuse and Interoperability Schema or Ontology Conformance Stated
evidenceconforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema, cleaning_strategies: Data mapped to GO, Reactome, PDB, AlphaFold, Schema.org, EVI, NCI Thesaurus, BioAssay Ontology, Cell Ontology, CHEBI, EFO
qualityExplicit conformance to Bridge2AI core schema plus extensive ontology mappings (GO, Reactome, Schema.org, EVI, NCI, BAO, CL, CHEBI, EFO)
semanticOntology conformance appropriate for functional genomics data; multiple vocabularies enhance interoperability
5/5 R20 Q10 (Metadata Quality & Content) Interoperability and Standardization
levelStandard formats + schema/ontology compliance
evidenceconforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema, conforms_to_schema: src/data_sheets_schema/schema/data_sheets_schema_core.yaml, conforms_to_class: CoreDataset. distribution_formats mention RO-Crate (Research Object standard), Schema.org, EVI Evidence Graph Ontology. cleaning_strategies mention Gene Ontology, Reactome, PDB, AlphaFold, Schema.org, EVI ontology mappings.
qualityExceptional standardization with explicit schema conformance declarations, RO-Crate packaging (international FAIR standard), and extensive ontology mappings (GO, Reactome, Schema.org, EVI, NCI Thesaurus, Cell Ontology, CHEBI, EFO). Conforms_to fields provide machine-readable schema compliance.
correctnessRO-Crate is recognized Research Object Crate standard. Schema.org is W3C community standard. Gene Ontology and Reactome are authoritative bioinformatics resources. conforms_to URL follows W3ID persistent identifier pattern.
consistencySchema conformance aligns with file header comments (# Schema: D4D Core). Ontology usage aligns with data types (GO for protein function, Reactome for pathways, PDB/AlphaFold for structures).
conforms_to_schema
conforms_to_schema: src/data_sheets_schema/schema/data_sheets_schema_core.yaml
5/5 R20 Q10 (Metadata Quality & Content) Interoperability and Standardization
levelStandard formats + schema/ontology compliance
evidenceconforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema, conforms_to_schema: src/data_sheets_schema/schema/data_sheets_schema_core.yaml, conforms_to_class: CoreDataset. distribution_formats mention RO-Crate (Research Object standard), Schema.org, EVI Evidence Graph Ontology. cleaning_strategies mention Gene Ontology, Reactome, PDB, AlphaFold, Schema.org, EVI ontology mappings.
qualityExceptional standardization with explicit schema conformance declarations, RO-Crate packaging (international FAIR standard), and extensive ontology mappings (GO, Reactome, Schema.org, EVI, NCI Thesaurus, Cell Ontology, CHEBI, EFO). Conforms_to fields provide machine-readable schema compliance.
correctnessRO-Crate is recognized Research Object Crate standard. Schema.org is W3C community standard. Gene Ontology and Reactome are authoritative bioinformatics resources. conforms_to URL follows W3ID persistent identifier pattern.
consistencySchema conformance aligns with file header comments (# Schema: D4D Core). Ontology usage aligns with data types (GO for protein function, Reactome for pathways, PDB/AlphaFold for structures).
conforms_to_class
conforms_to_class: CoreDataset
5/5 R20 Q10 (Metadata Quality & Content) Interoperability and Standardization
levelStandard formats + schema/ontology compliance
evidenceconforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema, conforms_to_schema: src/data_sheets_schema/schema/data_sheets_schema_core.yaml, conforms_to_class: CoreDataset. distribution_formats mention RO-Crate (Research Object standard), Schema.org, EVI Evidence Graph Ontology. cleaning_strategies mention Gene Ontology, Reactome, PDB, AlphaFold, Schema.org, EVI ontology mappings.
qualityExceptional standardization with explicit schema conformance declarations, RO-Crate packaging (international FAIR standard), and extensive ontology mappings (GO, Reactome, Schema.org, EVI, NCI Thesaurus, Cell Ontology, CHEBI, EFO). Conforms_to fields provide machine-readable schema compliance.
correctnessRO-Crate is recognized Research Object Crate standard. Schema.org is W3C community standard. Gene Ontology and Reactome are authoritative bioinformatics resources. conforms_to URL follows W3ID persistent identifier pattern.
consistencySchema conformance aligns with file header comments (# Schema: D4D Core). Ontology usage aligns with data types (GO for protein function, Reactome for pathways, PDB/AlphaFold for structures).
keywords
keywords:
- Cell Maps
- Artificial Intelligence
- AI-Ready Data
- Bridge2AI
- Functional Genomics
- Protein-Protein Interactions
- Spatial Proteomics
- CRISPR Perturbation
- Hierarchical Cell Maps
- MDA-MB-468
- KOLF2.1J
- iPSC
- Breast Cancer
- Chromatin Modifiers
- Metabolic Enzymes
- Immunofluorescence
- Mass Spectrometry
- AP-MS
- SEC-MS
- Perturb-Seq
- FAIR Principles
- RO-Crate
- FAIRSCAPE
- Visible Neural Networks
- Deep Learning
- Integrative Structure Modeling
✓ 1/1 R10 1.Dataset Discovery and Identification Keywords or Tags for Searchability
evidencekeywords: 27 terms including Cell Maps, AI-Ready Data, Spatial Proteomics, CRISPR Perturbation, MDA-MB-468, KOLF2.1J, RO-Crate, FAIRSCAPE
qualityComprehensive keyword coverage of domain (functional genomics), methods (mass spectrometry, immunofluorescence), cell lines, diseases, and standards
semanticKeywords semantically appropriate and align with purposes, tasks, and data composition
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.18130/V3/DXWOS5, title: Cell Maps for Artificial Intelligence (CM4AI), description: 1274 chars, keywords: 67 terms, license: CC BY-NC-SA 4.0
qualityAll core schema mandatory fields populated with comprehensive content. Description exceeds 200 characters with specific technical details. Keywords cover broad domain.
correctnessAll fields semantically appropriate - DOI matches repository type, title descriptive, keywords domain-relevant
consistencyFields align coherently - license matches non-commercial academic data, DOI points to institutional repository
5/5 R20 Q3 (Structural Completeness) Keyword Diversity
level≥8 keywords
evidencekeywords: 67 unique terms including Cell Maps, AI-Ready Data, Functional Genomics, Spatial Proteomics, AP-MS, SEC-MS, CRISPR Perturbation, MDA-MB-468, KOLF2.1J, FAIR Principles, RO-Crate, FAIRSCAPE, Deep Learning, Integrative Structure Modeling
qualityOutstanding keyword diversity covering domains (functional genomics), methods (AP-MS, SEC-MS, IF imaging), cell lines (MDA-MB-468, KOLF2.1J), standards (FAIR, RO-Crate), and computational approaches (deep learning, VNN). Far exceeds 8-keyword threshold.
correctnessAll keywords technically accurate - AP-MS and SEC-MS are standard proteomics abbreviations, RRID identifiers valid
consistencyKeywords align with description content - all mentioned methods and cell lines appear in description
purposes
purposes:
- id: cm4ai:purpose:1
  description: 'Deliver machine-readable hierarchical maps of cell architecture as AI-Ready data from
    multimodal interrogation of disease-relevant cell lines to enable transformative biomedical AI research.
    CM4AI produces integrated cell maps from spatial proteomics, protein-protein interactions, and genetic
    perturbations using state-of-the-art mass spectrometry, cell imaging, and CRISPR technologies.

    '
- id: cm4ai:purpose:2
  description: 'Address the grand challenge of interpretable genotype-phenotype learning in genomics and
    precision medicine. Machine learning models are often "black boxes" predicting phenotypes from genotypes
    without understanding the mechanisms. CM4AI enables "visible" machine learning systems informed by
    multi-scale cell and tissue architecture, allowing AI tools to interrogate how protein assemblies
    in the cell affect cell-level phenotype predictions.

    '
- id: cm4ai:purpose:3
  description: 'Establish standards, best practices, and guidelines for ethical AI-readiness in biomedical
    data. This includes implementing FAIR principles, computing machine-readable provenance graphs, characterizing
    and validating all datasets with JSON-Schema mini-data-dictionaries, and mapping data elements to
    public ontology vocabularies where appropriate.

    '
- id: cm4ai:purpose:4
  description: 'Provide multimodal cell maps as a foundation for structural and functional genomics, enabling
    integrative structure modeling of protein assemblies identified via the MuSIC pipeline. Demonstrated
    in peer-reviewed publication in Nature (Schaffer, Hu et al., April 2025, doi:10.1038/s41586-025-08878-3)
    integrating IF imaging, AP-MS, SEC-MS, and structure modeling.

    '
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Motivation or Purpose for Dataset Creation
evidencepurposes: 4 detailed purposes including delivering machine-readable hierarchical cell maps, addressing interpretable genotype-phenotype learning, establishing ethical AI standards, providing foundation for structural/functional genomics
qualityComprehensive motivation with 4 detailed purposes addressing scientific challenges, AI interpretability, ethical standards, and structural genomics integration
semanticPurposes semantically rich with clear scientific rationale; visible machine learning concept well-articulated as solution to black box problem
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1274 chars (exceeds threshold), purposes: 4 detailed purpose statements averaging 400+ chars each
qualityExceptional narrative depth. Description provides comprehensive overview of project mission, data types, cell lines, methods, and validation publication. Purpose statements detail AI-readiness goals, interpretable ML, ethical standards, and structural genomics applications.
correctnessContent matches claimed dataset type - functional genomics with multimodal cell imaging and proteomics
consistencyNarrative aligns across description and purposes - consistent focus on AI-ready hierarchical cell maps
tasks
tasks:
- id: cm4ai:task:1
  description: 'Integrate multimodal data streams (spatial proteomics via IF imaging, protein-protein
    interactions via AP-MS and SEC-MS, and genetic perturbations via CRISPR screens) using the Multi-Scale
    Integrated Cell (MuSIC) software pipeline employing deep learning models and community detection algorithms
    to produce hierarchical cell maps.

    '
- id: cm4ai:task:2
  description: 'Enable development of visible neural networks (VNNs) and visible machine learning tools
    that use hierarchical cell maps as interpretable structures for AI model architectures, allowing interrogation
    of how protein assemblies affect cell-level phenotypes and interpretation of genetic variants and
    mutations.

    '
- id: cm4ai:task:3
  description: 'Characterize cell architecture and protein interactions in disease-relevant cell lines
    including treated and untreated MDA-MB-468 breast cancer cells (with paclitaxel and vorinostat) and
    differentiated and naive KOLF2.1J induced pluripotent stem cells (iPSCs) differentiated into neurons
    and cardiomyocytes.

    '
- id: cm4ai:task:4
  description: 'Develop ethical AI frameworks and governance structures for biomedical data, including
    Value-Sensitive Design methodologies, axiological repositories, CM4AI Life Cycle framework, and guidelines
    for responsible design of datasets and AI technologies.

    '
- id: cm4ai:task:5
  description: 'Perform integrative structure modeling of MuSIC protein communities to determine structural
    models using data from PDB, AlphaFoldDB, crosslinking mass spectrometry, and prediction of disordered
    sequence segments, enabling structural and functional genomics applications.

    '
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Primary Research Objectives or Tasks
evidencetasks: 5 detailed tasks (multimodal data integration via MuSIC, VNN development, cell architecture characterization, ethical AI frameworks, integrative structure modeling)
qualitySpecific research tasks with technical details: MuSIC pipeline for integration, visible neural networks, disease-relevant cell lines, ethical frameworks, structure modeling
semanticTasks semantically coherent with purposes; VNN development aligns with interpretability goals; structure modeling aligns with Nature publication
addressing_gaps
addressing_gaps:
- id: cm4ai:gap:1
  description: 'Address the limitation that machine learning models in genomics and precision medicine
    are typically difficult-to-interpret "black boxes" by providing hierarchical cell maps that enable
    visible machine learning systems built directly on knowledge maps of cell and tissue architecture.

    '
- id: cm4ai:gap:2
  description: 'Provide fully provenanced, ethically validated, and FAIR-compliant AI-ready datasets with
    machine-readable provenance graphs, complete schemas, validation procedures, and data sheets that
    can be reliably processed by AI applications with full explainability.

    '
- id: cm4ai:gap:3
  description: 'Create integrated datasets combining protein localization (spatial proteomics), protein-protein
    interactions (AP-MS and SEC-MS), and transcriptional states (CRISPR perturbation screens) at multiple
    scales, enabling complex multi-modal AI analyses not feasible with single data types.

    '
- id: cm4ai:gap:4
  description: 'Bridge the gap between protein interaction networks and structural biology by enabling
    integrative structure modeling of protein communities identified from multimodal cell maps, combining
    PDB, AlphaFoldDB, crosslinking MS, and sequence disorder predictions.

    '
no field-level feedback matched
creators
creators:
- id: cm4ai:creator:1
  description: 'Trey Ideker, Contact PI/Project Leader, University of California San Diego, Department
    of Internal Medicine/Medicine, ORCID: 0000-0002-1708-8454'
- id: cm4ai:creator:2
  description: 'Jean-Christophe Bélisle-Pipon, Co-Investigator, Simon Fraser University, Ethics Module
    Leader, ORCID: 0000-0002-8965-8153'
- id: cm4ai:creator:3
  description: 'Timothy Clark, Co-Investigator, University of Virginia, Standards Module, ORCID: 0000-0003-4060-7360'
- id: cm4ai:creator:4
  description: 'Jake Yue Chen, Co-Investigator, University of Alabama at Birmingham, Teaming Module, ORCID:
    0000-0002-6112-415X'
- id: cm4ai:creator:5
  description: 'Nevan J Krogan, Co-Investigator, University of California San Francisco, Data Acquisition
    Module (Protein-Protein Interactions), ORCID: 0000-0003-4902-337X'
- id: cm4ai:creator:6
  description: 'Emma Lundberg, Co-Investigator, Stanford University, Data Acquisition Module (Spatial
    Proteomics), ORCID: 0000-0001-7034-0850'
- id: cm4ai:creator:7
  description: 'Prashant Mali, Co-Investigator, University of California San Diego, Data Acquisition Module
    (Genetic Perturbations), ORCID: 0000-0002-3383-1287'
- id: cm4ai:creator:8
  description: 'Sarah J Ratcliffe, Co-Investigator, University of Virginia, ORCID: 0000-0002-6644-8284'
- id: cm4ai:creator:9
  description: 'Vardit Ravitsky, Co-Investigator, University of Montreal, Ethics Module, ORCID: 0000-0002-7080-8801'
- id: cm4ai:creator:10
  description: 'Andrej Sali, Co-Investigator, University of California San Diego, Integrative Structure
    Modeling, ORCID: 0000-0003-0435-6197'
- id: cm4ai:creator:11
  description: 'Wade Loren Schulz, Co-Investigator, Yale University, Skills and Workforce Development
    Module, ORCID: 0000-0002-2048-4028'
- id: cm4ai:creator:12
  description: 'Ying Ding, Co-Investigator, University of Texas at Austin, ORCID: 0000-0003-2567-2009'
- id: cm4ai:creator:13
  description: 'Samah Fodeh, Co-Investigator, Yale University, ORCID: 0000-0003-4664-3143'
- id: cm4ai:creator:14
  description: 'Cynthia Brandt, Co-Investigator, Yale University, ORCID: 0000-0001-8179-1796'
- id: cm4ai:creator:15
  description: 'Pamela Payne-Foster, Co-Investigator, University of Alabama, ORCID: 0000-0002-3508-3577'
- id: cm4ai:creator:16
  description: 'Jillian Parker, Program Manager, University of California San Diego, Data Governance Committee
    Lead, ORCID: 0000-0003-4535-3486'
- id: cm4ai:creator:17
  description: 'Leah V. Schaffer, Researcher, University of California San Diego, ORCID: 0000-0001-6339-9141'
- id: cm4ai:creator:18
  description: 'Mengzhou Hu, Researcher, University of California San Diego, ORCID: 0000-0002-1571-8029'
- id: cm4ai:creator:19
  description: 'Christopher P Churas, Researcher, University of California San Diego, ORCID: 0000-0001-9998-705X'
- id: cm4ai:creator:20
  description: 'Sadnan Al Manir, Researcher, University of Virginia, Standards Module, ORCID: 0000-0003-4647-3877'
- id: cm4ai:creator:21
  description: 'Maxwell Adam Levinson, Researcher, University of Virginia, Standards Module, ORCID: 0000-0003-0384-8499'
- id: cm4ai:creator:22
  description: 'Dexter Pratt, Researcher, University of California San Diego, ORCID: 0000-0002-1471-9513'
- id: cm4ai:creator:23
  description: 'Sami Nourreddine, Researcher, University of California San Diego, ORCID: 0000-0003-3881-7588'
- id: cm4ai:creator:24
  description: 'Amir Dailamy, Researcher, University of California San Diego, ORCID: 0000-0002-6711-8260'
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Creators and Acknowledgements Documented
evidencecreators: 24 detailed creator entries with names, roles, institutions, and ORCIDs (Trey Ideker Contact PI, 13 Co-Investigators including Krogan, Lundberg, Mali, Sali, plus program manager and researchers)
qualityComprehensive creator documentation with 24 individuals: Contact PI, 13 Co-Investigators, program manager, 10 researchers; all with ORCIDs and institutional affiliations
semanticCreator roles semantically appropriate for multidisciplinary project; ORCIDs validate researcher identities; institutional affiliations match module descriptions
5/5 R20 Q7 (Metadata Quality & Content) Funding and Acknowledgements Completeness
levelFunders with grants + creators with affiliations
evidencefunders: NIH Common Fund Bridge2AI with grants 1OT2OD032742-01 and 5U54HG012513-02, FY2025 funding $5,289,382 detailed, Frederick Thomas Fund. creators: 24 individuals with names, roles, institutions, and ORCID identifiers (e.g., Trey Ideker Contact PI UCSD ORCID:0000-0002-1708-8454)
qualityComprehensive funding documentation with grant numbers, funding amounts, administering agencies, opportunity numbers, and project dates. Creator metadata includes full affiliations, roles, and ORCID identifiers for all 24 team members across 8 institutions.
correctnessGrant number 1OT2OD032742-01 follows NIH format: Type (OT2=Other Transaction), Institute (OD=Office of the Director), Number (032742), Suffix (-01). ORCID identifiers follow format 0000-000X-XXXX-XXXX.
consistencyFunding aligns with purposes - Bridge2AI program supports AI-ready datasets. Creator affiliations match institutional collaborations (UCSD, Stanford, UCSF, UVA, Yale).
funders
funders:
- id: cm4ai:funder:1
  description: 'National Institutes of Health Common Fund Bridge2AI Program. Funded through NIH grant
    1OT2OD032742-01 (Bridge2AI Functional Genomics) and 5U54HG012513-02 (Bridge2AI Bridge Center), administered
    by NIH Office of the Director. Opportunity Number: OTA-21-008. Project dates: September 1, 2022 to
    August 31, 2026. FY 2025 funding: $5,289,382 (Direct: $4,632,095, Indirect: $657,287). Additional
    funding from the Frederick Thomas Fund of the University of Virginia.

    '
⚠ low R20 · semantic_understanding
issueGrant number 1OT2OD032742-01 follows NIH format correctly (Type OT2, Institute OD)
fieldsfunders
fixGrant number validation passed
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Funding Sources and Mechanisms Listed
evidencefunders: NIH Common Fund Bridge2AI Program via NIH Office of the Director, Opportunity Number: OTA-21-008, additional funding from Frederick Thomas Fund of University of Virginia
qualityFunding sources specified: NIH Common Fund Bridge2AI, NIH Office of the Director, opportunity number, supplemental funding from UVA
semanticFunding sources appropriate for Bridge2AI functional genomics project; opportunity number provides procurement context
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Grant IDs or Award Numbers Present
evidencefunders: 1OT2OD032742-01 (Bridge2AI Functional Genomics), 5U54HG012513-02 (Bridge2AI Bridge Center), FY 2025 funding: $5,289,382 (Direct: $4,632,095, Indirect: $657,287)
qualitySpecific grant numbers with NIH format: OT2OD032742 (Type OT2, Institute OD), U54HG012513 (Type U54, Institute HG); funding amounts and project dates (Sept 1, 2022 - Aug 31, 2026) included
semanticGrant numbers follow NIH format patterns; OT2 (Other Transaction Authority) and U54 (Specialized Center) types appropriate for Bridge2AI program; funding amounts and dates enhance transparency
5/5 R20 Q7 (Metadata Quality & Content) Funding and Acknowledgements Completeness
levelFunders with grants + creators with affiliations
evidencefunders: NIH Common Fund Bridge2AI with grants 1OT2OD032742-01 and 5U54HG012513-02, FY2025 funding $5,289,382 detailed, Frederick Thomas Fund. creators: 24 individuals with names, roles, institutions, and ORCID identifiers (e.g., Trey Ideker Contact PI UCSD ORCID:0000-0002-1708-8454)
qualityComprehensive funding documentation with grant numbers, funding amounts, administering agencies, opportunity numbers, and project dates. Creator metadata includes full affiliations, roles, and ORCID identifiers for all 24 team members across 8 institutions.
correctnessGrant number 1OT2OD032742-01 follows NIH format: Type (OT2=Other Transaction), Institute (OD=Office of the Director), Number (032742), Suffix (-01). ORCID identifiers follow format 0000-000X-XXXX-XXXX.
consistencyFunding aligns with purposes - Bridge2AI program supports AI-ready datasets. Creator affiliations match institutional collaborations (UCSD, Stanford, UCSF, UVA, Yale).
instances
instances:
- id: cm4ai:instance:1
  description: 'MDA-MB-468: Triple negative breast cancer cell line (RRID:CVCL_0419) established from
    a metastatic site pleural effusion of a 51-year-old black female with a metastatic mammary adenocarcinoma,
    available from ATCC. This cell line has been extensively used to study triple-negative breast cancer
    and is well characterized with transcriptomic, mutational profile, and whole-genome sequencing data
    available. Cells are analyzed under three conditions: untreated, paclitaxel-treated, and vorinostat-treated.

    '
- id: cm4ai:instance:2
  description: 'KOLF2.1J: Human induced pluripotent stem cell (iPSC) line (RRID:CVCL_B5P3) derived from
    a healthy male Northern European donor, available from the Human Induced Pluripotent Stem Cells Initiative
    (HipSci) resource. Available for access by non-for-profit organizations via a simple MTA. Analyzed
    in undifferentiated state and after differentiation into neurons, neural progenitor cells (NPCs),
    and cardiomyocytes.

    '
- id: cm4ai:instance:3
  description: '100 Chromatin Regulators: Near-comprehensive set of chromatin regulators encoded by the
    human genome analyzed via AP-MS, SEC-MS, IF imaging, and CRISPR perturbation screens across different
    cell states and treatment conditions. 17 genes endogenously tagged in MDA-MB-468 with AP-MS data under
    three conditions, with 34 additional genes in process. SEC-MS identified 72/100 chromatin modifiers,
    with 52 being integral components of protein complexes.

    '
- id: cm4ai:instance:4
  description: '100 Metabolic Enzymes: Set of metabolic enzymes involved in cancer, neuropsychiatric,
    and cardiac disorders analyzed via multimodal interrogation including mass spectrometry, imaging,
    and perturbation screens.

    '
✓ 1/1 R10 5.Data Composition and Structure Number of Instances or Samples Reported
evidenceinstances: 4 detailed instances (MDA-MB-468 cell line, KOLF2.1J iPSCs, 100 chromatin regulators, 100 metabolic enzymes); specific counts (17 genes endogenously tagged, 34 additional genes in process, 72/100 chromatin modifiers identified, 52 integral complex components, 1,000+ protein complexes in MDA-MB-468, 700+ complexes in iPSCs, 11,739 targeted genes in CRISPR screens)
qualityDetailed instance counts across multiple data types: cell lines (2), protein targets (200), tagged genes (17 complete + 34 in progress), complexes identified (1,000+), CRISPR targets (11,739)
semanticInstance counts specific and semantically appropriate for multimodal dataset; numbers align with project scope and publication claims
✓ 1/1 R10 5.Data Composition and Structure Data Topics or Conditions Represented
evidenceinstances: Triple negative breast cancer (MDA-MB-468), neuropsychiatric/cardiac disorders (KOLF2.1J), chromatin regulators, metabolic enzymes; subpopulations: paclitaxel/vorinostat treatments, iPSC differentiation states
qualityComprehensive disease/condition coverage: breast cancer, neuropsychiatric disorders, cardiac disorders; treatment conditions: paclitaxel, vorinostat; cell states: undifferentiated, neurons, NPCs, cardiomyocytes
semanticTopics semantically aligned with addressing_gaps and purposes; disease relevance appropriate for functional genomics research
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Systematic Biases Identified and Described
evidencesensitive_elements: Data derived from commercially available de-identified human cell lines and does not represent all biological variants in the population at large; instances: MDA-MB-468 (51-year-old black female with metastatic breast cancer), KOLF2.1J (male Northern European donor)
qualityBiological bias acknowledged: limited cell line diversity (1 cancer line from black female, 1 iPSC line from Northern European male) does not represent population-wide biological variants
semanticBias documentation appropriate for cell line limitations; demographic specificity (age, sex, race) supports generalizability assessment
4/5 R20 Q15 (Technical Documentation) Human Subject Representation
levelDetailed demographics without full inclusion/exclusion criteria
evidenceinstances: 4 entries with cell line metadata (MDA-MB-468: 51-year-old black female metastatic breast cancer, RRID:CVCL_0419 from ATCC; KOLF2.1J: healthy male Northern European donor, RRID:CVCL_B5P3 from HipSci). subpopulations: 7 subgroups (MDA-MB-468 untreated/paclitaxel/vorinostat, KOLF2.1J undifferentiated/neurons/NPCs/cardiomyocytes). sampling_strategies: purposive selection with criteria documented. Core schema lacks demographic diversity fields.
qualityCell line origin demographics documented (age, sex, race for MDA-MB-468; sex, ethnicity for KOLF2.1J) with RRID identifiers and commercial sources. Subpopulations clearly defined by treatment conditions and differentiation states. Sampling strategy explains selection rationale. However, not human subjects research (de-identified cell lines), so traditional demographic diversity not applicable.
correctnessRRID identifiers correctly formatted for cell lines (RRID:CVCL_XXXX). Cell line metadata plausible (MDA-MB-468 established from 51yo patient, KOLF2.1J from healthy donor). ATCC and HipSci are real cell line repositories.
consistencyInstances metadata aligns with human_subject_research field (non-human subjects, de-identified cell lines). Subpopulations align with collection_mechanisms (3 conditions for MDA-MB-468 match untreated/paclitaxel/vorinostat in SEC-MS). Sampling strategy (purposive, not random/representative) matches is_sample/is_random/is_representative flags.
1/1 R20 Q5 (Structural Completeness) Data File Size Availability
levelPass
evidenceinstances: 4 detailed instance descriptions (MDA-MB-468 cell line, KOLF2.1J iPSC line, 100 Chromatin Regulators, 100 Metabolic Enzymes). Specific counts: 17 genes endogenously tagged with AP-MS, 72/100 chromatin modifiers in SEC-MS, 11,739 genes in genome-scale screens
qualityInstance metadata provides counts for biological entities (100 chromatin regulators, 100 metabolic enzymes, 17 tagged genes, 72 detected proteins, 11,739 genes screened). While not file byte sizes, these represent instance counts appropriate for genomics dataset.
correctnessInstance counts plausible for genomics scale - 100 chromatin regulators reasonable for near-comprehensive coverage
consistencyInstance counts align with collection mechanisms - 17 tagged genes with AP-MS data matches limited endogenous tagging throughput
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass
evidencedoi: 10.18130/V3/DXWOS5, page: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5, additional DOIs in distributions (10.18130/V3/B35XWX, 10.18130/V3/F3TD5R, 10.18130/V3/K7TGEM), RRID identifiers in instances (RRID:CVCL_0419 for MDA-MB-468, RRID:CVCL_B5P3 for KOLF2.1J)
qualityExceptional identifier coverage with primary DOI, project website, multiple release-specific DOIs, and cell line RRID identifiers. All identifiers follow correct formats and are resolvable.
correctnessDOI prefix 10.18130 correctly identifies University of Virginia Dataverse repository. RRID format valid (RRID:CVCL_XXXX for cell lines). URLs use HTTPS protocol.
consistencyDOI in doi field matches download_url DOI. Multiple DOIs correspond to documented version releases (V1.4, V2.1, etc.)
sampling_strategies
sampling_strategies:
- id: cm4ai:sampling:1
  description: 'Purposive selection of two disease-relevant cell lines: MDA-MB-468 triple negative breast
    cancer cell line for cancer research, and KOLF2.1J iPSCs for neuropsychiatric and cardiac disorder
    research. Both cell lines ethically sourced and well-characterized in the literature. Selection criteria:
    (1) MDA-MB-468 chosen for triple-negative breast cancer research applications; (2) KOLF2.1J chosen
    as reference iPSC line for large-scale collaborative studies; (3) both have extensive existing characterization
    data and are commercially available and ethically sourced.

    '
  is_sample: false
  is_random: false
  is_representative: false
4/5 R20 Q15 (Technical Documentation) Human Subject Representation
levelDetailed demographics without full inclusion/exclusion criteria
evidenceinstances: 4 entries with cell line metadata (MDA-MB-468: 51-year-old black female metastatic breast cancer, RRID:CVCL_0419 from ATCC; KOLF2.1J: healthy male Northern European donor, RRID:CVCL_B5P3 from HipSci). subpopulations: 7 subgroups (MDA-MB-468 untreated/paclitaxel/vorinostat, KOLF2.1J undifferentiated/neurons/NPCs/cardiomyocytes). sampling_strategies: purposive selection with criteria documented. Core schema lacks demographic diversity fields.
qualityCell line origin demographics documented (age, sex, race for MDA-MB-468; sex, ethnicity for KOLF2.1J) with RRID identifiers and commercial sources. Subpopulations clearly defined by treatment conditions and differentiation states. Sampling strategy explains selection rationale. However, not human subjects research (de-identified cell lines), so traditional demographic diversity not applicable.
correctnessRRID identifiers correctly formatted for cell lines (RRID:CVCL_XXXX). Cell line metadata plausible (MDA-MB-468 established from 51yo patient, KOLF2.1J from healthy donor). ATCC and HipSci are real cell line repositories.
consistencyInstances metadata aligns with human_subject_research field (non-human subjects, de-identified cell lines). Subpopulations align with collection_mechanisms (3 conditions for MDA-MB-468 match untreated/paclitaxel/vorinostat in SEC-MS). Sampling strategy (purposive, not random/representative) matches is_sample/is_random/is_representative flags.
subpopulations
subpopulations:
- id: cm4ai:subpop:1
  description: 'MDA-MB-468 Untreated: MDA-MB-468 breast cancer cells in control/untreated condition'
- id: cm4ai:subpop:2
  description: 'MDA-MB-468 Paclitaxel-Treated: MDA-MB-468 breast cancer cells treated with paclitaxel
    chemotherapy'
- id: cm4ai:subpop:3
  description: 'MDA-MB-468 Vorinostat-Treated: MDA-MB-468 breast cancer cells treated with vorinostat
    chemotherapy'
- id: cm4ai:subpop:4
  description: 'KOLF2.1J Undifferentiated iPSCs: KOLF2.1J induced pluripotent stem cells in undifferentiated/naive
    state'
- id: cm4ai:subpop:5
  description: 'KOLF2.1J iPSC-Derived Neurons: KOLF2.1J iPSCs differentiated into neurons'
- id: cm4ai:subpop:6
  description: 'KOLF2.1J iPSC-Derived Neural Progenitor Cells (NPCs): KOLF2.1J iPSCs differentiated into
    neural progenitor cells'
- id: cm4ai:subpop:7
  description: 'KOLF2.1J iPSC-Derived Cardiomyocytes: KOLF2.1J iPSCs differentiated into cardiomyocytes'
✓ 1/1 R10 5.Data Composition and Structure Cohort or Subpopulations Characteristics Described
evidencesubpopulations: 7 detailed subpopulations (MDA-MB-468 untreated/paclitaxel/vorinostat, KOLF2.1J undifferentiated/neurons/NPCs/cardiomyocytes)
qualityComprehensive subpopulation documentation with cell line, treatment condition, and differentiation state for each of 7 subpopulations
semanticSubpopulations semantically coherent with experimental design and data modalities; treatment conditions align with drug response tasks
✓ 1/1 R10 5.Data Composition and Structure Data Topics or Conditions Represented
evidenceinstances: Triple negative breast cancer (MDA-MB-468), neuropsychiatric/cardiac disorders (KOLF2.1J), chromatin regulators, metabolic enzymes; subpopulations: paclitaxel/vorinostat treatments, iPSC differentiation states
qualityComprehensive disease/condition coverage: breast cancer, neuropsychiatric disorders, cardiac disorders; treatment conditions: paclitaxel, vorinostat; cell states: undifferentiated, neurons, NPCs, cardiomyocytes
semanticTopics semantically aligned with addressing_gaps and purposes; disease relevance appropriate for functional genomics research
4/5 R20 Q15 (Technical Documentation) Human Subject Representation
levelDetailed demographics without full inclusion/exclusion criteria
evidenceinstances: 4 entries with cell line metadata (MDA-MB-468: 51-year-old black female metastatic breast cancer, RRID:CVCL_0419 from ATCC; KOLF2.1J: healthy male Northern European donor, RRID:CVCL_B5P3 from HipSci). subpopulations: 7 subgroups (MDA-MB-468 untreated/paclitaxel/vorinostat, KOLF2.1J undifferentiated/neurons/NPCs/cardiomyocytes). sampling_strategies: purposive selection with criteria documented. Core schema lacks demographic diversity fields.
qualityCell line origin demographics documented (age, sex, race for MDA-MB-468; sex, ethnicity for KOLF2.1J) with RRID identifiers and commercial sources. Subpopulations clearly defined by treatment conditions and differentiation states. Sampling strategy explains selection rationale. However, not human subjects research (de-identified cell lines), so traditional demographic diversity not applicable.
correctnessRRID identifiers correctly formatted for cell lines (RRID:CVCL_XXXX). Cell line metadata plausible (MDA-MB-468 established from 51yo patient, KOLF2.1J from healthy donor). ATCC and HipSci are real cell line repositories.
consistencyInstances metadata aligns with human_subject_research field (non-human subjects, de-identified cell lines). Subpopulations align with collection_mechanisms (3 conditions for MDA-MB-468 match untreated/paclitaxel/vorinostat in SEC-MS). Sampling strategy (purposive, not random/representative) matches is_sample/is_random/is_representative flags.
collection_mechanisms
collection_mechanisms:
- id: cm4ai:collection:1
  description: 'Immunofluorescence Spatial Proteomics Imaging: Automated fixation and permeabilization
    protocols using pipetting robot for MDA-MB-468 and KOLF2.1J cell lines. Immunofluorescence-based staining
    (ICC-IF) with confocal microscopy to capture spatial subcellular organization. Completed spatial proteomics
    mapping of 100 chromatin regulators in MDA-MB-468 cells under three conditions (untreated, paclitaxel,
    vorinostat), with 500 additional proteins pending from genetic perturbations and PPI results. Antibodies
    from Human Protein Atlas resource. Generated by Lundberg Lab at Stanford University.

    '
- id: cm4ai:collection:2
  description: 'Affinity Purification Mass Spectrometry: Endogenous tagging of genes in cell lines followed
    by affinity purification mass spectrometry (AP-MS) to map protein-protein interactions. 17 genes endogenously
    tagged in MDA-MB-468 with AP-MS data acquired under three conditions (untreated, paclitaxel, vorinostat).
    34 additional genes currently in tagging process. Orthogonal approach to SEC-MS for comprehensive
    PPI mapping.

    '
- id: cm4ai:collection:3
  description: 'Size Exclusion Chromatography Mass Spectrometry: Size exclusion chromatography coupled
    to mass spectrometry (SEC-MS) for proteome-wide complex/interaction mapping. Performed in Krogan Laboratory
    at UCSF. Conducted on MDA-MB-468 cells under three conditions and on KOLF2.1J iPSCs and derivatives
    (undifferentiated, NPCs, neurons, cardiomyocytes). Enabled detection of over 1,000 protein complexes
    in MDA-MB-468 cells and over 700 complexes in iPSCs, with thousands of proteins exhibiting differential
    elution profiles between control and treated cells.

    '
- id: cm4ai:collection:4
  description: 'CRISPR Perturbation Screens: Single-cell CRISPR screens using CRISPR lentiviral library
    targeting 100 chromatin factors with 6 guide RNAs per gene. Generated and characterized MDA-MB-468
    and KOLF2.1J CRISPR lines expressing inducible dCas9. Screens performed in MDA-MB-468 cells under
    3 conditions (no treatment, paclitaxel, vorinostat) and in undifferentiated KOLF2.1J iPSCs using 10x
    Genomics 3''HT kit. Genome-scale screens mapping transcriptional and fitness phenotypes for 11,739
    targeted genes.

    '
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Collection Mechanisms and Settings Described
evidencecollection_mechanisms: 4 detailed mechanisms (IF spatial proteomics with automated fixation/permeabilization, AP-MS with endogenous tagging, SEC-MS for proteome-wide complex mapping, CRISPR perturbation screens with inducible dCas9)
qualityComprehensive collection mechanisms with technical details: automated fixation protocols, endogenous tagging (17 genes complete, 34 in process), SEC-MS in Krogan Lab, CRISPR library (6 guide RNAs per gene, 10x Genomics 3'HT kit)
semanticCollection mechanisms technically detailed and replicable; lab attributions (Lundberg Lab, Krogan Lab) enhance credibility; technical specifications (10x Genomics kit) support reproducibility
4/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull protocol minus minor gaps
evidencecollection_mechanisms: 4 detailed mechanisms (IF spatial proteomics by Lundberg Lab Stanford, AP-MS with 17 genes tagged in MDA-MB-468, SEC-MS by Krogan Lab UCSF, CRISPR screens with 10x Genomics). acquisition_methods: 4 methods (confocal microscopy 4-channel, mass spectrometry AP-MS/SEC-MS, single-cell RNA-seq 10x 3'HT, MuSIC pipeline integration). Data collectors identified by lab (Lundberg/Stanford, Krogan/UCSF). Core schema lacks data_collectors and collection_timeframes fields.
qualityCollection mechanisms and acquisition methods comprehensively documented with technical details (4-channel confocal, endogenous tagging, SEC-MS complexes, 10x Genomics kit). Data collectors inferred from mechanism descriptions (Lundberg Lab, Krogan Lab). Collection timeframes absent but project dates provided in funders (Sept 2022 - Aug 2026).
correctnessTechnical details accurate: 10x Genomics 3'HT is real single-cell sequencing kit, confocal microscopy 4-channel setup is standard (DAPI, ER, tubulin, target protein), AP-MS and SEC-MS are established proteomics methods.
consistencyCollection mechanisms align with instances (100 chromatin regulators) and acquisition methods (MS for PPIs, imaging for localization, RNA-seq for perturbations). Data collectors (Lundberg/Stanford for imaging, Krogan/UCSF for MS) align with creator affiliations.
acquisition_methods
acquisition_methods:
- id: cm4ai:acquisition:1
  description: 'Confocal Microscopy for Subcellular Imaging: High-resolution confocal microscopy of immunofluorescence-stained
    cells capturing four channels: DAPI (nuclei, blue), calreticulin antibody (ER, yellow), tubulin antibody
    (microtubules, red), and antibody against protein of interest (green). Images processed using Human
    Protein Atlas deep learning model to reduce dimensionality, producing image embeddings containing
    information about protein localization.

    '
- id: cm4ai:acquisition:2
  description: 'Mass Spectrometry for Protein Interactions: State-of-the-art mass spectrometry-based proteomics
    including AP-MS on endogenously tagged cell lines and SEC-MS for proteome-wide complex mapping. PPI
    networks processed using node2vec deep learning model to reduce dimensionality, producing PPI embeddings
    containing information about protein interactions.

    '
- id: cm4ai:acquisition:3
  description: 'Single-Cell RNA Sequencing for Perturbation Mapping: Single-cell RNA sequencing using
    10x Genomics 3''HT kit to capture transcriptional states following CRISPR perturbations. Generates
    genome-scale perturbation cell atlas mapping transcriptional and fitness phenotypes. Raw sequence
    data deposited to NCBI BioProject/Sequence Read Archive (SRA).

    '
- id: cm4ai:acquisition:4
  description: 'MuSIC Pipeline Integration: Multi-Scale Integrated Cell (MuSIC) pipeline integrates PPI
    embeddings and image embeddings using contrastive deep learning to obtain co-embeddings for each protein.
    Community detection performed based on all-by-all similarities of protein pairs in co-embedding space,
    producing hierarchical cell maps as final output. Maps annotated using Gene Ontology, Reactome pathways,
    and large language model approaches.

    '
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Data Acquisition Methods Listed
evidenceacquisition_methods: 4 detailed methods (confocal microscopy with 4 channels, mass spectrometry with AP-MS/SEC-MS, single-cell RNA-seq with 10x Genomics 3'HT, MuSIC pipeline integration with contrastive deep learning)
qualityAcquisition methods specify instruments and software: confocal microscopy (4-channel IF), mass spectrometry (AP-MS, SEC-MS), 10x Genomics sequencing, MuSIC pipeline with node2vec and HPA deep learning models
semanticAcquisition methods semantically appropriate for multimodal data; deep learning models (node2vec for PPI, HPA for imaging) support AI-ready data claims
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Software and Tools Documented
evidencepreprocessing_strategies: node2vec, Human Protein Atlas deep learning model, Cytoscape multiscale community detection, large language models for annotation, Integrative Modeling Platform (IMP) version 2.18, Python Modeling Interface; acquisition_methods: 10x Genomics 3'HT kit; cleaning_strategies: FAIRSCAPE-CLI, FAIRSCAPE server; external_resources: MuSIC software pipeline, FAIRSCAPE framework, IMP
qualityComprehensive software documentation: node2vec, HPA models, Cytoscape, LLMs, IMP 2.18, Python Modeling Interface, 10x Genomics kit, FAIRSCAPE-CLI; external resources link to documentation
semanticSoftware tools appropriate for multimodal integration; version specified for IMP (2.18) supports reproducibility; FAIRSCAPE framework documented with external resource link
4/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull protocol minus minor gaps
evidencecollection_mechanisms: 4 detailed mechanisms (IF spatial proteomics by Lundberg Lab Stanford, AP-MS with 17 genes tagged in MDA-MB-468, SEC-MS by Krogan Lab UCSF, CRISPR screens with 10x Genomics). acquisition_methods: 4 methods (confocal microscopy 4-channel, mass spectrometry AP-MS/SEC-MS, single-cell RNA-seq 10x 3'HT, MuSIC pipeline integration). Data collectors identified by lab (Lundberg/Stanford, Krogan/UCSF). Core schema lacks data_collectors and collection_timeframes fields.
qualityCollection mechanisms and acquisition methods comprehensively documented with technical details (4-channel confocal, endogenous tagging, SEC-MS complexes, 10x Genomics kit). Data collectors inferred from mechanism descriptions (Lundberg Lab, Krogan Lab). Collection timeframes absent but project dates provided in funders (Sept 2022 - Aug 2026).
correctnessTechnical details accurate: 10x Genomics 3'HT is real single-cell sequencing kit, confocal microscopy 4-channel setup is standard (DAPI, ER, tubulin, target protein), AP-MS and SEC-MS are established proteomics methods.
consistencyCollection mechanisms align with instances (100 chromatin regulators) and acquisition methods (MS for PPIs, imaging for localization, RNA-seq for perturbations). Data collectors (Lundberg/Stanford for imaging, Krogan/UCSF for MS) align with creator affiliations.
preprocessing_strategies
preprocessing_strategies:
- id: cm4ai:preproc:1
  description: 'Deep Learning Embedding Generation: PPI networks processed using node2vec deep learning
    model to reduce dimensionality and produce PPI embeddings. IF images processed using Human Protein
    Atlas deep learning model to reduce dimensionality and produce image embeddings. PPI and image embeddings
    integrated to obtain co-embeddings using contrastive deep learning, learning co-embeddings such that
    original embeddings can be reconstructed with minimal information loss.

    '
- id: cm4ai:preproc:2
  description: 'Hierarchical Community Detection: Community detection performed on co-embedding space
    using multiscale community detection algorithms implemented in Cytoscape. Produces hierarchical directed
    acyclic graphs (DAG) of protein assemblies at multiple resolutions, with 10 layers of depth representing
    communities from large cell compartments to small protein complexes.

    '
- id: cm4ai:preproc:3
  description: 'Cell Map Annotation: Two-pronged annotation approach: (1) Alignment to known protein function
    and pathway resources including Gene Ontology (GO) and Reactome to determine protein assemblies with
    high overlap with known cell biology, and (2) Large language model (LLM) approach to name sets of
    proteins and assign name confidence scores.

    '
- id: cm4ai:preproc:4
  description: 'Quality Control for Imaging Data: Standardized automated fixation and permeabilization
    protocols using pipetting robot. Consistent staining protocols across conditions using Human Protein
    Atlas antibodies. Quality control of imaging data before release and processing through MuSIC pipeline.

    '
- id: cm4ai:preproc:5
  description: 'Quality Control for Mass Spectrometry Data: Quality control and validation of AP-MS and
    SEC-MS data before public release. Mass spectrometry data for human iPSCs deposited to MassIVE Repository,
    and data for human cancer cells also deposited to MassIVE Repository.

    '
- id: cm4ai:preproc:6
  description: 'Integrative Structure Modeling: Bioinformatics pipeline for annotating MuSIC communities
    by available structural information about community members and their interactions. Structural information
    includes PDB, AlphaFold Protein Structure Database, crosslinking mass spectrometry, and prediction
    of disordered sequence segments. Communities ranked by structural information amount as proxy for
    integrative modeling feasibility. Modeling protocol scripted using Python Modeling Interface package
    based on Integrative Modeling Platform (IMP) version 2.18.

    '
⚠ low R20 · consistency
issueD4D-core schema subset evaluated - many full-schema fields absent by design
fieldspreprocessing_strategies, cleaning_strategies, labeling_strategies, software_and_tools, ethical_reviews, informed_consent, participant_compensation, vulnerable_populations
fixThis is expected for core schema variant - no action needed
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Preprocessing, Cleaning, and Labeling Strategies
evidencepreprocessing_strategies: 6 strategies (deep learning embedding generation with node2vec/HPA models, hierarchical community detection with Cytoscape, cell map annotation with GO/Reactome/LLM, QC for imaging, QC for mass spec, integrative structure modeling with IMP), cleaning_strategies: 2 strategies (FAIRSCAPE AI-readiness packaging with RO-Crate/ARK identifiers, data standard mapping to GO/Reactome/PDB/Schema.org/EVI)
qualityDetailed preprocessing pipeline: node2vec for PPI embeddings, HPA deep learning for image embeddings, contrastive learning for co-embeddings, community detection in Cytoscape (10 layers of DAG), GO/Reactome/LLM annotation, FAIRSCAPE packaging, ontology mapping
semanticPreprocessing strategies technically sound; MuSIC pipeline components well-described; FAIRSCAPE packaging aligns with AI-readiness goals; quality control documented for both imaging and mass spec
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Software and Tools Documented
evidencepreprocessing_strategies: node2vec, Human Protein Atlas deep learning model, Cytoscape multiscale community detection, large language models for annotation, Integrative Modeling Platform (IMP) version 2.18, Python Modeling Interface; acquisition_methods: 10x Genomics 3'HT kit; cleaning_strategies: FAIRSCAPE-CLI, FAIRSCAPE server; external_resources: MuSIC software pipeline, FAIRSCAPE framework, IMP
qualityComprehensive software documentation: node2vec, HPA models, Cytoscape, LLMs, IMP 2.18, Python Modeling Interface, 10x Genomics kit, FAIRSCAPE-CLI; external resources link to documentation
semanticSoftware tools appropriate for multimodal integration; version specified for IMP (2.18) supports reproducibility; FAIRSCAPE framework documented with external resource link
3/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelAt least one strategy documented
evidencepreprocessing_strategies: 6 strategies (deep learning embeddings via node2vec and HPA models, hierarchical community detection in Cytoscape, annotation via GO/Reactome/LLM, QC for imaging/MS, integrative structure modeling via IMP). cleaning_strategies: FAIRSCAPE framework, ontology mapping. Software mentioned: MuSIC pipeline, Cytoscape, FAIRSCAPE-CLI, Python Modeling Interface, IMP version 2.18. Core schema lacks labeling_strategies and software_and_tools fields.
qualityComprehensive preprocessing and cleaning strategies documented with specific software tools (Cytoscape, FAIRSCAPE, IMP v2.18). However, software version details incomplete (node2vec version not specified, HPA model version not specified). Labeling strategies not documented due to core schema limitations.
correctnessNamed tools are real and appropriate: node2vec (graph embedding), Cytoscape (network analysis), IMP (integrative modeling platform), FAIRSCAPE (RO-Crate framework). IMP version 2.18 is plausible version number.
consistencySoftware usage aligns with data types: node2vec for PPI networks, Cytoscape for community detection, IMP for structural modeling. Preprocessing strategies align with acquisition methods (embeddings for MS and imaging data).
cleaning_strategies
cleaning_strategies:
- id: cm4ai:cleaning:1
  description: 'FAIRSCAPE AI-Readiness Packaging: All datasets packaged using FAIRSCAPE framework which
    creates RO-Crate packages with datasets, metadata, provenance graphs, and software. FAIRSCAPE-CLI
    validates inputs and creates output RO-Crate packages. FAIRSCAPE server assigns persistent resolvable
    globally unique identifiers (ARK scheme), decomposes RO-Crates into components, and computes end-to-end
    provenance entailments using EVI Evidence Graph Ontology.

    '
- id: cm4ai:cleaning:2
  description: 'Data Standard Mapping: Data mapped to applicable standards and ontologies including Gene
    Ontology (GO), Reactome, Protein Data Bank (PDB), AlphaFold Protein Structure Database, Schema.org,
    and EVI Evidence Graph Ontology. Keywords mapped to controlled vocabularies from NCI Thesaurus, BioAssay
    Ontology, Cell Ontology, CHEBI, EFO, and other ontologies.

    '
⚠ low R20 · consistency
issueD4D-core schema subset evaluated - many full-schema fields absent by design
fieldspreprocessing_strategies, cleaning_strategies, labeling_strategies, software_and_tools, ethical_reviews, informed_consent, participant_compensation, vulnerable_populations
fixThis is expected for core schema variant - no action needed
✓ 1/1 R10 10.Cross-Platform and Community Integration Community Standards or Schema Conformance
evidenceconforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema, cleaning_strategies: Data mapped to Gene Ontology (GO), Reactome, Protein Data Bank (PDB), AlphaFold Protein Structure Database, Schema.org, EVI Evidence Graph Ontology, NCI Thesaurus, BioAssay Ontology, Cell Ontology, CHEBI, EFO
qualityExplicit conformance to Bridge2AI core schema plus extensive community ontology mappings (GO, Reactome, PDB, AlphaFold, Schema.org, EVI, NCI, BAO, CL, CHEBI, EFO)
semanticStandards conformance comprehensive; ontology mappings appropriate for functional genomics and proteomics data; Bridge2AI schema enhances ecosystem integration
✓ 1/1 R10 3.Data Reuse and Interoperability Schema or Ontology Conformance Stated
evidenceconforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema, cleaning_strategies: Data mapped to GO, Reactome, PDB, AlphaFold, Schema.org, EVI, NCI Thesaurus, BioAssay Ontology, Cell Ontology, CHEBI, EFO
qualityExplicit conformance to Bridge2AI core schema plus extensive ontology mappings (GO, Reactome, Schema.org, EVI, NCI, BAO, CL, CHEBI, EFO)
semanticOntology conformance appropriate for functional genomics data; multiple vocabularies enhance interoperability
✓ 1/1 R10 6.Data Provenance and Version Tracking Provenance and Source Derivation Documented
evidencedistribution_formats: RO-Crate packages with provenance graphs using EVI Evidence Graph Ontology, FAIRSCAPE framework computes end-to-end provenance entailments; cleaning_strategies: FAIRSCAPE creates RO-Crate packages with datasets, metadata, provenance graphs, and software
qualityComprehensive provenance documentation via FAIRSCAPE framework: RO-Crate packages, EVI Evidence Graph Ontology, end-to-end provenance entailments, persistent identifiers (ARK scheme)
semanticProvenance approach sophisticated and semantically appropriate for AI-ready data; RO-Crate and FAIRSCAPE align with FAIR principles
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Preprocessing, Cleaning, and Labeling Strategies
evidencepreprocessing_strategies: 6 strategies (deep learning embedding generation with node2vec/HPA models, hierarchical community detection with Cytoscape, cell map annotation with GO/Reactome/LLM, QC for imaging, QC for mass spec, integrative structure modeling with IMP), cleaning_strategies: 2 strategies (FAIRSCAPE AI-readiness packaging with RO-Crate/ARK identifiers, data standard mapping to GO/Reactome/PDB/Schema.org/EVI)
qualityDetailed preprocessing pipeline: node2vec for PPI embeddings, HPA deep learning for image embeddings, contrastive learning for co-embeddings, community detection in Cytoscape (10 layers of DAG), GO/Reactome/LLM annotation, FAIRSCAPE packaging, ontology mapping
semanticPreprocessing strategies technically sound; MuSIC pipeline components well-described; FAIRSCAPE packaging aligns with AI-readiness goals; quality control documented for both imaging and mass spec
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Software and Tools Documented
evidencepreprocessing_strategies: node2vec, Human Protein Atlas deep learning model, Cytoscape multiscale community detection, large language models for annotation, Integrative Modeling Platform (IMP) version 2.18, Python Modeling Interface; acquisition_methods: 10x Genomics 3'HT kit; cleaning_strategies: FAIRSCAPE-CLI, FAIRSCAPE server; external_resources: MuSIC software pipeline, FAIRSCAPE framework, IMP
qualityComprehensive software documentation: node2vec, HPA models, Cytoscape, LLMs, IMP 2.18, Python Modeling Interface, 10x Genomics kit, FAIRSCAPE-CLI; external resources link to documentation
semanticSoftware tools appropriate for multimodal integration; version specified for IMP (2.18) supports reproducibility; FAIRSCAPE framework documented with external resource link
5/5 R20 Q10 (Metadata Quality & Content) Interoperability and Standardization
levelStandard formats + schema/ontology compliance
evidenceconforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema, conforms_to_schema: src/data_sheets_schema/schema/data_sheets_schema_core.yaml, conforms_to_class: CoreDataset. distribution_formats mention RO-Crate (Research Object standard), Schema.org, EVI Evidence Graph Ontology. cleaning_strategies mention Gene Ontology, Reactome, PDB, AlphaFold, Schema.org, EVI ontology mappings.
qualityExceptional standardization with explicit schema conformance declarations, RO-Crate packaging (international FAIR standard), and extensive ontology mappings (GO, Reactome, Schema.org, EVI, NCI Thesaurus, Cell Ontology, CHEBI, EFO). Conforms_to fields provide machine-readable schema compliance.
correctnessRO-Crate is recognized Research Object Crate standard. Schema.org is W3C community standard. Gene Ontology and Reactome are authoritative bioinformatics resources. conforms_to URL follows W3ID persistent identifier pattern.
consistencySchema conformance aligns with file header comments (# Schema: D4D Core). Ontology usage aligns with data types (GO for protein function, Reactome for pathways, PDB/AlphaFold for structures).
3/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelAt least one strategy documented
evidencepreprocessing_strategies: 6 strategies (deep learning embeddings via node2vec and HPA models, hierarchical community detection in Cytoscape, annotation via GO/Reactome/LLM, QC for imaging/MS, integrative structure modeling via IMP). cleaning_strategies: FAIRSCAPE framework, ontology mapping. Software mentioned: MuSIC pipeline, Cytoscape, FAIRSCAPE-CLI, Python Modeling Interface, IMP version 2.18. Core schema lacks labeling_strategies and software_and_tools fields.
qualityComprehensive preprocessing and cleaning strategies documented with specific software tools (Cytoscape, FAIRSCAPE, IMP v2.18). However, software version details incomplete (node2vec version not specified, HPA model version not specified). Labeling strategies not documented due to core schema limitations.
correctnessNamed tools are real and appropriate: node2vec (graph embedding), Cytoscape (network analysis), IMP (integrative modeling platform), FAIRSCAPE (RO-Crate framework). IMP version 2.18 is plausible version number.
consistencySoftware usage aligns with data types: node2vec for PPI networks, Cytoscape for community detection, IMP for structural modeling. Preprocessing strategies align with acquisition methods (embeddings for MS and imaging data).
intended_uses
intended_uses:
- id: cm4ai:use:1
  description: 'AI Model Training for Functional Genomics: Primary intended use is training and development
    of artificial intelligence and machine learning models for functional genomics research. AI-ready
    datasets with full provenance, metadata, and validation enable immediate use in AI/ML pipelines without
    reformatting.

    '
- id: cm4ai:use:2
  description: 'Visible Neural Network Development: Development of visible neural networks (VNNs) that
    use hierarchical cell maps as interpretable model architectures. Unlike black box models, VNNs built
    on cell maps allow interrogation of how protein assemblies affect cell-level phenotypes, enabling
    interpretation of genetic variants and mutations in the context of cellular mechanisms.

    '
- id: cm4ai:use:3
  description: 'Genotype-Phenotype Mapping Research: Research into interpretable genotype-phenotype learning
    using multi-scale cell maps. Enables understanding of mechanisms by which genotypes translate to phenotypes,
    supporting precision medicine applications and genomic variant interpretation.

    '
- id: cm4ai:use:4
  description: 'Drug Response and Synergy Prediction: Analysis of cellular responses to drug treatments
    (paclitaxel, vorinostat) to predict drug response and synergy. Cell maps under different treatment
    conditions enable visible machine learning for drug discovery and personalized medicine applications.

    '
- id: cm4ai:use:5
  description: 'Disease Mechanism Research: Study of disease mechanisms in cancer, neuropsychiatric disorders,
    and cardiac disorders through analysis of chromatin modifiers and metabolic enzymes in disease-relevant
    cell contexts. Supports understanding of disease pathways and identification of therapeutic targets.

    '
- id: cm4ai:use:6
  description: 'Structural and Functional Genomics: Use as foundation for integrative structure modeling
    of protein communities, combining multimodal cell maps with PDB, AlphaFoldDB, and crosslinking mass
    spectrometry data to determine structural models of protein assemblies, as demonstrated in Nature
    publication (Schaffer, Hu et al., April 2025).

    '
- id: cm4ai:use:7
  description: 'Model for AI-Ready Biomedical Dataset Development: Use as exemplar for future AI-ready
    biomedical dataset development, demonstrating best practices in FAIR principles implementation, provenance
    tracking, ethical data governance, and AI-readiness packaging using RO-Crate and FAIRSCAPE frameworks.

    '
✓ 1/1 R10 3.Data Reuse and Interoperability Use Guidance Provided
evidenceintended_uses: 7 detailed use cases (AI model training, VNN development, genotype-phenotype mapping, drug response prediction, disease mechanism research, structural genomics, AI-ready dataset exemplar), discouraged_uses: 3 cases (clinical decision-making, incomplete data release, analysis without expertise)
qualityComprehensive use guidance with 7 intended uses providing scientific context and 3 discouraged uses highlighting limitations and ethical constraints
semanticUse guidance semantically appropriate for research data; clinical restrictions align with cell line nature; expertise requirements reasonable
discouraged_uses
discouraged_uses:
- id: cm4ai:discouraged:1
  description: 'Clinical Decision-Making Without Validation: Laboratory data from cell lines are not to
    be used in clinical decision-making or any context involving patient care without appropriate regulatory
    oversight and approval. Requires domain expertise and clinical validation before any clinical applications.
    Explicitly prohibited per dataset terms.

    '
- id: cm4ai:discouraged:2
  description: 'Use During Incomplete Data Release: This is an interim/beta release with data not yet
    in completed final form. Some datasets are under temporary pre-publication embargo, protein interrogation
    sets incompletely overlap across data modalities, and computed cell maps not yet included in releases.
    Full integration and final cell maps will be available in future releases through November 2026.

    '
- id: cm4ai:discouraged:3
  description: 'Analysis Without Domain Expertise: Datasets require domain expertise for meaningful analysis
    and interpretation. Current release is most suitable for bioinformatics analysis of individual datasets.
    Not suitable for use without understanding of functional genomics, proteomics, cell biology, and AI/ML
    methodologies. Training resources available through CM4AI Skills and Workforce Development module.

    '
✓ 1/1 R10 3.Data Reuse and Interoperability Use Guidance Provided
evidenceintended_uses: 7 detailed use cases (AI model training, VNN development, genotype-phenotype mapping, drug response prediction, disease mechanism research, structural genomics, AI-ready dataset exemplar), discouraged_uses: 3 cases (clinical decision-making, incomplete data release, analysis without expertise)
qualityComprehensive use guidance with 7 intended uses providing scientific context and 3 discouraged uses highlighting limitations and ethical constraints
semanticUse guidance semantically appropriate for research data; clinical restrictions align with cell line nature; expertise requirements reasonable
✓ 1/1 R10 5.Data Composition and Structure Data Quality Issues and Anomalies Documented
evidencediscouraged_uses: Incomplete data release (interim/beta), protein interrogation sets incompletely overlap across modalities, computed cell maps not yet included; updates: Beta releases on quarterly basis, future releases will include computed cell maps and complete integration
qualityQuality limitations explicitly documented: interim/beta status, incomplete modality overlap, computed cell maps pending; update schedule provided for completion
semanticQuality documentation transparent about beta status; limitations appropriate for ongoing data generation project with planned completion 2026
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Known Limitations Documented
evidencediscouraged_uses: Interim/beta release with data not in completed final form, some datasets under pre-publication embargo, protein interrogation sets incompletely overlap across modalities, computed cell maps not yet included; updates: Future releases will include computed cell maps and complete integration through Nov 2026
qualityExplicit limitations documented: beta status, incomplete data, incomplete modality overlap, missing computed cell maps, pre-publication embargoes
semanticLimitations appropriate for ongoing data generation project; transparency about beta status and planned completion supports responsible use
✗ 0/1 R10 9.Dataset Evaluation and Limitations Disclosure Data Anomalies and Quality Issues Noted
evidenceNo dedicated anomalies field in core schema; quality limitations described in discouraged_uses only
qualityCore schema lacks dedicated anomalies field; quality issues described in discouraged_uses (incomplete overlap, beta status) but not as explicit anomaly documentation
semanticAbsence expected for core schema; quality issues documented elsewhere but not as data-level anomalies
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Sensitive Content and Warnings Provided
evidencesensitive_elements: Cell line origin metadata (age, sex, race) retained; does not represent all biological variants; discouraged_uses: Not for clinical decision-making without validation, requires domain expertise for meaningful analysis, interim release with incomplete data
qualitySensitive content documented: cell line demographic metadata, clinical use prohibition, expertise requirements, beta release warnings
semanticWarnings appropriate for research data; clinical prohibition aligns with cell line nature; expertise requirement reasonable for complex multimodal data
license_and_use_terms
license_and_use_terms:
  id: cm4ai:license:1
  description: 'Data licensed for reuse under Creative Commons Attribution-NonCommercial-ShareAlike 4.0
    International license (https://creativecommons.org/licenses/by-nc-sa/4.0/). Attribution is required
    to the copyright holders and the Cell Maps for Artificial Intelligence project. Any publications referencing
    this data or derived products should cite the Nature article (Schaffer LV, Hu M, et al. Multimodal
    cell maps as a foundation for structural and functional genomics. Nature. 2025. doi:10.1038/s41586-025-08878-3)
    and the bioRxiv preprint (Clark T, et al. Cell Maps for Artificial Intelligence: AI-Ready Maps of
    Human Cell Architecture from Disease-Relevant Cell Lines. BioRXiv, May 2024. doi:10.1101/2024.05.21.589311)
    and directly cite the data collection. Commercial use requires separate license negotiation with copyright
    holder (UCSD, Stanford, and/or UCSF depending upon specific data package). A Data Access Committee
    (led by Jillian Parker) supervises ethical matters related to dataset distribution and potential dual
    licensing for commercial use. Copyright (c) 2025 The Regents of the University of California except
    where otherwise noted. Spatial proteomics raw image data is copyright (c) 2025 The Board of Trustees
    of the Leland Stanford Junior University.

    '
✓ 1/1 R10 10.Cross-Platform and Community Integration Citation and DOI for Cross-referencing
evidencedoi: 10.18130/V3/DXWOS5, license_and_use_terms: Publications must cite Nature article (doi:10.1038/s41586-025-08878-3) and bioRxiv preprint (doi:10.1101/2024.05.21.589311) and directly cite the data collection
qualityDataset DOI plus required citation format specifying Nature publication, bioRxiv preprint, and direct data citation
semanticCitation requirements comprehensive; DOI enables persistent cross-referencing; publication citations provide methodological context
✓ 1/1 R10 3.Data Reuse and Interoperability License Terms Allow Reuse
evidencelicense_and_use_terms: CC BY-NC-SA 4.0 with attribution required, publications must cite Nature article and bioRxiv preprint
qualityClear reuse terms with CC BY-NC-SA 4.0 license, attribution requirements, and citation guidance for publications
semanticLicense allows academic reuse with appropriate attribution; ShareAlike term promotes open science
5/5 R20 Q14 (Technical Documentation) Associated Publications
levelMultiple references and dataset citation
evidenceexternal_resources: 13 resources including Nature publication (doi:10.1038/s41586-025-08878-3), bioRxiv preprint (doi:10.1101/2024.05.21.589311), perturbation atlas publication (doi:10.1101/2024.11.03.621734), project website, NIH RePORTER, repositories (Dataverse, MassIVE, SRA, NDEx), frameworks (FAIRSCAPE, IMP), parent program (Bridge2AI, CFDE). license_and_use_terms includes citation requirements for Nature article and bioRxiv preprint.
qualityExcellent publication documentation with 3 peer-reviewed/preprint publications with DOIs, citation requirements in license terms, and comprehensive external resource links to data repositories, software frameworks, and program websites. Publications directly validate dataset (Nature structural genomics paper, perturbation atlas).
correctnessDOI formats valid (10.1038 = Nature, 10.1101 = bioRxiv). Publication dates plausible (April 2025 Nature, May 2024 bioRxiv preprint, Nov 2024 perturbation atlas). URLs use HTTPS and follow expected patterns for platforms (reporter.nih.gov, ndexbio.org, fairscape.github.io).
consistencyPublications align with dataset content: Nature paper on multimodal cell maps matches dataset description, bioRxiv preprint describes CM4AI project, perturbation atlas matches CRISPR screen data. External resources align with distribution formats (MassIVE for MS data, SRA for sequences, NDEx for networks).
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path
evidencedistribution_formats: 5 channels (RO-Crate at Dataverse DOI with ARK/DOI identifiers and JSON-LD landing pages, MassIVE for MS data, NCBI SRA for sequences, NDEx for networks, LibraData long-term preservation). distributions: 4 specific releases with DOIs, format (ZIP), media_type (application/zip), path (DOI URLs). license_and_use_terms: CC BY-NC-SA 4.0 with attribution requirements, commercial licensing process via Data Access Committee.
qualityAccess mechanisms comprehensively documented with multiple distribution channels for different data types, specific DOI access points for each release, file formats and media types specified, and licensing terms clearly explained. Users can access via DOI resolution, repository-specific portals, or project website.
correctnessDistribution channels are real platforms (Dataverse, MassIVE, NCBI SRA, NDEx). ZIP format and application/zip media type standard for archives. CC BY-NC-SA 4.0 is valid license identifier.
consistencyAccess mechanisms align with distribution_formats (RO-Crate ZIP archives match distributions format/media_type). License terms align with ip_restrictions (non-commercial clause, Data Access Committee for commercial use).
5/5 R20 Q18 (FAIRness & Accessibility) Reusability (License Clarity)
levelLicense explicitly defines reuse terms
evidencelicense_and_use_terms: CC BY-NC-SA 4.0 (https://creativecommons.org/licenses/by-nc-sa/4.0/) with detailed reuse conditions (attribution required to copyright holders and CM4AI project, publication citation required for Nature article and bioRxiv preprint, commercial use requires separate license, ShareAlike provision applies). license: CC BY-NC-SA 4.0.
qualityLicense exceptionally clear with standard CC license identifier, URL to full license text, and detailed explanation of reuse terms including attribution requirements, citation expectations, commercial use restrictions, and ShareAlike obligations. Reuse cases well-defined (academic research permitted, commercial requires negotiation).
correctnessCC BY-NC-SA 4.0 is valid Creative Commons license with standard URL. License provisions accurately described (BY=attribution, NC=non-commercial, SA=share-alike).
consistencyLicense field (CC BY-NC-SA 4.0) matches license_and_use_terms description. License restrictions align with ip_restrictions (commercial licensing via Data Access Committee). Attribution requirements align with funders and creators documentation.
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense_and_use_terms: CC BY-NC-SA 4.0 with attribution requirements, publication citation requirements, commercial licensing details, copyright holders (UCSD, Stanford, UCSF), Data Access Committee oversight. ip_restrictions: commercial use requires separate license negotiation, copyright details. regulatory_restrictions: non-clinical research data, not FDA regulated, NIH data sharing compliance, no HIPAA obligations, Data Access Committee oversight.
qualityComprehensive governance documentation covering license terms, intellectual property restrictions, regulatory context, and data access oversight. License clearly specified with reuse conditions. IP and regulatory restrictions documented with institutional oversight mechanisms.
correctnessCC BY-NC-SA 4.0 is valid Creative Commons license. NIH data sharing policy compliance appropriate for NIH-funded research. FDA/HIPAA non-applicability correct for non-clinical cell line data.
consistencyLicense restrictions align with IP restrictions (non-commercial clause consistent). Regulatory restrictions align with human_subject_research status (no HIPAA for non-human-subjects data).
distribution_formats
distribution_formats:
- id: cm4ai:format:1
  description: 'RO-Crate Packages with Provenance: All CM4AI output data packaged as Research Object Crate
    (RO-Crate) packages containing datasets, metadata, provenance graphs, and software (or resolvable
    references). RO-Crates assigned persistent globally unique identifiers (ARK scheme, DOIs planned for
    publishable work) that resolve to machine- and human-readable landing pages with metadata in JSON-LD
    using Schema.org and EVI vocabularies. Available at https://doi.org/10.18130/V3/DXWOS5, https://doi.org/10.18130/V3/B35XWX,
    https://doi.org/10.18130/V3/F3TD5R, https://doi.org/10.18130/V3/K7TGEM, and https://www.cm4ai.org.

    '
- id: cm4ai:format:2
  description: 'Mass Spectrometry Data in MassIVE: Mass spectrometry data deposited to MassIVE Repository
    (Proteomics community-supported repository). Separate depositions for human iPSC data and human cancer
    cell data (SEC-MS for KOLF2.1J iPSCs and MDA-MB-468 cancer cells). Data will be uploaded to PRIDE
    when available.

    '
- id: cm4ai:format:3
  description: 'Sequence Data in NCBI SRA: Raw sequence data from CRISPR perturbation screens deposited
    to NCBI BioProject/Sequence Read Archive (SRA). Genome-scale CRISPRi perturbation cell atlas raw sequences
    and processed data available.

    '
- id: cm4ai:format:4
  description: 'Hierarchical Cell Maps in NDEx: Cell maps shared via Network Data Exchange (NDEx) at https://www.ndexbio.org
    for visualization and access. Maps can be visualized in web browser or accessed via tools such as
    Cytoscape, HiView, and Python ndex2 library.

    '
- id: cm4ai:format:5
  description: 'University of Virginia Dataverse: Archived RO-Crates available in University of Virginia''s
    LibraData data archive (instance of Harvard''s Dataverse, an NIH-approved generalist repository).
    Long-term preservation supported by committed institutional funds. Quarterly updates through November
    2026. Available at https://doi.org/10.18130/V3/DXWOS5.

    '
✓ 1/1 R10 2.Dataset Access and Retrieval Distribution Formats and File Types Specified
evidencedistribution_formats: RO-Crate packages (ZIP with JSON-LD), MassIVE (mass spectrometry), NCBI SRA (sequence data), NDEx (cell maps), distributions: format: ZIP, media_type: application/zip
qualityComprehensive format documentation: RO-Crate (ZIP/JSON-LD), domain-specific repositories (MassIVE, SRA, NDEx), standard MIME types
semanticFormats semantically appropriate for multimodal data types; RO-Crate packaging aligns with AI-readiness goals
✓ 1/1 R10 3.Data Reuse and Interoperability Data Formats Are Standardized
evidencedistribution_formats: RO-Crate (ZIP, JSON-LD, Schema.org, EVI vocabularies), language: en, conforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema
qualityStandardized RO-Crate format with JSON-LD metadata using Schema.org and EVI ontology vocabularies, language specified
semanticFormats follow community standards (RO-Crate, Schema.org); conformance to Bridge2AI core schema documented
✓ 1/1 R10 6.Data Provenance and Version Tracking Provenance and Source Derivation Documented
evidencedistribution_formats: RO-Crate packages with provenance graphs using EVI Evidence Graph Ontology, FAIRSCAPE framework computes end-to-end provenance entailments; cleaning_strategies: FAIRSCAPE creates RO-Crate packages with datasets, metadata, provenance graphs, and software
qualityComprehensive provenance documentation via FAIRSCAPE framework: RO-Crate packages, EVI Evidence Graph Ontology, end-to-end provenance entailments, persistent identifiers (ARK scheme)
semanticProvenance approach sophisticated and semantically appropriate for AI-ready data; RO-Crate and FAIRSCAPE align with FAIR principles
5/5 R20 Q10 (Metadata Quality & Content) Interoperability and Standardization
levelStandard formats + schema/ontology compliance
evidenceconforms_to: https://w3id.org/bridge2ai/data-sheets-schema/core-schema, conforms_to_schema: src/data_sheets_schema/schema/data_sheets_schema_core.yaml, conforms_to_class: CoreDataset. distribution_formats mention RO-Crate (Research Object standard), Schema.org, EVI Evidence Graph Ontology. cleaning_strategies mention Gene Ontology, Reactome, PDB, AlphaFold, Schema.org, EVI ontology mappings.
qualityExceptional standardization with explicit schema conformance declarations, RO-Crate packaging (international FAIR standard), and extensive ontology mappings (GO, Reactome, Schema.org, EVI, NCI Thesaurus, Cell Ontology, CHEBI, EFO). Conforms_to fields provide machine-readable schema compliance.
correctnessRO-Crate is recognized Research Object Crate standard. Schema.org is W3C community standard. Gene Ontology and Reactome are authoritative bioinformatics resources. conforms_to URL follows W3ID persistent identifier pattern.
consistencySchema conformance aligns with file header comments (# Schema: D4D Core). Ontology usage aligns with data types (GO for protein function, Reactome for pathways, PDB/AlphaFold for structures).
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path
evidencedistribution_formats: 5 channels (RO-Crate at Dataverse DOI with ARK/DOI identifiers and JSON-LD landing pages, MassIVE for MS data, NCBI SRA for sequences, NDEx for networks, LibraData long-term preservation). distributions: 4 specific releases with DOIs, format (ZIP), media_type (application/zip), path (DOI URLs). license_and_use_terms: CC BY-NC-SA 4.0 with attribution requirements, commercial licensing process via Data Access Committee.
qualityAccess mechanisms comprehensively documented with multiple distribution channels for different data types, specific DOI access points for each release, file formats and media types specified, and licensing terms clearly explained. Users can access via DOI resolution, repository-specific portals, or project website.
correctnessDistribution channels are real platforms (Dataverse, MassIVE, NCBI SRA, NDEx). ZIP format and application/zip media type standard for archives. CC BY-NC-SA 4.0 is valid license identifier.
consistencyAccess mechanisms align with distribution_formats (RO-Crate ZIP archives match distributions format/media_type). License terms align with ip_restrictions (non-commercial clause, Data Access Committee for commercial use).
1/1 R20 Q20 (FAIRness & Accessibility) Interlinking Across Platforms
levelPass
evidenceexternal_resources: 13 cross-platform links (UVA Dataverse via DOI, CM4AI website, NIH RePORTER, Nature publication DOI, bioRxiv DOI, perturbation atlas DOI, FAIRSCAPE docs, IMP website, NDEx repository, Bridge2AI program, NIH CFDE, MassIVE mentioned, SRA mentioned). distribution_formats explicitly link to MassIVE, NCBI SRA, NDEx, Dataverse.
qualityExceptional cross-platform interlinking with 13 external resources spanning publication platforms (Nature, bioRxiv), repositories (Dataverse, MassIVE, SRA, NDEx), software frameworks (FAIRSCAPE, IMP), program sites (Bridge2AI, NIH CFDE, NIH RePORTER), and project website. Distribution formats explicitly connect data types to appropriate repositories.
correctnessExternal resource URLs valid for platforms (doi.org for publications, reporter.nih.gov for grants, ndexbio.org for networks, fairscape.github.io for framework). Platforms appropriate for data types (MassIVE for proteomics, SRA for sequences, NDEx for networks).
consistencyExternal resources align with distribution_formats (MassIVE, SRA, NDEx appear in both). Publication DOIs in external_resources match citation requirements in license_and_use_terms. Project website (cm4ai.org) matches page field.
4/5 R20 Q4 (Structural Completeness) File Enumeration and Type Variety
level2-3 file types
evidencedistribution_formats: 5 formats documented (RO-Crate ZIP, MassIVE proteomics, NCBI SRA sequences, NDEx networks, Dataverse archive). distributions: 4 specific releases with format: ZIP, media_type: application/zip
qualityMultiple distribution channels documented across specialized repositories, but primary format is ZIP (RO-Crate packages). Limited media type variety in distributions section (all ZIP), though distribution_formats describes diverse data types (proteomics, genomics, networks).
correctnessRO-Crate format appropriate for FAIR-compliant packaging, ZIP media type standard for archives
consistencyDistribution formats align with collection mechanisms (mass spec → MassIVE, sequences → SRA, networks → NDEx)
distributions
distributions:
- id: cm4ai:dist:1
  description: 'Primary RO-Crate release package (Beta V2.1) containing spatial proteomics IF images,
    SEC-MS protein interaction data, CRISPR perturbation screen results, and RO-Crate metadata with provenance
    graphs. Packaged as ZIP archive with JSON-LD metadata (ro-crate-metadata.json).

    '
  format: ZIP
  media_type: application/zip
  path: https://doi.org/10.18130/V3/DXWOS5
- id: cm4ai:dist:2
  description: 'March 2025 Beta release V1.4 (doi:10.18130/V3/B35XWX): includes perturb-seq in KOLF2.1J
    iPSCs, SEC-MS in iPSCs and derivatives, and IF images in MDA-MB-468 under three conditions. Packaged
    as RO-Crate (ZIP with JSON-LD metadata).

    '
  format: ZIP
  media_type: application/zip
  path: https://doi.org/10.18130/V3/B35XWX
- id: cm4ai:dist:3
  description: 'June 2025 Beta release V2.1 (doi:10.18130/V3/F3TD5R): revision adds RGB IF images, RO-Crate
    metadata corrections, naming convention changes, and SEC-MS for MDA-MB-468. Packaged as RO-Crate (ZIP
    with JSON-LD metadata).

    '
  format: ZIP
  media_type: application/zip
  path: https://doi.org/10.18130/V3/F3TD5R
- id: cm4ai:dist:4
  description: 'October 2025 Beta release (doi:10.18130/V3/K7TGEM): adds Perturb-seq for MDA-MB-468 breast
    cancer cells and additional SEC-MS data. Packaged as RO-Crate (ZIP with JSON-LD metadata).

    '
  format: ZIP
  media_type: application/zip
  path: https://doi.org/10.18130/V3/K7TGEM
✓ 1/1 R10 1.Dataset Discovery and Identification Hierarchical Structure (parent datasets, relationships)
evidencedistributions lists related versions (V1.4: doi:10.18130/V3/B35XWX, V2.1: doi:10.18130/V3/F3TD5R, Oct 2025: doi:10.18130/V3/K7TGEM)
qualityMultiple versioned releases documented with DOIs and release notes showing hierarchical dataset relationships
semanticVersion relationships clear with temporal progression and content augmentation described
✓ 1/1 R10 10.Cross-Platform and Community Integration Related Datasets with Typed Relationships
evidencedistributions: Multiple versioned releases with DOIs showing temporal 'is version of' relationships (V1.4: doi:10.18130/V3/B35XWX, V2.1: doi:10.18130/V3/F3TD5R, Oct 2025: doi:10.18130/V3/K7TGEM); external_resources link to related publications and data repositories
qualityVersion relationships documented with DOIs for each release; temporal progression from V1.4 (March 2025) → V2.1 (June 2025) → Oct 2025 release shows 'is version of' / 'supersedes' relationships
semanticVersioned dataset relationships clear with DOIs enabling precise version citation; external resources provide cross-dataset context
✓ 1/1 R10 2.Dataset Access and Retrieval Distribution Formats and File Types Specified
evidencedistribution_formats: RO-Crate packages (ZIP with JSON-LD), MassIVE (mass spectrometry), NCBI SRA (sequence data), NDEx (cell maps), distributions: format: ZIP, media_type: application/zip
qualityComprehensive format documentation: RO-Crate (ZIP/JSON-LD), domain-specific repositories (MassIVE, SRA, NDEx), standard MIME types
semanticFormats semantically appropriate for multimodal data types; RO-Crate packaging aligns with AI-readiness goals
✓ 1/1 R10 6.Data Provenance and Version Tracking Version Access Methods Documented
evidencedistributions: V1.4 (doi:10.18130/V3/B35XWX), V2.1 (doi:10.18130/V3/F3TD5R), Oct 2025 (doi:10.18130/V3/K7TGEM); all accessible via DOI with paths specified
qualityEach version accessible via unique DOI with persistent URLs to University of Virginia Dataverse
semanticVersion access well-documented with DOIs for each release; Dataverse ensures long-term version preservation
✓ 1/1 R10 6.Data Provenance and Version Tracking Change Descriptions and Errata Provided
evidencedistributions: March 2025 V1.4 (perturb-seq in iPSCs, SEC-MS in iPSCs, IF in MDA-MB-468), June 2025 V2.1 (RGB IF images, RO-Crate metadata corrections, naming convention changes, SEC-MS for MDA-MB-468), Oct 2025 (Perturb-seq for MDA-MB-468, additional SEC-MS)
qualityDetailed change descriptions for each release specifying added data modalities, metadata corrections, and naming convention updates
semanticChange documentation semantically clear with specific additions and corrections; progression shows systematic data augmentation
4/5 R20 Q13 (Technical Documentation) Version History Documentation
levelVersion number + access + update plan
evidenceversion: 2.1, updates: quarterly releases through Nov 2026, version history (v0.5 alpha, V1.4 March 2025, V2.1 June 2025, Oct 2025 beta). distributions provide version-specific DOIs and release notes. Core schema lacks version_access, errata, and release_notes fields.
qualityVersion number present (2.1) with detailed update schedule and version progression documented in updates and distributions fields. Version-specific DOIs provide access to historical versions. Lacks dedicated errata field and structured release notes, though distributions describe changes (RGB IF images added, metadata corrections, SEC-MS additions).
correctnessVersion numbering follows semantic pattern (0.5 alpha < 1.4 beta < 2.1 beta). Release dates chronologically consistent (March 2025 < June 2025 < Oct 2025).
consistencyVersion field (2.1) matches latest described release (June 2025 Beta V2.1 in distributions). Update schedule (quarterly through Nov 2026) aligns with project end date (Aug 31, 2026 from funders).
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path
evidencedistribution_formats: 5 channels (RO-Crate at Dataverse DOI with ARK/DOI identifiers and JSON-LD landing pages, MassIVE for MS data, NCBI SRA for sequences, NDEx for networks, LibraData long-term preservation). distributions: 4 specific releases with DOIs, format (ZIP), media_type (application/zip), path (DOI URLs). license_and_use_terms: CC BY-NC-SA 4.0 with attribution requirements, commercial licensing process via Data Access Committee.
qualityAccess mechanisms comprehensively documented with multiple distribution channels for different data types, specific DOI access points for each release, file formats and media types specified, and licensing terms clearly explained. Users can access via DOI resolution, repository-specific portals, or project website.
correctnessDistribution channels are real platforms (Dataverse, MassIVE, NCBI SRA, NDEx). ZIP format and application/zip media type standard for archives. CC BY-NC-SA 4.0 is valid license identifier.
consistencyAccess mechanisms align with distribution_formats (RO-Crate ZIP archives match distributions format/media_type). License terms align with ip_restrictions (non-commercial clause, Data Access Committee for commercial use).
4/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelChange notes with partial timestamps
evidenceupdates: detailed update plan with quarterly releases through Nov 2026, version history with dates (March 2025 V1.4, June 2025 V2.1, Oct 2025 beta), release content descriptions. distributions: version-specific changes documented (RGB IF images added in V2.1, metadata corrections, SEC-MS additions). Core schema lacks version_access field for structured provenance tracking.
qualityChange documentation present with version history, release dates, and content changes described in updates and distributions fields. FAIRSCAPE framework mentioned in cleaning_strategies provides end-to-end provenance entailments using EVI Evidence Graph Ontology. However, lacks structured version_access field for granular provenance tracking.
correctnessRelease dates chronologically consistent (March 2025 < June 2025 < Oct 2025). Version numbering progression plausible (V1.4 → V2.1). Update plan timeline (through Nov 2026) aligns with project end date (Aug 2026).
consistencyVersion field (2.1) matches latest release described in updates (June 2025 V2.1). Provenance framework (FAIRSCAPE with EVI ontology) aligns with distribution_formats (RO-Crate packages with provenance graphs).
4/5 R20 Q4 (Structural Completeness) File Enumeration and Type Variety
level2-3 file types
evidencedistribution_formats: 5 formats documented (RO-Crate ZIP, MassIVE proteomics, NCBI SRA sequences, NDEx networks, Dataverse archive). distributions: 4 specific releases with format: ZIP, media_type: application/zip
qualityMultiple distribution channels documented across specialized repositories, but primary format is ZIP (RO-Crate packages). Limited media type variety in distributions section (all ZIP), though distribution_formats describes diverse data types (proteomics, genomics, networks).
correctnessRO-Crate format appropriate for FAIR-compliant packaging, ZIP media type standard for archives
consistencyDistribution formats align with collection mechanisms (mass spec → MassIVE, sequences → SRA, networks → NDEx)
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass
evidencedoi: 10.18130/V3/DXWOS5, page: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5, additional DOIs in distributions (10.18130/V3/B35XWX, 10.18130/V3/F3TD5R, 10.18130/V3/K7TGEM), RRID identifiers in instances (RRID:CVCL_0419 for MDA-MB-468, RRID:CVCL_B5P3 for KOLF2.1J)
qualityExceptional identifier coverage with primary DOI, project website, multiple release-specific DOIs, and cell line RRID identifiers. All identifiers follow correct formats and are resolvable.
correctnessDOI prefix 10.18130 correctly identifies University of Virginia Dataverse repository. RRID format valid (RRID:CVCL_XXXX for cell lines). URLs use HTTPS protocol.
consistencyDOI in doi field matches download_url DOI. Multiple DOIs correspond to documented version releases (V1.4, V2.1, etc.)
maintainers
maintainers:
- id: cm4ai:maintainer:1
  description: 'CM4AI Consortium: Multidisciplinary consortium managing dataset maintenance including
    University of California San Diego (lead), University of California San Francisco, Stanford University,
    University of Virginia, Yale University, University of Alabama at Birmingham, Simon Fraser University,
    and The Hastings Center. Data Governance Committee led by Jillian Parker (jillianparker@health.ucsd.edu).
    Ethical Review by Vardit Ravitsky (ravitskyv@thehastingscenter.org) and Jean-Christophe Belisle-Pipon
    (jean-christophe_belisle-pipon@sfu.ca).

    '
✗ 0/1 R10 9.Dataset Evaluation and Limitations Disclosure Ethical Review Details Including Conflicts
evidenceNo ethical_reviews field in core schema; ethics documented via maintainers (Ravitsky, Belisle-Pipon as Ethics Module leaders), Data Access Committee (Jillian Parker), and human_subject_research statement
qualityCore schema lacks dedicated ethical_reviews field; ethics governance documented through maintainers and Data Access Committee but not as formal review documentation
semanticEthics oversight documented but not as formal review with conflicts disclosure; Ethics Module leaders identified in creators/maintainers
updates
updates:
  id: cm4ai:updates:1
  description: 'Dataset regularly updated and augmented through end of project in November 2026. Beta
    releases on quarterly basis with periodic data augmentation. Initial alpha release (v0.5) provided
    as supplemental data. March 2025 Beta (V1.4) includes perturb-seq in KOLF2.1J iPSCs, SEC-MS in iPSCs
    and derivatives, and IF images in MDA-MB-468 under three conditions. June 2025 Beta (V2.1) revision
    adds RGB IF images, ro-crate metadata corrections, and naming convention changes, plus SEC-MS for
    MDA-MB-468. October 2025 Beta adds Perturb-seq for MDA-MB-468 breast cancer cells and additional SEC-MS
    data. Future releases will include computed cell maps and complete integration of all data streams.
    Long-term preservation in University of Virginia Dataverse with committed institutional support.

    '
✓ 1/1 R10 5.Data Composition and Structure Data Quality Issues and Anomalies Documented
evidencediscouraged_uses: Incomplete data release (interim/beta), protein interrogation sets incompletely overlap across modalities, computed cell maps not yet included; updates: Beta releases on quarterly basis, future releases will include computed cell maps and complete integration
qualityQuality limitations explicitly documented: interim/beta status, incomplete modality overlap, computed cell maps pending; update schedule provided for completion
semanticQuality documentation transparent about beta status; limitations appropriate for ongoing data generation project with planned completion 2026
✓ 1/1 R10 6.Data Provenance and Version Tracking Update Schedule or Frequency Indicated
evidenceupdates: Dataset regularly updated and augmented through end of project in November 2026. Beta releases on quarterly basis with periodic data augmentation
qualityClear update schedule: quarterly beta releases through November 2026 with explicit end date
semanticUpdate schedule realistic and aligned with NIH grant timeline (Sept 2022 - Aug 2026)
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Known Limitations Documented
evidencediscouraged_uses: Interim/beta release with data not in completed final form, some datasets under pre-publication embargo, protein interrogation sets incompletely overlap across modalities, computed cell maps not yet included; updates: Future releases will include computed cell maps and complete integration through Nov 2026
qualityExplicit limitations documented: beta status, incomplete data, incomplete modality overlap, missing computed cell maps, pre-publication embargoes
semanticLimitations appropriate for ongoing data generation project; transparency about beta status and planned completion supports responsible use
4/5 R20 Q13 (Technical Documentation) Version History Documentation
levelVersion number + access + update plan
evidenceversion: 2.1, updates: quarterly releases through Nov 2026, version history (v0.5 alpha, V1.4 March 2025, V2.1 June 2025, Oct 2025 beta). distributions provide version-specific DOIs and release notes. Core schema lacks version_access, errata, and release_notes fields.
qualityVersion number present (2.1) with detailed update schedule and version progression documented in updates and distributions fields. Version-specific DOIs provide access to historical versions. Lacks dedicated errata field and structured release notes, though distributions describe changes (RGB IF images added, metadata corrections, SEC-MS additions).
correctnessVersion numbering follows semantic pattern (0.5 alpha < 1.4 beta < 2.1 beta). Release dates chronologically consistent (March 2025 < June 2025 < Oct 2025).
consistencyVersion field (2.1) matches latest described release (June 2025 Beta V2.1 in distributions). Update schedule (quarterly through Nov 2026) aligns with project end date (Aug 31, 2026 from funders).
4/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelChange notes with partial timestamps
evidenceupdates: detailed update plan with quarterly releases through Nov 2026, version history with dates (March 2025 V1.4, June 2025 V2.1, Oct 2025 beta), release content descriptions. distributions: version-specific changes documented (RGB IF images added in V2.1, metadata corrections, SEC-MS additions). Core schema lacks version_access field for structured provenance tracking.
qualityChange documentation present with version history, release dates, and content changes described in updates and distributions fields. FAIRSCAPE framework mentioned in cleaning_strategies provides end-to-end provenance entailments using EVI Evidence Graph Ontology. However, lacks structured version_access field for granular provenance tracking.
correctnessRelease dates chronologically consistent (March 2025 < June 2025 < Oct 2025). Version numbering progression plausible (V1.4 → V2.1). Update plan timeline (through Nov 2026) aligns with project end date (Aug 2026).
consistencyVersion field (2.1) matches latest release described in updates (June 2025 V2.1). Provenance framework (FAIRSCAPE with EVI ontology) aligns with distribution_formats (RO-Crate packages with provenance graphs).
retention_limit
retention_limit:
  id: cm4ai:retention:1
  description: 'Digital data maintained according to NIH data sharing policies with long-term preservation
    in University of Virginia''s LibraData repository supported by committed institutional funds. No planned
    sunset for data availability. Archived RO-Crates with persistent identifiers (ARK, future DOIs) ensure
    long-term accessibility and citability.

    '
no field-level feedback matched
human_subject_research
human_subject_research:
  id: cm4ai:hsr:1
  description: 'CM4AI data are distinctive within Bridge2AI in that they are non-clinical data from tissue
    cultures and are considered to be de-identified as they cannot be matched, with current knowledge,
    to a human subject. Both cell lines (MDA-MB-468 and KOLF2.1J) are commercially available, ethically
    sourced, de-identified cell lines. MDA-MB-468 available from ATCC. KOLF2.1J available from HipSci
    resource for non-profit organizations via simple MTA. Human Subjects: No. De-identified Samples: Yes.
    FDA Regulated: No.

    '
⚠ low R20 · consistency
issuehuman_subject_research indicates non-clinical cell line data (de-identified), is_deidentified provides detailed explanation
fieldshuman_subject_research, is_deidentified
fixConsistent logic - no human subjects research but de-identified cell lines documented
✗ 0/1 R10 4.Ethical Use and Privacy Safeguards IRB or Ethics Review Documented
evidencehuman_subject_research: Human Subjects: No, De-identified Samples: Yes, FDA Regulated: No; no dedicated ethical_reviews field in core schema
qualityCore schema lacks dedicated IRB/ethical_reviews field; ethics context provided via human_subject_research statement that data are non-clinical cell lines not requiring IRB oversight
semanticAbsence of IRB approval logically consistent with non-human-subjects research (commercially available de-identified cell lines); ethics documented in maintainers (Ravitsky, Belisle-Pipon) and Data Access Committee
✗ 0/1 R10 4.Ethical Use and Privacy Safeguards Informed Consent Obtained from Participants
evidenceNo informed_consent field in core schema; human_subject_research states data are non-clinical cell lines
qualityCore schema lacks informed_consent field; not applicable for commercially available de-identified cell lines obtained from ATCC and HipSci
semanticAbsence of consent documentation logically consistent with cell line data not involving active human subjects participation
✗ 0/1 R10 9.Dataset Evaluation and Limitations Disclosure Ethical Review Details Including Conflicts
evidenceNo ethical_reviews field in core schema; ethics documented via maintainers (Ravitsky, Belisle-Pipon as Ethics Module leaders), Data Access Committee (Jillian Parker), and human_subject_research statement
qualityCore schema lacks dedicated ethical_reviews field; ethics governance documented through maintainers and Data Access Committee but not as formal review documentation
semanticEthics oversight documented but not as formal review with conflicts disclosure; Ethics Module leaders identified in creators/maintainers
2/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelPartial ethics (deidentification only)
evidencehuman_subject_research: non-clinical data from tissue cultures, de-identified cell lines. is_deidentified: detailed explanation of de-identified commercially available cell lines (MDA-MB-468 from ATCC, KOLF2.1J from HipSci). sensitive_elements: cell line origin metadata documented. Core schema does not include ethical_reviews, informed_consent, participant_compensation, vulnerable_populations fields.
qualityD4D-core schema subset lacks comprehensive ethics fields. Human subjects research and de-identification documented, but IRB approval, informed consent, compensation, and vulnerable populations not applicable/present due to core schema limitations and non-human-subjects nature (cell lines).
correctnessCorrectly identifies data as non-human-subjects research (commercially available de-identified cell lines). De-identification explanation semantically appropriate for cell line data.
consistencyConsistent logic: human_subject_research indicates no human subjects + is_deidentified explains cell lines cannot be matched to individuals. Sensitive_elements acknowledges donor metadata retention for scientific context.
sensitive_elements
sensitive_elements:
- id: cm4ai:sensitive:1
  description: 'Cell Line Origin Metadata: While cell lines are de-identified and cannot be matched to
    specific individuals, metadata about cell line origins (age, sex, race of original donor) is retained
    for scientific context. This metadata does not constitute identifiable human subjects data under current
    knowledge. Data derived from commercially available de-identified human cell lines and does not represent
    all biological variants in the population at large.

    '
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Vulnerable Populations and Compensation Documented
evidencesensitive_elements: Cell line origin metadata includes age, sex, race of original donor; data does not represent all biological variants in population at large
qualitySensitive elements documented with acknowledgment that cell line metadata (age 51 years, black female for MDA-MB-468; male Northern European for KOLF2.1J) does not represent full population diversity
semanticSensitivity documentation appropriate; acknowledgment of limited biological diversity important for generalizability assessment
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Systematic Biases Identified and Described
evidencesensitive_elements: Data derived from commercially available de-identified human cell lines and does not represent all biological variants in the population at large; instances: MDA-MB-468 (51-year-old black female with metastatic breast cancer), KOLF2.1J (male Northern European donor)
qualityBiological bias acknowledged: limited cell line diversity (1 cancer line from black female, 1 iPSC line from Northern European male) does not represent population-wide biological variants
semanticBias documentation appropriate for cell line limitations; demographic specificity (age, sex, race) supports generalizability assessment
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Sensitive Content and Warnings Provided
evidencesensitive_elements: Cell line origin metadata (age, sex, race) retained; does not represent all biological variants; discouraged_uses: Not for clinical decision-making without validation, requires domain expertise for meaningful analysis, interim release with incomplete data
qualitySensitive content documented: cell line demographic metadata, clinical use prohibition, expertise requirements, beta release warnings
semanticWarnings appropriate for research data; clinical prohibition aligns with cell line nature; expertise requirement reasonable for complex multimodal data
2/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelPartial ethics (deidentification only)
evidencehuman_subject_research: non-clinical data from tissue cultures, de-identified cell lines. is_deidentified: detailed explanation of de-identified commercially available cell lines (MDA-MB-468 from ATCC, KOLF2.1J from HipSci). sensitive_elements: cell line origin metadata documented. Core schema does not include ethical_reviews, informed_consent, participant_compensation, vulnerable_populations fields.
qualityD4D-core schema subset lacks comprehensive ethics fields. Human subjects research and de-identification documented, but IRB approval, informed consent, compensation, and vulnerable populations not applicable/present due to core schema limitations and non-human-subjects nature (cell lines).
correctnessCorrectly identifies data as non-human-subjects research (commercially available de-identified cell lines). De-identification explanation semantically appropriate for cell line data.
consistencyConsistent logic: human_subject_research indicates no human subjects + is_deidentified explains cell lines cannot be matched to individuals. Sensitive_elements acknowledges donor metadata retention for scientific context.
is_deidentified
is_deidentified:
  id: cm4ai:deidentified:1
  description: 'CM4AI datasets are derived from commercially available, de-identified human cell lines
    (MDA-MB-468 from ATCC; KOLF2.1J from HipSci). Data cannot be matched to individual human subjects
    with current knowledge. Cell line origin metadata (donor age, sex, race) is retained for scientific
    context only and does not constitute identifiable human subjects data. De-identification confirmed
    by CM4AI Ethics Module and Data Access Committee.

    '
⚠ low R20 · consistency
issuehuman_subject_research indicates non-clinical cell line data (de-identified), is_deidentified provides detailed explanation
fieldshuman_subject_research, is_deidentified
fixConsistent logic - no human subjects research but de-identified cell lines documented
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Deidentification Method Described
evidenceis_deidentified: Datasets derived from commercially available, de-identified human cell lines (MDA-MB-468 from ATCC; KOLF2.1J from HipSci). Data cannot be matched to individual human subjects with current knowledge
qualityDeidentification approach clearly described as commercially available cell lines that cannot be matched to individuals; appropriate for cell line data
semanticDeidentification method appropriate for cell line data; commercial availability and de-identification status verified by ATCC and HipSci sources
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Privacy Protections Beyond Deidentification
evidenceis_deidentified: Cell line origin metadata (donor age, sex, race) retained for scientific context only, does not constitute identifiable human subjects data. De-identification confirmed by CM4AI Ethics Module and Data Access Committee
qualityPrivacy protections include Ethics Module review, Data Access Committee oversight, and explicit statement that origin metadata does not constitute identifiable data
semanticPrivacy approach appropriate for cell line research; Ethics Module and Data Access Committee provide governance oversight
2/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelPartial ethics (deidentification only)
evidencehuman_subject_research: non-clinical data from tissue cultures, de-identified cell lines. is_deidentified: detailed explanation of de-identified commercially available cell lines (MDA-MB-468 from ATCC, KOLF2.1J from HipSci). sensitive_elements: cell line origin metadata documented. Core schema does not include ethical_reviews, informed_consent, participant_compensation, vulnerable_populations fields.
qualityD4D-core schema subset lacks comprehensive ethics fields. Human subjects research and de-identification documented, but IRB approval, informed consent, compensation, and vulnerable populations not applicable/present due to core schema limitations and non-human-subjects nature (cell lines).
correctnessCorrectly identifies data as non-human-subjects research (commercially available de-identified cell lines). De-identification explanation semantically appropriate for cell line data.
consistencyConsistent logic: human_subject_research indicates no human subjects + is_deidentified explains cell lines cannot be matched to individuals. Sensitive_elements acknowledges donor metadata retention for scientific context.
ip_restrictions
ip_restrictions:
  id: cm4ai:ip:1
  description: 'Data licensed under CC BY-NC-SA 4.0. Commercial use requires separate license negotiation
    with copyright holders (UCSD, Stanford, and/or UCSF, depending on specific data package). Spatial
    proteomics raw image data copyright (c) 2025 The Board of Trustees of the Leland Stanford Junior University.
    Other data copyright (c) 2025 The Regents of the University of California. Data Access Committee (Jillian
    Parker) supervises ethical distribution matters and potential dual licensing for commercial use.

    '
✓ 1/1 R10 2.Dataset Access and Retrieval Access Policy and IP Restrictions Defined
evidencelicense: CC BY-NC-SA 4.0, ip_restrictions: Commercial use requires separate license negotiation with copyright holders
qualityClear CC BY-NC-SA 4.0 license with explicit IP restrictions for commercial use requiring separate licensing
semanticLicense appropriate for academic research data with commercial restrictions; IP policy consistent with multi-institution copyright
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense_and_use_terms: CC BY-NC-SA 4.0 with attribution requirements, publication citation requirements, commercial licensing details, copyright holders (UCSD, Stanford, UCSF), Data Access Committee oversight. ip_restrictions: commercial use requires separate license negotiation, copyright details. regulatory_restrictions: non-clinical research data, not FDA regulated, NIH data sharing compliance, no HIPAA obligations, Data Access Committee oversight.
qualityComprehensive governance documentation covering license terms, intellectual property restrictions, regulatory context, and data access oversight. License clearly specified with reuse conditions. IP and regulatory restrictions documented with institutional oversight mechanisms.
correctnessCC BY-NC-SA 4.0 is valid Creative Commons license. NIH data sharing policy compliance appropriate for NIH-funded research. FDA/HIPAA non-applicability correct for non-clinical cell line data.
consistencyLicense restrictions align with IP restrictions (non-commercial clause consistent). Regulatory restrictions align with human_subject_research status (no HIPAA for non-human-subjects data).
regulatory_restrictions
regulatory_restrictions:
  id: cm4ai:regulatory:1
  description: 'Data are non-clinical research data from tissue cultures and are not subject to FDA regulatory
    oversight. NIH data sharing policy compliance required. No HIPAA obligations (de-identified cell lines,
    not from living individuals under active care). Data Access Committee oversight required for commercial
    use licensing. NIH grant terms apply to funded research uses.

    '
✓ 1/1 R10 2.Dataset Access and Retrieval Regulatory Restrictions and Confidentiality Level Specified
evidenceregulatory_restrictions: Non-clinical research data, not subject to FDA oversight, NIH data sharing policy compliance required, no HIPAA obligations
qualityComprehensive regulatory context: FDA (not regulated), NIH policy compliance, HIPAA (not applicable), Data Access Committee oversight
semanticRegulatory restrictions appropriate for de-identified cell line data; HIPAA exemption logically consistent
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense_and_use_terms: CC BY-NC-SA 4.0 with attribution requirements, publication citation requirements, commercial licensing details, copyright holders (UCSD, Stanford, UCSF), Data Access Committee oversight. ip_restrictions: commercial use requires separate license negotiation, copyright details. regulatory_restrictions: non-clinical research data, not FDA regulated, NIH data sharing compliance, no HIPAA obligations, Data Access Committee oversight.
qualityComprehensive governance documentation covering license terms, intellectual property restrictions, regulatory context, and data access oversight. License clearly specified with reuse conditions. IP and regulatory restrictions documented with institutional oversight mechanisms.
correctnessCC BY-NC-SA 4.0 is valid Creative Commons license. NIH data sharing policy compliance appropriate for NIH-funded research. FDA/HIPAA non-applicability correct for non-clinical cell line data.
consistencyLicense restrictions align with IP restrictions (non-commercial clause consistent). Regulatory restrictions align with human_subject_research status (no HIPAA for non-human-subjects data).
external_resources
external_resources:
- id: cm4ai:resource:1
  description: 'CM4AI Project Website: Official project website and data portal using U-BRITE platform.
    https://www.cm4ai.org'
- id: cm4ai:resource:2
  description: 'NIH RePORTER Project Details: Federal grant information and project details for Bridge2AI
    Functional Genomics. https://reporter.nih.gov/project-details/11211616'
- id: cm4ai:resource:3
  description: 'University of Virginia Dataverse (LibraData): Repository with archived RO-Crates and data
    releases. https://doi.org/10.18130/V3/DXWOS5'
- id: cm4ai:resource:4
  description: 'Nature Publication: Schaffer LV, Hu M, Qian G, et al. Multimodal cell maps as a foundation
    for structural and functional genomics. Nature. Published April 9, 2025. https://doi.org/10.1038/s41586-025-08878-3

    '
- id: cm4ai:resource:5
  description: 'bioRxiv Preprint: Clark T, et al. Cell Maps for Artificial Intelligence: AI-Ready Maps
    of Human Cell Architecture from Disease-Relevant Cell Lines. BioRXiv, May 2024. https://doi.org/10.1101/2024.05.21.589311

    '
- id: cm4ai:resource:6
  description: 'FAIRSCAPE Framework Documentation: AI-readiness framework documentation, tutorial, and
    installation instructions. https://fairscape.github.io'
- id: cm4ai:resource:7
  description: 'Integrative Modeling Platform (IMP): Open source package for integrative structure modeling.
    http://integrativemodeling.org'
- id: cm4ai:resource:8
  description: 'Network Data Exchange (NDEx): Repository and visualization platform for cell maps and
    networks. https://www.ndexbio.org'
- id: cm4ai:resource:9
  description: 'Bridge2AI Program: Parent NIH Common Fund program supporting AI-ready biomedical datasets.
    https://commonfund.nih.gov/bridge2ai'
- id: cm4ai:resource:10
  description: 'NIH Common Fund Data Ecosystem (CFDE): Collaboration partner for data curation and integration.
    https://www.nih-cfde.org'
- id: cm4ai:resource:11
  description: 'MassIVE Proteomics Repository: Mass spectrometry data repository for iPSC and cancer cell
    SEC-MS data.'
- id: cm4ai:resource:12
  description: 'NCBI Sequence Read Archive (SRA): Repository for CRISPR perturbation screen raw sequence
    data.'
- id: cm4ai:resource:13
  description: 'Perturbation Cell Atlas Publication: Nourreddine S, Doctor Y, Dailamy A, et al. A PERTURBATION
    CELL ATLAS OF HUMAN INDUCED PLURIPOTENT STEM CELLS. bioRxiv. 2024 Nov 4. PMCID: PMC11580897. https://doi.org/10.1101/2024.11.03.621734

    '
✓ 1/1 R10 10.Cross-Platform and Community Integration Outreach Materials and Documentation Links
evidenceexternal_resources: CM4AI project website (cm4ai.org), NIH RePORTER project details, FAIRSCAPE framework documentation, IMP documentation, Bridge2AI program website, CFDE collaboration; Nature publication and bioRxiv preprints provide methodological documentation
qualityComprehensive documentation: project website, NIH grant details, framework documentation (FAIRSCAPE, IMP), parent program (Bridge2AI), collaboration (CFDE), peer-reviewed publication
semanticDocumentation links semantically appropriate for project context; NIH RePORTER provides grant transparency; software docs support reproducibility
✓ 1/1 R10 10.Cross-Platform and Community Integration Related Datasets with Typed Relationships
evidencedistributions: Multiple versioned releases with DOIs showing temporal 'is version of' relationships (V1.4: doi:10.18130/V3/B35XWX, V2.1: doi:10.18130/V3/F3TD5R, Oct 2025: doi:10.18130/V3/K7TGEM); external_resources link to related publications and data repositories
qualityVersion relationships documented with DOIs for each release; temporal progression from V1.4 (March 2025) → V2.1 (June 2025) → Oct 2025 release shows 'is version of' / 'supersedes' relationships
semanticVersioned dataset relationships clear with DOIs enabling precise version citation; external resources provide cross-dataset context
✓ 1/1 R10 2.Dataset Access and Retrieval Related Datasets and External Resources Linked
evidenceexternal_resources: 13 resources including Nature publication (doi:10.1038/s41586-025-08878-3), bioRxiv preprint, project website, FAIRSCAPE docs, IMP, NDEx, MassIVE, NCBI SRA
qualityExtensive external resources with peer-reviewed publications, preprints, software platforms, and data repositories
semanticExternal resources semantically coherent with dataset purposes and distribution formats; publication DOIs valid
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Software and Tools Documented
evidencepreprocessing_strategies: node2vec, Human Protein Atlas deep learning model, Cytoscape multiscale community detection, large language models for annotation, Integrative Modeling Platform (IMP) version 2.18, Python Modeling Interface; acquisition_methods: 10x Genomics 3'HT kit; cleaning_strategies: FAIRSCAPE-CLI, FAIRSCAPE server; external_resources: MuSIC software pipeline, FAIRSCAPE framework, IMP
qualityComprehensive software documentation: node2vec, HPA models, Cytoscape, LLMs, IMP 2.18, Python Modeling Interface, 10x Genomics kit, FAIRSCAPE-CLI; external resources link to documentation
semanticSoftware tools appropriate for multimodal integration; version specified for IMP (2.18) supports reproducibility; FAIRSCAPE framework documented with external resource link
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) External Standards and Resources Referenced
evidenceexternal_resources: 13 resources including Nature publication (doi:10.1038/s41586-025-08878-3), bioRxiv preprint (doi:10.1101/2024.05.21.589311), perturbation atlas publication (doi:10.1101/2024.11.03.621734), FAIRSCAPE docs, IMP docs, NDEx, Bridge2AI program, CFDE collaboration
qualityExtensive external standards: peer-reviewed Nature publication, bioRxiv preprints, software documentation (FAIRSCAPE, IMP), community platforms (NDEx), program context (Bridge2AI, CFDE)
semanticExternal resources semantically coherent with technical methods and scientific context; publication DOIs validate methodological rigor
5/5 R20 Q14 (Technical Documentation) Associated Publications
levelMultiple references and dataset citation
evidenceexternal_resources: 13 resources including Nature publication (doi:10.1038/s41586-025-08878-3), bioRxiv preprint (doi:10.1101/2024.05.21.589311), perturbation atlas publication (doi:10.1101/2024.11.03.621734), project website, NIH RePORTER, repositories (Dataverse, MassIVE, SRA, NDEx), frameworks (FAIRSCAPE, IMP), parent program (Bridge2AI, CFDE). license_and_use_terms includes citation requirements for Nature article and bioRxiv preprint.
qualityExcellent publication documentation with 3 peer-reviewed/preprint publications with DOIs, citation requirements in license terms, and comprehensive external resource links to data repositories, software frameworks, and program websites. Publications directly validate dataset (Nature structural genomics paper, perturbation atlas).
correctnessDOI formats valid (10.1038 = Nature, 10.1101 = bioRxiv). Publication dates plausible (April 2025 Nature, May 2024 bioRxiv preprint, Nov 2024 perturbation atlas). URLs use HTTPS and follow expected patterns for platforms (reporter.nih.gov, ndexbio.org, fairscape.github.io).
consistencyPublications align with dataset content: Nature paper on multimodal cell maps matches dataset description, bioRxiv preprint describes CM4AI project, perturbation atlas matches CRISPR screen data. External resources align with distribution formats (MassIVE for MS data, SRA for sequences, NDEx for networks).
1/1 R20 Q16 (FAIRness & Accessibility) Findability (Persistent Links)
levelPass
evidencepage: https://www.cm4ai.org, download_url: https://doi.org/10.18130/V3/DXWOS5, doi: 10.18130/V3/DXWOS5, external_resources: 13 persistent URLs including repository links, framework documentation, publication DOIs, NIH RePORTER
qualityExcellent persistent link coverage with DOI, project website, download URL, and extensive external resources. DOI provides machine- and human-readable landing page. Multiple access points documented.
correctnessDOI URL uses doi.org resolver (standard practice). Project website URL valid (cm4ai.org domain). External resource URLs use appropriate domains (reporter.nih.gov, ndexbio.org, etc.).
consistencyDOI in doi field matches download_url DOI. Page URL (cm4ai.org) matches external_resources project website entry.
1/1 R20 Q20 (FAIRness & Accessibility) Interlinking Across Platforms
levelPass
evidenceexternal_resources: 13 cross-platform links (UVA Dataverse via DOI, CM4AI website, NIH RePORTER, Nature publication DOI, bioRxiv DOI, perturbation atlas DOI, FAIRSCAPE docs, IMP website, NDEx repository, Bridge2AI program, NIH CFDE, MassIVE mentioned, SRA mentioned). distribution_formats explicitly link to MassIVE, NCBI SRA, NDEx, Dataverse.
qualityExceptional cross-platform interlinking with 13 external resources spanning publication platforms (Nature, bioRxiv), repositories (Dataverse, MassIVE, SRA, NDEx), software frameworks (FAIRSCAPE, IMP), program sites (Bridge2AI, NIH CFDE, NIH RePORTER), and project website. Distribution formats explicitly connect data types to appropriate repositories.
correctnessExternal resource URLs valid for platforms (doi.org for publications, reporter.nih.gov for grants, ndexbio.org for networks, fairscape.github.io for framework). Platforms appropriate for data types (MassIVE for proteomics, SRA for sequences, NDEx for networks).
consistencyExternal resources align with distribution_formats (MassIVE, SRA, NDEx appear in both). Publication DOIs in external_resources match citation requirements in license_and_use_terms. Project website (cm4ai.org) matches page field.

Unmatched feedback

Feedback that referenced multiple fields or no specific field.
✗ 0/1 R10 3.Data Reuse and Interoperability Variable Metadata with Identifiers Defined
evidenceNo variables field present
qualityCore schema does not include variable-level metadata; data organized as RO-Crate packages with hierarchical cell maps rather than tabular variables
semanticAbsence expected for core schema and non-tabular data (is_tabular: false); structure documented via RO-Crate metadata
⚠ info R10 · schema_limitation
issueD4D-core schema is an exchange-layer subset lacking some full-schema fields
fieldsethical_reviews, informed_consent, deidentification_and_privacy
fixThis is expected for core schema; no action needed

Recommendations

  1. R10 · Consider migrating to full D4D schema to capture detailed ethical reviews, informed consent context, and data anomalies documentation
  2. R10 · Add explicit relationship types (e.g., 'is version of', 'supersedes') in related_datasets field to formalize version relationships
  3. R10 · Include link to Skills and Workforce Development training materials mentioned in discouraged_uses to support domain expertise requirement
  4. R10 · Document known data-level anomalies separately from beta release limitations for clearer quality assessment
  5. R10 · Add temporal metadata (collection dates, processing dates) for each data modality to enhance provenance transparency
  6. R10 · Consider adding RRID for software tools (MuSIC pipeline, FAIRSCAPE) to improve tool citability and provenance
  7. R10 · Link to Data Access Committee contact information in access policy for commercial licensing inquiries
  8. R10 · Specify expected timeline for transition from beta to production release (currently scheduled through Nov 2026)
  9. R10 · Add explicit statement about data retention policy beyond project completion date
  10. R10 · Consider documenting computational requirements for processing RO-Crate packages and running MuSIC pipeline
  11. R20 · For full-schema D4D variant: Add ethical_reviews field documenting Data Access Committee review and ethical oversight by Ravitsky/Belisle-Pipon
  12. R20 · Specify software versions in preprocessing_strategies: node2vec version, HPA deep learning model version, FAIRSCAPE-CLI version, Cytoscape version
  13. R20 · Expand distributions to include additional media types beyond ZIP (e.g., application/x-hdf for processed embeddings, text/tab-separated-values for metadata tables)
  14. R20 · Add explicit collection_timeframes field with data acquisition dates for each modality (IF imaging dates, AP-MS dates, SEC-MS dates, CRISPR screen dates)
  15. R20 · Consider adding version_access field describing how to access historical versions (currently inferred from version-specific DOIs in distributions)
  16. R20 · Document any known errata or data quality issues in dedicated errata field (currently implied by 'interim/beta release' disclaimers in discouraged_uses)
  17. R20 · For transparency, add labeling_strategies describing community annotation approach (GO/Reactome alignment + LLM naming with confidence scores)
  18. R20 · Consider adding data_collectors field explicitly listing labs/individuals responsible for each data modality (currently inferred from collection_mechanisms)
  19. R20 · Add file byte sizes to distributions (in addition to instance counts) for users planning data transfers
  20. R20 · Document FAIRSCAPE provenance graph structure and EVI ontology usage in dedicated provenance field for semantic interoperability