Rubric20-Semantic Evaluation Report

71.0/84
84.5% Overall Score · Grade: B

Category Performance

Category
Structural Completeness
21/21
100.0%
Category
Metadata Quality & Content
21/21
100.0%
Category
Technical Documentation
21/25
84.0%
Category
FAIRness & Accessibility
15/17
88.2%

Question-Level Assessment

Structural Completeness

Q1. Field Completeness numeric
5/5
Assessment
All mandatory fields present with comprehensive, high-quality content. ID uses project homepage URL, title is descriptive, description is detailed (576 characters), keywords extensive (26 unique terms), and license fully documented with terms and contact information.
Evidence Found
id: https://chorus4ai.org/, title: 'Patient-Focused Collaborative Hospital Repository Uniting Standards (CHoRUS) for Equitable AI', description: 576 chars, keywords: 26 keywords, license_and_use_terms: comprehensive with detailed terms
Q2. Entry Length Adequacy numeric
5/5
Assessment
Excellent narrative depth. Main description is comprehensive at 576 characters covering dataset scope, scale, methodology, and current status. Purpose statements average 282 characters each with detailed rationale and context.
Evidence Found
description: 576 chars, purposes (4 entries): avg 282 chars per entry
Q3. Keyword Diversity numeric
5/5
Assessment
Outstanding keyword coverage with 26 diverse terms spanning dataset name, domain (critical care), data types (EHR, waveforms, imaging, EEG), standards (OMOP, DICOM, OHNLP), and cross-cutting themes (equity, ethics, privacy).
Evidence Found
keywords: 26 unique keywords including 'CHoRUS', 'Bridge2AI', 'critical care', 'OMOP Common Data Model', 'health equity', 'social determinants of health', 'multi-modal data', 'ethical AI', 'workforce development', 'privacy preservation', 'bias mitigation'
Q4. File Enumeration and Type Variety numeric
5/5
Assessment
Excellent format diversity reflecting multi-modal dataset. Five distinct distribution formats, each using appropriate international standards: OMOP for structured EHR, WFDB for waveforms, DICOM for imaging, OHNLP for text, EDF+/Persyst for EEG.
Evidence Found
distribution_formats: 5 formats listed - (1) OMOP Common Data Model, (2) WFDB waveform format, (3) DICOM imaging format, (4) OHNLP tokenized text, (5) EDF+ and Persyst EEG formats
Q5. Data File Size Availability pass_fail
1/1
Assessment
Clear instance count metadata. Dataset size documented as 23,400 unique admissions with target of 100,000+ patients. Metadata header also documents source concatenated file size (79K) and source file count (7).
Evidence Found
instances: 23,400 unique admissions (as of November 2024), target: 100,000+ critically ill patients. Source file metrics in header: 79K size, 7 source files

Metadata Quality & Content

Q6. Dataset Identification Metadata pass_fail
1/1
Assessment
Multiple persistent identifiers present. Project homepage URL used as primary ID and page. NIH RePORTER project number (10472824) provides federal grant tracking. Publication DOI available (though format has minor issue). No dataset-specific DOI yet, but project URL is persistent.
Evidence Found
id: https://chorus4ai.org/, page: https://chorus4ai.org/, external_resources: NIH RePORTER URL (https://reporter.nih.gov/project-details/10472824), publication DOI (http://doi:10.1007/s12028-024-02007)
Q7. Funding and Acknowledgements Completeness numeric
5/5
Assessment
Comprehensive funding and creator documentation. Funder entry includes grant number (1OT2OD032701-01), opportunity number (OTA-21-008), study section, fiscal year, total funding amount ($5,880,300 all direct costs), and project timeline. All 19 creators have institutional affiliations and specific roles.
Evidence Found
funders: NIH Common Fund Bridge2AI Program with grant 1OT2OD032701-01, fiscal year 2022, total funding $5,880,300, project dates Sept 2022 - Nov 2026. creators: 19 creators listed, all with institutional affiliations (MGH, UF, UT Health, Tufts, Emory) and specific roles (Contact PI, Principal Investigator, Lecturer, Workshop Lead)
Q8. Ethical and Privacy Declarations numeric
5/5
Assessment
Outstanding ethical documentation. Human subjects research clearly acknowledged with multi-level ethics review (institutional IRBs at 14 centers, community ethics focus groups, legal advisory teams). Comprehensive sensitive element enumeration with privacy protections documented for each data type. Strong emphasis on privacy preservation, bias mitigation, and Social Determinants of Health.
Evidence Found
human_subject_research: involves_human_subjects=true, ethics_review_board: IRBs at 14 data acquisition centers + community-facing ethics focus groups + legal/ethical advisory teams. sensitive_elements: 4 detailed entries covering clinical data (HIPAA-protected), physiological monitoring, medical imaging (de-identified), Social Determinants of Health. Privacy approaches: de-identification, tokenization, controlled access, privacy-preserving transformations, risk assessment
Q9. Access Requirements and Governance Documentation numeric
5/5
Assessment
Excellent governance documentation. Controlled access model clearly defined with institutional email requirement, signed licensing agreement, secure enclave infrastructure. Prohibition on re-identification attempts explicitly stated. Contact information provided for access requests. Appropriate for sensitive clinical data containing PHI.
Evidence Found
license: 'Controlled Access with Data Use Agreement', license_and_use_terms: detailed requirements (institutional email, signed licensing agreement, secure enclave access), contact emails provided (dbold@emory.edu, jared.houghtaling@tuftsmedicine.org). Explicit terms: prohibition on re-identification, compliance with ethical/legal requirements, HIPAA protection
Q10. Interoperability and Standardization numeric
5/5
Assessment
Exceptional interoperability through consistent use of international standards. Every data modality has documented schema conformance: OMOP CDM enables OHDSI tool stack, WFDB follows PhysioNet standards, DICOM for imaging, OHNLP for NLP, EDF+ for EEG. Metadata schema validation documented as part of cleaning strategies.
Evidence Found
Standard formats documented for all modalities: OMOP Common Data Model (structured EHR), WFDB format (waveforms, PhysioNet schema extended), DICOM (imaging with metadata schema), OHNLP open source schema (tokenized notes), EDF+ and Persyst open source schemas (EEG). Explicit schema conformance and metadata validation mentioned for each format.

Technical Documentation

Q11. Tool and Software Transparency numeric
4/5
Assessment
Very good technical documentation with comprehensive strategy descriptions. Each preprocessing, cleaning, and labeling strategy includes detailed preprocessing_details, cleaning_details, or labeling_details arrays. Software tools mentioned (OHDSI, OHNLP, WFDB) but specific version numbers not provided. GitHub organization (chorus-ai) with 28 repositories documented but individual tool versions missing. Minor deduction for lack of specific software versions.
Evidence Found
preprocessing_strategies: 5 detailed strategies (OMOP transformation, OHNLP tokenization, WFDB waveform standardization, DICOM de-identification, re-identification limitation). cleaning_strategies: 2 strategies (multi-center harmonization with SOPs, metadata schema validation). labeling_strategies: 1 strategy (visualization and annotation environment). Software mentioned: OMOP/OHDSI tools, OHNLP toolkit, PhysioNet WFDB. GitHub organization with 28 repositories documented.
Q12. Collection Protocol Clarity numeric
5/5
Assessment
Excellent collection protocol documentation. Five comprehensive collection mechanisms each with detailed descriptions. Eight acquisition methods specify exact data types, storage formats, and metadata schemas. Timeframes clearly documented (Sept 2022 - Nov 2026 with current status Nov 2024). Data collectors identified (14 acquisition centers, consortium structure). Retrospective study design explicit.
Evidence Found
collection_mechanisms: 5 detailed mechanisms (retrospective EHR extraction, waveform telemetry capture, medical imaging acquisition, EEG recording extraction, standardized data transformation). acquisition_methods: 8 detailed methods with specifics on data sources, formats, and access. Timeframes: project dates Sept 1, 2022 to Nov 30, 2026, current status as of November 2024. Data collectors: 14 data acquisition centers, multi-institutional consortium.
Q13. Version History Documentation numeric
4/5
Assessment
Good update and provenance documentation. Ongoing expansion clearly described with current metrics (23,400 admissions), target (100,000+), and timeline (through Nov 2026). GitHub project management tracks deliverables and site updates. However, no explicit version numbering scheme (e.g., v1.0, v2.0) or formal release notes documented. Deduction for lack of discrete version numbers and formal errata tracking.
Evidence Found
updates: 'Ongoing data collection and expansion' with current status (23,400 admissions Nov 2024), target (100,000+ patients), timeline (through Nov 2026), frequency (continuous updates). GitHub project management for tracking. version_access: not explicitly documented but GitHub organization and secure enclave infrastructure mentioned. No explicit version numbering system.
Q14. Associated Publications numeric
4/5
Assessment
Very good external resource documentation with multiple references. One peer-reviewed publication documented with DOI (though URL format needs correction). NIH RePORTER link provides federal grant context. GitHub organization with 28 repositories. Links to related platforms (OHDSI, Bridge2AI, AIM-AHEAD). No formal dataset citation field or additional publication references. Minor deduction for single publication and DOI format issue.
Evidence Found
external_resources: 10 entries including published research DOI (http://doi:10.1007/s12028-024-02007), NIH RePORTER project details (https://reporter.nih.gov/project-details/10472824), project website (https://chorus4ai.org/), GitHub organization (https://github.com/chorus-ai), Bridge2AI program reference, OHDSI community reference, AIM-AHEAD partnership. No formal dataset citation field.
Q15. Human Subject Representation numeric
4/5
Assessment
Good demographic and subpopulation documentation. Instance counts clear (23,400 current, 100,000+ target). Subpopulations documented by hospital (14 sites for geographic diversity) and by data completeness. Sampling strategies emphasize balanced and diverse cohorts with Social Determinants of Health. Strong equity focus with community ethics input. However, specific demographic breakdowns (age ranges, race/ethnicity, gender distributions) not provided. Minor deduction for lack of detailed demographic statistics.
Evidence Found
instances: 23,400 unique admissions from critically ill patients across 14 hospitals, target 100,000+ patients. subpopulations: 2 documented (patients by hospital for geographic diversity, patients with complete multi-modal data). sampling_strategies: multi-center federated sampling for balanced/diverse cohort, contextual factors (geographic distance, Social Determinants of Health). Emphasis on diversity and equity throughout.

FAIRness & Accessibility

Q16. Findability (Persistent Links) pass_fail
1/1
Assessment
Multiple persistent URLs provided. Project homepage (chorus4ai.org) serves as primary landing page. GitHub organization provides persistent code/documentation repository. NIH RePORTER provides federal grant tracking. Bridge2AI program page. Publication DOI for peer-reviewed article.
Evidence Found
page: https://chorus4ai.org/, external_resources: https://bridge2ai.org/chorus, https://github.com/chorus-ai, https://reporter.nih.gov/project-details/10472824, publication DOI
Q17. Accessibility (Access Mechanism) numeric
5/5
Assessment
Excellent access mechanism documentation. Clear step-by-step process: (1) registration with institutional email, (2) complete registration form, (3) sign licensing agreement, (4) approval review, (5) receive access instructions via email, (6) access secure enclave. Specific contact emails provided. All distribution formats reference https://chorus4ai.org/ for access. Controlled access model appropriate for sensitive clinical data.
Evidence Found
distribution_formats: 5 formats documented with access_urls. license_and_use_terms: step-by-step process (institutional email registration → signed licensing agreement → approval review → email with access instructions → secure enclave access). Contact emails: dbold@emory.edu, jared.houghtaling@tuftsmedicine.org
Q18. Reusability (License Clarity) numeric
5/5
Assessment
Excellent license clarity with explicit reuse terms. Controlled access license appropriate for sensitive clinical data containing PHI. Six explicit license terms documented. Data use agreement specifies permitted uses. Intended uses clearly enumerated (AI/ML development, external validation, health equity research, clinical care improvement, education). Discouraged uses explicitly stated (clinical decision-making without validation, re-identification attempts). Reuse constraints justified by privacy and ethical requirements.
Evidence Found
license: 'Controlled Access with Data Use Agreement', license_terms: 6 explicit terms including institutional email requirement, signed licensing agreement, data use agreement specifies permitted uses, prohibition on re-identification, compliance requirements. Permitted uses documented in intended_uses (5 entries) and discouraged_uses (3 entries).
Q19. Data Integrity and Provenance numeric
3/5
Assessment
Good provenance and update documentation but lacks structured change log. Continuous update model clearly described with current status snapshots (Nov 2024: 23,400 admissions). GitHub organization provides version control infrastructure (28 repositories, Chorus_SOP documentation, project tracking). However, no structured change log with specific timestamps for dataset releases or detailed version history. Generation timestamp in metadata header. Deduction for lack of formal versioning scheme and detailed change logs.
Evidence Found
updates: frequency='Continuous updates through November 2026', update_details include current status (23,400 admissions Nov 2024), target metrics, GitHub project tracking, documentation updates in Chorus_SOP repository, ongoing semantic mapping validation. GitHub organization for version control. Timestamp in D4D file header: 'Generated: 2025-12-20'
Q20. Interlinking Across Platforms pass_fail
1/1
Assessment
Strong cross-platform interlinking. Ten external resource entries connect to diverse platforms: project website (chorus4ai.org), GitHub (chorus-ai organization), NIH RePORTER (federal grant tracking), Bridge2AI program page, peer-reviewed publication, OHDSI community (for OMOP CDM), AIM-AHEAD training partnership. Links span data repositories, code repositories, documentation sites, funding agencies, and research communities.
Evidence Found
external_resources: 10 entries linking to multiple platforms - CHoRUS website, GitHub organization (https://github.com/chorus-ai with 28 repos), NIH RePORTER (https://reporter.nih.gov/project-details/10472824), Bridge2AI program (https://bridge2ai.org/chorus), OHDSI community, AIM-AHEAD training partnership, published research DOI

Semantic Analysis Summary

Consistency Checks

Passed: 22
Failed: 0
Warnings: 4

Issues Detected

correctness MEDIUM
Fields: external_resources
Recommendation:
consistency LOW
Fields: funders
Recommendation:
consistency LOW
Fields: external_resources
Recommendation:
semantic_understanding LOW
Fields: external_resources
Recommendation:
Generated on 2025-12-23 12:34:22 using Bridge2AI Data Sheets Schema