Persistent Identifier (DOI, RRID, or URI)
1/1
Fields:
Rationale: id: https://chorus4ai.org/
Dataset Title and Description Completeness
1/1
Fields:
Rationale: title: Patient-Focused Collaborative Hospital Repository Uniting Standards (CHoRUS) for Equitable AI; description: 600+ character comprehensive description covering scope, data types, institutions, ethical approach, and scale
Keywords or Tags for Searchability
1/1
Fields:
Rationale: keywords: 31 entries including 'CHoRUS', 'Bridge2AI', 'critical care', 'acute illness', 'OMOP Common Data Model', 'EHR', 'waveform telemetry', 'medical imaging', 'health equity', 'federated access', 'ethical AI', 'trustworthy AI', domain-specific terms
Landing Page and Resources (page, hierarchical resources)
1/1
Fields:
Rationale: page: https://chorus4ai.org/; external_resources includes multiple CHoRUS resources (website, GitHub org, documentation, developer resources)
Hierarchical Structure (parent datasets, relationships)
0/1
Fields:
Rationale: No parent_datasets or related_datasets fields found
Access Policy and IP Restrictions Defined
1/1
Fields:
Rationale: license_and_use_terms: CHoRUS Controlled Access License with detailed description of registration requirements (institutional email, signed licensing agreement, approval process), contact emails provided; ip_restrictions field not present
Regulatory Restrictions and Confidentiality Level Specified
0/1
Fields:
Rationale: Ethical and legal frameworks mentioned in description and human_subject_research, but regulatory_restrictions and confidentiality_level schema fields not populated
Download URL or Platform Link Available
1/1
Fields:
Rationale: distribution_formats.access_urls: https://chorus4ai.org/ (appears in 5 format entries); page: https://chorus4ai.org/; contact emails for access requests
Distribution Formats and File Types Specified
1/1
Fields:
Rationale: distribution_formats: 5 entries (OMOP Common Data Model, WFDB waveform format, DICOM imaging format, OHNLP tokenized text, EDF+/Persyst EEG formats) with descriptions and access information
Related Datasets and External Resources Linked
1/1
Fields:
Rationale: external_resources: 10 entries including CHoRUS website, GitHub organization, documentation, OHDSI community, published research, contact emails, NIH RePORTER, Bridge2AI program, AIM-AHEAD partnership; related_datasets field not present
License Terms Allow Reuse
1/1
Fields:
Rationale: license_and_use_terms: CHoRUS Controlled Access License specifies data use agreement, permitted research uses, ethical/legal compliance requirements
Data Formats Are Standardized (encoding, format)
1/1
Fields:
Rationale: distribution_formats specify OMOP, WFDB, DICOM, OHNLP, EDF+, Persyst - all recognized international standards; format/encoding top-level fields not populated
Schema or Ontology Conformance Stated
1/1
Fields:
Rationale: preprocessing_strategies and distribution_formats reference OMOP Common Data Model, DICOM schema, WFDB/PhysioNet schema, OHNLP schema, EDF+ schema, Persyst schema; conforms_to field not populated
Variable Metadata with Identifiers Defined
0/1
Fields:
Rationale: OMOP schema metadata mentioned as 'published' but variables field (as schema list) not populated; no explicit variable-level documentation in D4D file
Use Guidance Provided (intended, prohibited uses)
1/1
Fields:
Rationale: intended_uses: 5 entries (AI/ML model development, external validation, health equity research, clinical care improvement, educational/training); discouraged_uses: 3 entries (uses during ongoing collection, clinical decision-making without validation, re-identification attempts); prohibited_uses field not present
Motivation or Purpose for Dataset Creation
1/1
Fields:
Rationale: purposes: 4 detailed entries (improve recovery from acute illness, create AI-ready critical care dataset, establish data standards and tools, promote diversity and health equity)
Primary Research Objectives or Tasks
1/1
Fields:
Rationale: tasks: 5 entries (characterize acute/critical care illness, predict complications, measure treatment response, external validation for marketplace adoption, label data for prediction targets)
Funding Sources and Mechanisms Listed
1/1
Fields:
Rationale: funders: NIH Common Fund Bridge2AI Program with detailed grant information (1OT2OD032701-01, NIH Office of the Director)
Grant IDs or Award Numbers Present
1/1
Fields:
Rationale: funders: grant 1OT2OD032701-01, opportunity number OTA-21-008, study section DCMM, fiscal year 2022, total funding $5,880,300 (2022)
Creators and Acknowledgements Documented
1/1
Fields:
Rationale: creators: 19 entries including Contact PI (Eric S. Rosenthal, MGH), 11 Principal Investigators, Program Lead, lecturers/instructors, workshop leads with affiliations and roles
Collection Mechanisms and Settings Described
1/1
Fields:
Rationale: collection_mechanisms: 5 entries (retrospective EHR extraction, waveform telemetry capture, medical imaging acquisition, EEG recording extraction, standardized data transformation) with detailed descriptions of systems and processes
Data Acquisition Methods Listed
1/1
Fields:
Rationale: acquisition_methods: 8 entries covering structured EHR (OMOP), medication administration, provider documentation, nursing flowsheets, clinical notes (OHNLP), medical imaging (DICOM), waveform telemetry (WFDB), EEG waveforms (EDF+/Persyst)
Preprocessing, Cleaning, and Labeling Strategies
1/1
Fields:
Rationale: preprocessing_strategies: 5 entries (OMOP transformation, clinical note tokenization, waveform standardization, imaging de-identification, re-identification limitation); cleaning_strategies: 2 entries (multi-center harmonization, metadata schema validation); labeling_strategies: 1 entry (visualization and annotation environment)
Software and Tools Documented
1/1
Fields:
Rationale: preprocessing/acquisition reference OHDSI tool stack, OHNLP toolkit, OMOP CDM, PhysioNet schema, DICOM schema; GitHub organization (chorus-ai) with 28 repositories including chorus_waveform, privacy_scan_tool, chorus-container-apps; software_and_tools field not populated
External Standards and Resources Referenced
1/1
Fields:
Rationale: external_resources: 10 entries; preprocessing references OMOP CDM, DICOM, WFDB, OHNLP, EDF+, Persyst schemas; OHDSI community included in external_resources; conforms_to field not populated
Dataset Published on a Recognized Platform
1/1
Fields:
Rationale: CHoRUS secure enclave mentioned in license_and_use_terms, distribution_formats, and maintainers; external_resources includes NIH RePORTER, Bridge2AI program; publisher field not populated
Citation and DOI for Cross-referencing
1/1
Fields:
Rationale: external_resources includes published research with DOI (http://doi:10.1007/s12028-024-02007); id field is URL not DOI; citation field not populated
Community Standards or Schema Conformance
1/1
Fields:
Rationale: OMOP Common Data Model, DICOM, WFDB/PhysioNet, OHNLP, EDF+, Persyst schemas referenced extensively; OHDSI community in external_resources; FAIR principles implied by Bridge2AI program; conforms_to field not populated
Outreach Materials and Documentation Links
1/1
Fields:
Rationale: external_resources: 10 entries (CHoRUS website, GitHub org with 28 repos, Chorus_SOP documentation, developer docs, NIH RePORTER, Bridge2AI, AIM-AHEAD partnership, OHDSI community, published research, contact emails); page field present
Related Datasets with Typed Relationships
0/1
Fields:
Rationale: No related_datasets field found; Bridge2AI program mentioned but no typed relationships to other Bridge2AI datasets
Recommendations
Obtain DOI for dataset via appropriate repository (Zenodo, institutional, or domain-specific) to enable persistent identification and formal citation
Add RRID if applicable for additional identifier support and life sciences repository integration
Populate version field with current version/snapshot number and implement formal versioning scheme for continuous updates
Add version_access class documenting how users access specific dataset versions or snapshots
Create known_limitations field documenting: (1) ongoing collection status and completeness, (2) multi-center harmonization challenges, (3) non-random federated sampling, (4) generalizability across different ICU types, (5) variable data availability by modality (imaging/EEG in process), (6) institutional EHR system variability
Add known_biases field identifying: (1) academic medical center bias (14 university-affiliated centers), (2) geographic bias (US-only), (3) temporal bias (specific collection period), (4) clinical bias (ICU patients not general population)
Document informed_consent procedures and model for retrospective EHR collection (waiver rationale, opt-out mechanisms, institutional policies)
Add vulnerable_populations entry for critically ill patients with ethical considerations for this inherently vulnerable population
Add publisher field: 'CHoRUS Consortium / National Institutes of Health Bridge2AI Program'
Create formatted citation field for standard reference in publications
Add conforms_to field listing: OMOP Common Data Model, DICOM, WFDB (extended PhysioNet schema), OHNLP, EDF+, Persyst, OHDSI standards, FAIR principles
Populate conforms_to_schema with structured references to schema versions used
Add variables field with structured variable metadata (can reference external OMOP documentation but provide summary)
Add is_tabular field based on OMOP structured format
Add format and encoding fields for top-level dataset characteristics
Create is_deidentified class instance documenting de-identification methodology and privacy transformations
Add participant_privacy list documenting: (1) de-identification approaches, (2) controlled access model, (3) re-identification prohibition, (4) community ethics input, (5) privacy-preserving transformations
Create ethical_reviews list with entries for each of 14 IRB approvals including institutions and approval scope
Add regulatory_restrictions and confidentiality_level based on HIPAA framework and controlled access requirements
Add prohibited_uses field distinguishing prohibited from discouraged uses (e.g., commercial use restrictions, re-identification prohibition)
Add related_datasets entries with typed relationships to other Bridge2AI datasets (AI-READI, VOICE, CM4AI) and parent Bridge2AI program
Create software_and_tools list extracting from narratives: OHDSI tool stack, OHNLP toolkit, chorus_waveform, privacy_scan_tool, chorus-container-apps, visualization/annotation environment
Add anomalies field if quality issues identified during multi-center harmonization and validation
Add content_warnings field if sensitive clinical content requires user warnings beyond general sensitive_elements
Document participant_compensation (likely not applicable for retrospective collection but clarify)
Add was_derived_from documenting retrospective EHR extraction provenance
Add errata field for tracking data quality issues or corrections discovered post-release
Consider adding parent_datasets or hierarchical resources structure if dataset has logical subdivisions (by site, by modality, by time period)
Generated on 2026-01-13 16:57:40 using Bridge2AI Data Sheets Schema