Answer the grand challenge of improving recovery from acute illness by developing high-resolution mu...
purpose-001
Improve recovery from acute illness
Develop a publicly available, AI-ready critical care dataset from more than 100,000 critically ill p...
purpose-002
Create AI-ready critical care dataset
Unify standards to harmonize multi-modal EHR, waveform, imaging, and text data. Develop software and...
purpose-003
Establish data standards and tools
Ensure comprehensive sets of patient conditions and clinical treatment strategies with appropriate c...
purpose-004
Promote diversity and health equity
ID
funder-001
Name
NIH Common Fund Bridge2AI Program
Description
Funded through National Institutes of Health grant 1OT2OD032701-01, administered by NIH Office of the Director. Opportunity Number: OTA-21-008. Study Section: Data Coordination, Mapping, and Modeling [DCMM]. Fiscal Year 2022. Total funding in 2022: $5,880,300 (all direct costs). Project dates: September 1, 2022 to November 30, 2026 (with no-cost extension approved).
📊
Composition
What do the instances represent?
ID
instance-001
Name
Critically ill patients
Description
Individual critically ill patients requiring acute or critical care admitted to intensive care units or similar hospital settings across 14 data acquisition centers. As of November 2024, dataset covers 14 different hospitals with 23,400 unique admissions. Target enrollment exceeds 100,000 critically ill patients. Retrospective data collection from patients with acute or critical illness.
Instance Type
Human subjects - critically ill patients admitted to participating hospitals between data collection timeframe (specific dates vary by site, ongoing collection through November 2026).
Description
ID
Name
Patients distributed across 14 different hospitals within the CHoRUS network, ensuring geographic an...
subpop-001
Critically ill patients by hospital
Subset of patients with complete data across multiple modalities: structured EHR, waveform telemetry...
subpop-002
Patients with complete multi-modal data
Access Urls
Description
ID
Name
https://chorus4ai.org/
Structured electronic health record data distributed in OMOP Common Data Model format. Enables use o...
format-001
OMOP Common Data Model
https://chorus4ai.org/
Waveform telemetry data distributed in WFDB (WaveForm DataBase) format following extended PhysioNet ...
format-002
WFDB waveform format
https://chorus4ai.org/
Medical imaging data distributed in DICOM format with comprehensive metadata. De-identification in p...
format-003
DICOM imaging format
https://chorus4ai.org/
Clinical notes distributed as OHNLP-tokenized text following open source schema. Stored locally at s...
format-004
OHNLP tokenized text
https://chorus4ai.org/
EEG waveform data distributed in EDF+ (European Data Format) and Persyst formats following open sour...
Patient-Focused Collaborative Hospital Repository Uniting Standards (CHoRUS) for Equitable AI
CHoRUS for Equitable AI is a Bridge2AI data generation project developing the most diverse, high-resolution, ethically sourced, AI-ready critical care dataset to answer the grand challenge of improving recovery from acute illness. The project spans 20 academic centers (14 data acquisition centers) and creates a publicly available dataset of over 100,000 critically ill patients with multi-modal data including structured EHR, waveform telemetry, medical imaging, EEG, and clinical notes. All data is standardized to the OMOP Common Data Model with additional formats (DICOM, WFDB, OHNLP tokenization) and includes comprehensive metadata schemas. Patient-focused efforts determine ethical and legal approaches to manage privacy and bias while accounting for Social Determinants of Health. A visualization and annotation environment labels data with targets important for prediction. The project emphasizes skills and workforce development for a next generation of diverse academic and community AI scientists through training programs and partnerships with AIM-AHEAD. As of November 2024, the dataset covers 14 different hospitals with 23,400 unique admissions.
Address the absence of large-scale, diverse, high-resolution multi-center datasets for critical care...
gap-001
Lack of diverse multi-center critical care datasets
Overcome the lack of unified standards in critical care data by harmonizing multi-modal EHR, wavefor...
gap-002
Insufficient data standardization
Address ethical and legal challenges in AI through patient-focused efforts that determine approaches...
gap-003
Privacy and bias concerns in AI
Bridge the gap in diverse AI/ML workforce through comprehensive educational approaches, training pro...
gap-004
Limited AI workforce diversity
Role
Name
ORCID
Affiliation
Contributor
Eric S. Rosenthal
creator-001
-
Contributor
Azra Bihorac
creator-002
-
Contributor
Ashley Cordes
creator-003
-
Contributor
Gilles Clermont
creator-004
-
Contributor
Gari David Clifford
creator-005
-
Contributor
Barbara J. Evans
creator-006
-
Contributor
Xiao Hu
creator-007
-
Contributor
Rishikesan Kamaleswaran
creator-008
-
Contributor
Yulia A. Levites Strekalova
creator-009
-
Contributor
Parisa Rashidi
creator-010
-
Contributor
Cynthia Rudin
creator-011
-
Contributor
Ishan Canty Williams
creator-012
-
Contributor
Andrew Ewing Williams
creator-013
-
Contributor
Xiaoqian Jiang
creator-014
-
Contributor
Morteza Zabihi
creator-015
-
Contributor
Zhenhong Hu
creator-016
-
Contributor
Debora Simmons
creator-017
-
Contributor
Andrew Williams
creator-018
-
Contributor
Aliyah Geer
creator-019
-
Description
ID
Name
Primary dataset with controlled access requiring data use agreement and licensing. Includes OMOP-sta...
subset-001
Controlled Access Dataset
Clinical notes extracted and tokenized using OHNLP (Open Health Natural Language Processing) toolkit...
subset-002
Clinical Notes (Tokenized)
Imaging data from PACS (Picture Archiving and Communication System) in DICOM format. De-identificati...
subset-003
Medical Imaging
Electroencephalography (EEG) waveform data from hospital databases. Stored in EDF+ (European Data Fo...
subset-004
EEG Waveforms
Subsets of the dataset currently being used for training activities and publications as the dataset ...
subset-005
Training and Publication Datasets
ID
sampling-001
Name
Multi-center federated sampling
Description
Federated access enables sampling methods to ensure a balanced and diverse cohort across 20 academic centers (14 data acquisition centers). Legal framework established for collecting data at scale with sampling to ensure comprehensive sets of patient conditions and clinical treatment strategies.
Is Sample
True
Is Random
False
Is Representative
True
Strategies
Federated multi-center data collection across 14 hospitals
Sampling for balanced and diverse patient populations
Inclusion of contextual factors (geographic distance to hospital, Social Determinants of Health)
Community-facing ethics focus groups to determine appropriate data for public sharing
Legal framework for data collection at scale
Description
ID
Name
Retrospective data collection from electronic health record systems at 14 data acquisition centers. ...
collection-001
Retrospective EHR extraction
Continuous waveform telemetry data captured from bedside monitors through gateway/middleware systems...
collection-002
Waveform telemetry capture
Medical imaging data acquired from hospital PACS (Picture Archiving and Communication System) and st...
collection-003
Medical imaging acquisition
Electroencephalography recordings extracted from hospital EEG databases in EDF+ and Persyst formats ...
collection-004
EEG recording extraction
Multi-center network capabilities to acquire, standardize, tokenize, store, visualize, and label dat...
collection-005
Standardized data transformation
Description
ID
Name
Demographics, medication administration, procedures, nursing flowsheets, and diagnoses acquired from...
acquisition-001
Structured EHR data (OMOP)
Dosing information time-stamped upon each infusion change or dose administration. Stored in OMOP for...
acquisition-002
Medication administration records
Procedures and diagnoses documented by healthcare providers. Stored in OMOP format with controlled a...
acquisition-003
Provider documentation
Nursing documentation captured at high frequency intervals. Stored in OMOP format with controlled ac...
acquisition-004
High-frequency nursing flowsheets
Clinical notes extracted and tokenized using OHNLP (Open Health Natural Language Processing) toolkit...
acquisition-005
Clinical notes (tokenized)
Imaging data from PACS systems in DICOM format. De-identification in process. Planned controlled acc...
acquisition-006
Medical imaging (DICOM)
Bedside monitor waveform data acquired through gateway/middleware systems. Stored in WFDB format wit...
acquisition-007
Waveform telemetry (WFDB)
EEG recordings from hospital databases in EDF+ and Persyst formats. Extraction in process. Planned c...
acquisition-008
EEG waveforms
Description
ID
Name
Preprocessing Details
All structured electronic health record data standardized to the OMOP (Observational Medical Outcome...
preproc-001
OMOP Common Data Model transformation
Transformation of source EHR data to OMOP CDM, Standardization of terminology and codes, ... (+3 more)
Clinical notes processed using OHNLP (Open Health Natural Language Processing) toolkit for extractio...
preproc-002
Clinical note tokenization
Text extraction from clinical notes, OHNLP tokenization pipeline, ... (+3 more)
Waveform telemetry data from diverse bedside monitoring systems standardized to WFDB (WaveForm DataB...
preproc-003
Waveform standardization
Conversion from proprietary monitor formats, Standardization to WFDB format, ... (+3 more)
DICOM imaging data undergoing de-identification process to remove patient identifiable information w...
Digital data maintained according to NIH data sharing policies and institutional requirements. Controlled access model ensures long-term availability for research while protecting patient privacy.
Retention Details
NIH data sharing policies govern retention
Institutional requirements at participating centers
Controlled access model for privacy protection
Secure enclave infrastructure for data storage
Long-term maintenance through CHoRUS Consortium
Description
ID
Name
Sensitive Elements Present
Sensitivity Details
Complete electronic health records including demographics, diagnoses, procedures, medications, nursi...
sensitive-001
Clinical and medical data
True
Demographics and patient identifiers (de-identified), Diagnoses and medical conditions, ... (+4 more)
Continuous waveform telemetry and EEG recordings capturing detailed physiological states of critical...
Diagnostic imaging studies in DICOM format. Undergoing de-identification to remove embedded patient ...
sensitive-003
Medical imaging
True
CT scans and X-rays, MRI and other imaging modalities, ... (+2 more)
Contextual factors including geographic information (distance to hospital) and social determinants o...
sensitive-004
Social Determinants of Health
True
Geographic location information, Social determinants of health variables, ... (+2 more)
Role
Name
ORCID
Affiliation
Contributor
CHoRUS Project Website
resource-001
-
Contributor
CHoRUS GitHub Organization
resource-002
-
Contributor
Chorus_SOP Documentation
resource-003
-
Contributor
CHoRUS Developer Documentation
resource-004
-
Contributor
NIH RePORTER Project Details
resource-005
-
Contributor
Bridge2AI Program
resource-006
-
Contributor
AIM-AHEAD Training Partnership
resource-007
-
Contributor
OHDSI Community
resource-008
-
Contributor
Published Research
resource-009
-
Contributor
Contact for Access
resource-010
-
🚀
Uses
What (other) tasks could the dataset be used for?
Description
ID
Name
Generate data for ML/AI applications aimed at characterizing acute and critical care illness pattern...
task-001
Characterize acute and critical care illness
Enable prediction of complications among patients with acute or critical illness using multi-modal d...
task-002
Predict complications in critically ill patients
Support measurement and analysis of treatment response among critically ill patients through high-fr...
task-003
Measure treatment response
Provision a holdout test set accessible for model external validation to aid marketplace adoption of...
task-004
External validation for marketplace adoption
Utilize visualization and annotation environment to label data with targets important for prediction...
task-005
Label data for prediction targets
Description
ID
Name
Primary intended use is development and training of artificial intelligence and machine learning mod...
use-001
AI/ML model development for critical care
Provision of holdout test set accessible for model external validation to aid marketplace adoption o...
use-002
External validation of AI models
Studies examining health equity, social determinants of health, and disparities in critical care out...
use-003
Research on health equity and disparities
Research aimed at improving clinical care delivery, treatment protocols, and patient outcomes in acu...
use-004
Clinical care improvement research
Training and education of next generation of diverse academic and community AI scientists through ha...
use-005
Educational and training purposes
Description
ID
Name
As data collection continues through November 2026 and quality assurance processes are ongoing, earl...
discouraged-001
Uses during ongoing data collection
Dataset is for research purposes. Any AI/ML models developed should undergo appropriate clinical val...
discouraged-002
Clinical decision-making without validation
Attempts to re-identify patients from de-identified data violate ethical principles, data use agreem...
discouraged-003
Re-identification attempts
ID
license-001
Name
CHoRUS Controlled Access License
Description
Dataset distributed under controlled access requiring institutional email registration and signed licensing agreement. Access granted after review and approval process. Participants must complete registration form with name, email (institutional, not personal), and institution. Once approved, users receive email with access instructions to CHoRUS secure enclave. Contact for access requests: dbold@emory.edu or jared.houghtaling@tuftsmedicine.org.
License Terms
Institutional email required for registration
Signed licensing agreement required
Controlled access through secure enclave
Data use agreement specifies permitted uses
Prohibition on re-identification attempts
Compliance with ethical and legal requirements
📤
Distribution
How will the dataset be distributed?
Controlled Access with Data Use Agreement
🔄
Maintenance
How will the dataset be maintained?
ID
updates-001
Name
Ongoing data collection and expansion
Description
Dataset updated continuously as data collection progresses at 14 acquisition centers. As of November 2024, covers 14 hospitals with 23,400 unique admissions. Target exceeds 100,000 critically ill patients. Project timeline extends through November 30, 2026 (with approved no-cost extension). Regular status updates tracked through GitHub project management system. Sites provide updates via GitHub interface or Google Form submissions.
Frequency
Continuous updates through November 2026
Update Details
Ongoing retrospective data collection at 14 sites
Current status: 23,400 unique admissions (as of November 2024)
Target: 100,000+ critically ill patients
Regular site status updates via GitHub/Google Forms
GitHub project tracking for deliverables
Documentation updates in Chorus_SOP repository
Software and tooling continuous development
Semantic mapping validation and expansion
👥
Human Subjects
Does the dataset relate to people?
ID
hsr-001
Name
CHoRUS Human Subjects Research
Description
Retrospective data collection from critically ill patients approved through institutional review processes. Community-facing ethics focus groups conducted to determine what data is appropriate for public sharing. Legal framework established for collecting data at scale. Patient-focused efforts determine ethical and legal approaches to manage privacy and bias while accounting for Social Determinants of Health. Project draws expertise from law, ethics, health services, biomedical science, engineering, and scientific journal publications disciplines.
Involves Human Subjects
True
Ethics Review Board
Institutional review boards at 14 data acquisition centers
Community-facing ethics focus groups
Legal and ethical advisory teams
Privacy and accountability review processes
Generated on 2025-12-20 19:23:28 using Bridge2AI Data Sheets Schema