CHORUS Dataset Documentation

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

DescriptionIDName
Answer the grand challenge of improving recovery from acute illness by developing high-resolution mu...purpose-001Improve recovery from acute illness
Develop a publicly available, AI-ready critical care dataset from more than 100,000 critically ill p...purpose-002Create AI-ready critical care dataset
Unify standards to harmonize multi-modal EHR, waveform, imaging, and text data. Develop software and...purpose-003Establish data standards and tools
Ensure comprehensive sets of patient conditions and clinical treatment strategies with appropriate c...purpose-004Promote diversity and health equity
  • ID
    funder-001
    Name
    NIH Common Fund Bridge2AI Program
    Description
    Funded through National Institutes of Health grant 1OT2OD032701-01, administered by NIH Office of the Director. Opportunity Number: OTA-21-008. Study Section: Data Coordination, Mapping, and Modeling [DCMM]. Fiscal Year 2022. Total funding in 2022: $5,880,300 (all direct costs). Project dates: September 1, 2022 to November 30, 2026 (with no-cost extension approved).
📊

Composition

What do the instances represent?

  1. ID
    instance-001
    Name
    Critically ill patients
    Description
    Individual critically ill patients requiring acute or critical care admitted to intensive care units or similar hospital settings across 14 data acquisition centers. As of November 2024, dataset covers 14 different hospitals with 23,400 unique admissions. Target enrollment exceeds 100,000 critically ill patients. Retrospective data collection from patients with acute or critical illness.
    Instance Type
    Human subjects - critically ill patients admitted to participating hospitals between data collection timeframe (specific dates vary by site, ongoing collection through November 2026).
DescriptionIDName
Patients distributed across 14 different hospitals within the CHoRUS network, ensuring geographic an...subpop-001Critically ill patients by hospital
Subset of patients with complete data across multiple modalities: structured EHR, waveform telemetry...subpop-002Patients with complete multi-modal data
Access UrlsDescriptionIDName
https://chorus4ai.org/Structured electronic health record data distributed in OMOP Common Data Model format. Enables use o...format-001OMOP Common Data Model
https://chorus4ai.org/Waveform telemetry data distributed in WFDB (WaveForm DataBase) format following extended PhysioNet ...format-002WFDB waveform format
https://chorus4ai.org/Medical imaging data distributed in DICOM format with comprehensive metadata. De-identification in p...format-003DICOM imaging format
https://chorus4ai.org/Clinical notes distributed as OHNLP-tokenized text following open source schema. Stored locally at s...format-004OHNLP tokenized text
https://chorus4ai.org/EEG waveform data distributed in EDF+ (European Data Format) and Persyst formats following open sour...format-005EDF+ and Persyst EEG formats
🔍

Collection Process

How was the data acquired?

CHoRUS
Patient-Focused Collaborative Hospital Repository Uniting Standards (CHoRUS) for Equitable AI
CHoRUS for Equitable AI is a Bridge2AI data generation project developing the most diverse, high-resolution, ethically sourced, AI-ready critical care dataset to answer the grand challenge of improving recovery from acute illness. The project spans 20 academic centers (14 data acquisition centers) and creates a publicly available dataset of over 100,000 critically ill patients with multi-modal data including structured EHR, waveform telemetry, medical imaging, EEG, and clinical notes. All data is standardized to the OMOP Common Data Model with additional formats (DICOM, WFDB, OHNLP tokenization) and includes comprehensive metadata schemas. Patient-focused efforts determine ethical and legal approaches to manage privacy and bias while accounting for Social Determinants of Health. A visualization and annotation environment labels data with targets important for prediction. The project emphasizes skills and workforce development for a next generation of diverse academic and community AI scientists through training programs and partnerships with AIM-AHEAD. As of November 2024, the dataset covers 14 different hospitals with 23,400 unique admissions.
en
  • CHoRUS
  • Bridge2AI
  • critical care
  • acute illness
  • AI-ready dataset
  • OMOP Common Data Model
  • electronic health records
  • EHR
  • waveform telemetry
  • medical imaging
  • DICOM
  • EEG
  • clinical notes
  • OHNLP
  • health equity
  • social determinants of health
  • federated access
  • multi-modal data
  • high-resolution data
  • ethical AI
  • trustworthy AI
  • workforce development
  • data standardization
  • privacy preservation
  • bias mitigation
DescriptionIDName
Address the absence of large-scale, diverse, high-resolution multi-center datasets for critical care...gap-001Lack of diverse multi-center critical care datasets
Overcome the lack of unified standards in critical care data by harmonizing multi-modal EHR, wavefor...gap-002Insufficient data standardization
Address ethical and legal challenges in AI through patient-focused efforts that determine approaches...gap-003Privacy and bias concerns in AI
Bridge the gap in diverse AI/ML workforce through comprehensive educational approaches, training pro...gap-004Limited AI workforce diversity
RoleNameORCIDAffiliation
ContributorEric S. Rosenthalcreator-001-
ContributorAzra Bihoraccreator-002-
ContributorAshley Cordescreator-003-
ContributorGilles Clermontcreator-004-
ContributorGari David Cliffordcreator-005-
ContributorBarbara J. Evanscreator-006-
ContributorXiao Hucreator-007-
ContributorRishikesan Kamaleswarancreator-008-
ContributorYulia A. Levites Strekalovacreator-009-
ContributorParisa Rashidicreator-010-
ContributorCynthia Rudincreator-011-
ContributorIshan Canty Williamscreator-012-
ContributorAndrew Ewing Williamscreator-013-
ContributorXiaoqian Jiangcreator-014-
ContributorMorteza Zabihicreator-015-
ContributorZhenhong Hucreator-016-
ContributorDebora Simmonscreator-017-
ContributorAndrew Williamscreator-018-
ContributorAliyah Geercreator-019-
DescriptionIDName
Primary dataset with controlled access requiring data use agreement and licensing. Includes OMOP-sta...subset-001Controlled Access Dataset
Clinical notes extracted and tokenized using OHNLP (Open Health Natural Language Processing) toolkit...subset-002Clinical Notes (Tokenized)
Imaging data from PACS (Picture Archiving and Communication System) in DICOM format. De-identificati...subset-003Medical Imaging
Electroencephalography (EEG) waveform data from hospital databases. Stored in EDF+ (European Data Fo...subset-004EEG Waveforms
Subsets of the dataset currently being used for training activities and publications as the dataset ...subset-005Training and Publication Datasets
  1. ID
    sampling-001
    Name
    Multi-center federated sampling
    Description
    Federated access enables sampling methods to ensure a balanced and diverse cohort across 20 academic centers (14 data acquisition centers). Legal framework established for collecting data at scale with sampling to ensure comprehensive sets of patient conditions and clinical treatment strategies.
    Is Sample
    • True
    Is Random
    • False
    Is Representative
    • True
    Strategies
    • Federated multi-center data collection across 14 hospitals
    • Sampling for balanced and diverse patient populations
    • Inclusion of contextual factors (geographic distance to hospital, Social Determinants of Health)
    • Community-facing ethics focus groups to determine appropriate data for public sharing
    • Legal framework for data collection at scale
DescriptionIDName
Retrospective data collection from electronic health record systems at 14 data acquisition centers. ...collection-001Retrospective EHR extraction
Continuous waveform telemetry data captured from bedside monitors through gateway/middleware systems...collection-002Waveform telemetry capture
Medical imaging data acquired from hospital PACS (Picture Archiving and Communication System) and st...collection-003Medical imaging acquisition
Electroencephalography recordings extracted from hospital EEG databases in EDF+ and Persyst formats ...collection-004EEG recording extraction
Multi-center network capabilities to acquire, standardize, tokenize, store, visualize, and label dat...collection-005Standardized data transformation
DescriptionIDName
Demographics, medication administration, procedures, nursing flowsheets, and diagnoses acquired from...acquisition-001Structured EHR data (OMOP)
Dosing information time-stamped upon each infusion change or dose administration. Stored in OMOP for...acquisition-002Medication administration records
Procedures and diagnoses documented by healthcare providers. Stored in OMOP format with controlled a...acquisition-003Provider documentation
Nursing documentation captured at high frequency intervals. Stored in OMOP format with controlled ac...acquisition-004High-frequency nursing flowsheets
Clinical notes extracted and tokenized using OHNLP (Open Health Natural Language Processing) toolkit...acquisition-005Clinical notes (tokenized)
Imaging data from PACS systems in DICOM format. De-identification in process. Planned controlled acc...acquisition-006Medical imaging (DICOM)
Bedside monitor waveform data acquired through gateway/middleware systems. Stored in WFDB format wit...acquisition-007Waveform telemetry (WFDB)
EEG recordings from hospital databases in EDF+ and Persyst formats. Extraction in process. Planned c...acquisition-008EEG waveforms
DescriptionIDNamePreprocessing Details
All structured electronic health record data standardized to the OMOP (Observational Medical Outcome...preproc-001OMOP Common Data Model transformationTransformation of source EHR data to OMOP CDM, Standardization of terminology and codes, ... (+3 more)
Clinical notes processed using OHNLP (Open Health Natural Language Processing) toolkit for extractio...preproc-002Clinical note tokenizationText extraction from clinical notes, OHNLP tokenization pipeline, ... (+3 more)
Waveform telemetry data from diverse bedside monitoring systems standardized to WFDB (WaveForm DataB...preproc-003Waveform standardizationConversion from proprietary monitor formats, Standardization to WFDB format, ... (+3 more)
DICOM imaging data undergoing de-identification process to remove patient identifiable information w...preproc-004Medical imaging de-identificationDICOM header de-identification, Removal of embedded patient information, ... (+3 more)
Data transformed using approaches that limit re-identification while maintaining analytical utility....preproc-005Data re-identification limitationApplication of de-identification algorithms, Privacy-preserving transformations, ... (+3 more)
Cleaning DetailsDescriptionIDName
Semantic mapping validation by clinical experts, Standard operating protocol implementation, ... (+3 more)Data from 14 acquisition centers harmonized through validated semantic mappings and standard operati...cleaning-001Multi-center data harmonization
Schema compliance verification, Metadata completeness checks, ... (+3 more)All data modalities validated against published metadata schemas (OMOP, DICOM, WFDB, OHNLP, EDF+, Pe...cleaning-002Metadata schema validation
  1. ID
    labeling-001
    Name
    Visualization and annotation environment
    Description
    Custom visualization and annotation environment developed to label data with targets important for prediction tasks in critical care AI applications.
    Labeling Details
    • Interactive visualization tools
    • Annotation interface for clinical experts
    • Labeling of prediction targets
    • Quality control of annotations
    • Documentation of labeling protocols
DescriptionIDMaintainer DetailsName
Multi-institutional consortium managing dataset maintenance including data acquisition centers, coor...maintainer-001Massachusetts General Hospital (lead institution), University of Florida (data acquisition and coordination), ... (+7 more)CHoRUS Consortium
Active GitHub organization (chorus-ai) housing repositories for software, semantic mappings, standar...maintainer-002GitHub organization: https://github.com/chorus-ai, Chorus_SOP repository (centralized documentation), ... (+5 more)CHoRUS GitHub Organization
ID
retention-001
Name
Long-term dataset retention
Description
Digital data maintained according to NIH data sharing policies and institutional requirements. Controlled access model ensures long-term availability for research while protecting patient privacy.
Retention Details
  • NIH data sharing policies govern retention
  • Institutional requirements at participating centers
  • Controlled access model for privacy protection
  • Secure enclave infrastructure for data storage
  • Long-term maintenance through CHoRUS Consortium
DescriptionIDNameSensitive Elements PresentSensitivity Details
Complete electronic health records including demographics, diagnoses, procedures, medications, nursi...sensitive-001Clinical and medical dataTrueDemographics and patient identifiers (de-identified), Diagnoses and medical conditions, ... (+4 more)
Continuous waveform telemetry and EEG recordings capturing detailed physiological states of critical...sensitive-002Physiological monitoring dataTrueContinuous cardiac waveforms, Respiratory monitoring data, ... (+3 more)
Diagnostic imaging studies in DICOM format. Undergoing de-identification to remove embedded patient ...sensitive-003Medical imagingTrueCT scans and X-rays, MRI and other imaging modalities, ... (+2 more)
Contextual factors including geographic information (distance to hospital) and social determinants o...sensitive-004Social Determinants of HealthTrueGeographic location information, Social determinants of health variables, ... (+2 more)
RoleNameORCIDAffiliation
ContributorCHoRUS Project Websiteresource-001-
ContributorCHoRUS GitHub Organizationresource-002-
ContributorChorus_SOP Documentationresource-003-
ContributorCHoRUS Developer Documentationresource-004-
ContributorNIH RePORTER Project Detailsresource-005-
ContributorBridge2AI Programresource-006-
ContributorAIM-AHEAD Training Partnershipresource-007-
ContributorOHDSI Communityresource-008-
ContributorPublished Researchresource-009-
ContributorContact for Accessresource-010-
🚀

Uses

What (other) tasks could the dataset be used for?

DescriptionIDName
Generate data for ML/AI applications aimed at characterizing acute and critical care illness pattern...task-001Characterize acute and critical care illness
Enable prediction of complications among patients with acute or critical illness using multi-modal d...task-002Predict complications in critically ill patients
Support measurement and analysis of treatment response among critically ill patients through high-fr...task-003Measure treatment response
Provision a holdout test set accessible for model external validation to aid marketplace adoption of...task-004External validation for marketplace adoption
Utilize visualization and annotation environment to label data with targets important for prediction...task-005Label data for prediction targets
DescriptionIDName
Primary intended use is development and training of artificial intelligence and machine learning mod...use-001AI/ML model development for critical care
Provision of holdout test set accessible for model external validation to aid marketplace adoption o...use-002External validation of AI models
Studies examining health equity, social determinants of health, and disparities in critical care out...use-003Research on health equity and disparities
Research aimed at improving clinical care delivery, treatment protocols, and patient outcomes in acu...use-004Clinical care improvement research
Training and education of next generation of diverse academic and community AI scientists through ha...use-005Educational and training purposes
DescriptionIDName
As data collection continues through November 2026 and quality assurance processes are ongoing, earl...discouraged-001Uses during ongoing data collection
Dataset is for research purposes. Any AI/ML models developed should undergo appropriate clinical val...discouraged-002Clinical decision-making without validation
Attempts to re-identify patients from de-identified data violate ethical principles, data use agreem...discouraged-003Re-identification attempts
ID
license-001
Name
CHoRUS Controlled Access License
Description
Dataset distributed under controlled access requiring institutional email registration and signed licensing agreement. Access granted after review and approval process. Participants must complete registration form with name, email (institutional, not personal), and institution. Once approved, users receive email with access instructions to CHoRUS secure enclave. Contact for access requests: dbold@emory.edu or jared.houghtaling@tuftsmedicine.org.
License Terms
  • Institutional email required for registration
  • Signed licensing agreement required
  • Controlled access through secure enclave
  • Data use agreement specifies permitted uses
  • Prohibition on re-identification attempts
  • Compliance with ethical and legal requirements
📤

Distribution

How will the dataset be distributed?

Controlled Access with Data Use Agreement
🔄

Maintenance

How will the dataset be maintained?

ID
updates-001
Name
Ongoing data collection and expansion
Description
Dataset updated continuously as data collection progresses at 14 acquisition centers. As of November 2024, covers 14 hospitals with 23,400 unique admissions. Target exceeds 100,000 critically ill patients. Project timeline extends through November 30, 2026 (with approved no-cost extension). Regular status updates tracked through GitHub project management system. Sites provide updates via GitHub interface or Google Form submissions.
Frequency
Continuous updates through November 2026
Update Details
  • Ongoing retrospective data collection at 14 sites
  • Current status: 23,400 unique admissions (as of November 2024)
  • Target: 100,000+ critically ill patients
  • Regular site status updates via GitHub/Google Forms
  • GitHub project tracking for deliverables
  • Documentation updates in Chorus_SOP repository
  • Software and tooling continuous development
  • Semantic mapping validation and expansion
👥

Human Subjects

Does the dataset relate to people?

ID
hsr-001
Name
CHoRUS Human Subjects Research
Description
Retrospective data collection from critically ill patients approved through institutional review processes. Community-facing ethics focus groups conducted to determine what data is appropriate for public sharing. Legal framework established for collecting data at scale. Patient-focused efforts determine ethical and legal approaches to manage privacy and bias while accounting for Social Determinants of Health. Project draws expertise from law, ethics, health services, biomedical science, engineering, and scientific journal publications disciplines.
Involves Human Subjects
True
Ethics Review Board
  • Institutional review boards at 14 data acquisition centers
  • Community-facing ethics focus groups
  • Legal and ethical advisory teams
  • Privacy and accountability review processes
Generated on 2025-12-20 19:23:28 using Bridge2AI Data Sheets Schema