CHORUS Dataset Documentation

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

DescriptionIDName
Answer the grand challenge of improving recovery from acute illness by developing high-resolution multi-center datasets as a critical first step towards actionable and trustworthy AI in critical care. Address the urgent need for infrastructure to support artificial intelligence and machine learning (AI/ML) in critical care settings.
purpose-001Improve recovery from acute illness
Develop a publicly available, AI-ready critical care dataset from more than 100,000 critically ill patients while ensuring methods promote privacy, accountability, and clinical benefit. Generate the most diverse, high-resolution, ethically sourced dataset for AI/ML applications in acute and critical care.
purpose-002Create AI-ready critical care dataset
Unify standards to harmonize multi-modal EHR, waveform, imaging, and text data. Develop software and tooling to interact with and extract insight from clinical data in diverse formats. Create validated semantic mappings for connecting clinical data in various source formats to international standards (OMOP Common Data Model).
purpose-003Establish data standards and tools
Ensure comprehensive sets of patient conditions and clinical treatment strategies with appropriate contextual factors such as geographic distance to nearest hospital and Social Determinants of Health. Develop the skills and workforce for a next generation of diverse academic and community AI scientists through comprehensive training and education programs.
purpose-004Promote diversity and health equity
  • ID
    funder-001
    Name
    NIH Common Fund Bridge2AI Program
    Description
    Funded through National Institutes of Health grant 1OT2OD032701-01, administered by NIH Office of the Director. Opportunity Number: OTA-21-008. Study Section: Data Coordination, Mapping, and Modeling [DCMM]. Fiscal Year 2022. Total funding in 2022: $5,880,300 (all direct costs). Project dates: September 1, 2022 to November 30, 2026 (with no-cost extension approved).
📊

Composition

What do the instances represent?

  1. ID
    instance-001
    Name
    Critically ill patients
    Description
    Individual critically ill patients requiring acute or critical care admitted to intensive care units or similar hospital settings across 14 data acquisition centers. As of November 2024, dataset covers 14 different hospitals with 23,400 unique admissions. Target enrollment exceeds 100,000 critically ill patients. Retrospective data collection from patients with acute or critical illness.
    Instance Type
    Human subjects - critically ill patients admitted to participating hospitals between data collection timeframe (specific dates vary by site, ongoing collection through November 2026).
DescriptionIDName
Patients distributed across 14 different hospitals within the CHoRUS network, ensuring geographic and institutional diversity in critical care settings.
subpop-001Critically ill patients by hospital
Subset of patients with complete data across multiple modalities: structured EHR, waveform telemetry, imaging, clinical notes, and EEG (availability varies).
subpop-002Patients with complete multi-modal data
Access UrlsDescriptionIDName
https://chorus4ai.org/
Structured electronic health record data distributed in OMOP Common Data Model format. Enables use of OHDSI tool stack for analysis. Available in secure enclave with controlled access.
format-001OMOP Common Data Model
https://chorus4ai.org/
Waveform telemetry data distributed in WFDB (WaveForm DataBase) format following extended PhysioNet schema. Available in secure enclave with controlled access.
format-002WFDB waveform format
https://chorus4ai.org/
Medical imaging data distributed in DICOM format with comprehensive metadata. De-identification in process, planned for controlled access in secure enclave.
format-003DICOM imaging format
https://chorus4ai.org/
Clinical notes distributed as OHNLP-tokenized text following open source schema. Stored locally at sites except tokens, planned for controlled access.
format-004OHNLP tokenized text
https://chorus4ai.org/
EEG waveform data distributed in EDF+ (European Data Format) and Persyst formats following open source schemas. Extraction in process, planned for controlled access.
format-005EDF+ and Persyst EEG formats
🔍

Collection Process

How was the data acquired?

CHoRUS
Patient-Focused Collaborative Hospital Repository Uniting Standards (CHoRUS) for Equitable AI
CHoRUS for Equitable AI is a Bridge2AI data generation project developing the most diverse, high-resolution, ethically sourced, AI-ready critical care dataset to answer the grand challenge of improving recovery from acute illness. The project spans 20 academic centers (14 data acquisition centers) and creates a publicly available dataset of over 100,000 critically ill patients with multi-modal data including structured EHR, waveform telemetry, medical imaging, EEG, and clinical notes. All data is standardized to the OMOP Common Data Model with additional formats (DICOM, WFDB, OHNLP tokenization) and includes comprehensive metadata schemas. Patient-focused efforts determine ethical and legal approaches to manage privacy and bias while accounting for Social Determinants of Health. A visualization and annotation environment labels data with targets important for prediction. The project emphasizes skills and workforce development for a next generation of diverse academic and community AI scientists through training programs and partnerships with AIM-AHEAD. As of November 2024, the dataset covers 14 different hospitals with 23,400 unique admissions.
en
  • CHoRUS
  • Bridge2AI
  • critical care
  • acute illness
  • AI-ready dataset
  • OMOP Common Data Model
  • electronic health records
  • EHR
  • waveform telemetry
  • medical imaging
  • DICOM
  • EEG
  • clinical notes
  • OHNLP
  • health equity
  • social determinants of health
  • federated access
  • multi-modal data
  • high-resolution data
  • ethical AI
  • trustworthy AI
  • workforce development
  • data standardization
  • privacy preservation
  • bias mitigation
DescriptionIDName
Address the absence of large-scale, diverse, high-resolution multi-center datasets for critical care AI/ML by creating a dataset spanning 20 academic centers with 100,000+ critically ill patients and ensuring balanced, diverse cohorts through federated access and sampling methods.
gap-001Lack of diverse multi-center critical care datasets
Overcome the lack of unified standards in critical care data by harmonizing multi-modal EHR, waveform, imaging, and text data to OMOP Common Data Model and other international standards (DICOM, WFDB, OHNLP).
gap-002Insufficient data standardization
Address ethical and legal challenges in AI through patient-focused efforts that determine approaches to manage privacy and bias while accounting for Social Determinants of Health and performing community-facing ethics focus groups.
gap-003Privacy and bias concerns in AI
Bridge the gap in diverse AI/ML workforce through comprehensive educational approaches, training programs (including AIM-AHEAD partnership), and cultivation of expertise in lay and scientific communities to improve AI literacy and utilization.
gap-004Limited AI workforce diversity
RoleNameORCIDAffiliation
ContributorEric S. Rosenthalcreator-001-
ContributorAzra Bihoraccreator-002-
ContributorAshley Cordescreator-003-
ContributorGilles Clermontcreator-004-
ContributorGari David Cliffordcreator-005-
ContributorBarbara J. Evanscreator-006-
ContributorXiao Hucreator-007-
ContributorRishikesan Kamaleswarancreator-008-
ContributorYulia A. Levites Strekalovacreator-009-
ContributorParisa Rashidicreator-010-
ContributorCynthia Rudincreator-011-
ContributorIshan Canty Williamscreator-012-
ContributorAndrew Ewing Williamscreator-013-
ContributorXiaoqian Jiangcreator-014-
ContributorMorteza Zabihicreator-015-
ContributorZhenhong Hucreator-016-
ContributorDebora Simmonscreator-017-
ContributorAndrew Williamscreator-018-
ContributorAliyah Geercreator-019-
DescriptionIDName
Primary dataset with controlled access requiring data use agreement and licensing. Includes OMOP-standardized structured EHR data, waveform telemetry, and associated metadata. As of November 2024, OMOP and telemetry data available in secure enclave. Registration required with institution email, and all participants must sign licensing agreement.
subset-001Controlled Access Dataset
Clinical notes extracted and tokenized using OHNLP (Open Health Natural Language Processing) toolkit. Stored locally at contributing sites (except tokens) with planned controlled access. Uses OHNLP open source schema for standardization.
subset-002Clinical Notes (Tokenized)
Imaging data from PACS (Picture Archiving and Communication System) in DICOM format. De-identification in process as of November 2024, planned for controlled access. DICOM schema metadata available.
subset-003Medical Imaging
Electroencephalography (EEG) waveform data from hospital databases. Stored in EDF+ (European Data Format) and Persyst formats. Extraction in process as of November 2024, planned for controlled access. Open source EDF+ and Persyst schema available.
subset-004EEG Waveforms
Subsets of the dataset currently being used for training activities and publications as the dataset undergoes expansion and quality assurance processes.
subset-005Training and Publication Datasets
  1. ID
    sampling-001
    Name
    Multi-center federated sampling
    Description
    Federated access enables sampling methods to ensure a balanced and diverse cohort across 20 academic centers (14 data acquisition centers). Legal framework established for collecting data at scale with sampling to ensure comprehensive sets of patient conditions and clinical treatment strategies.
    Sample
    True
    Random Sampling
    False
    Representative Sample
    True
    Strategies
    • Federated multi-center data collection across 14 hospitals
    • Sampling for balanced and diverse patient populations
    • Inclusion of contextual factors (geographic distance to hospital, Social Determinants of Health)
    • Community-facing ethics focus groups to determine appropriate data for public sharing
    • Legal framework for data collection at scale
DescriptionIDName
Retrospective data collection from electronic health record systems at 14 data acquisition centers. Data extracted includes demographics, medication administration (dosing time-stamped upon each infusion change or dose administration), procedures, nursing flowsheets (high-frequency documentation), diagnoses, and clinical notes.
collection-001Retrospective EHR extraction
Continuous waveform telemetry data captured from bedside monitors through gateway/middleware systems. Stored in WFDB (WaveForm DataBase) format following PhysioNet schema (extended).
collection-002Waveform telemetry capture
Medical imaging data acquired from hospital PACS (Picture Archiving and Communication System) and stored in DICOM format with comprehensive metadata following DICOM schema.
collection-003Medical imaging acquisition
Electroencephalography recordings extracted from hospital EEG databases in EDF+ and Persyst formats with metadata following open source schemas.
collection-004EEG recording extraction
Multi-center network capabilities to acquire, standardize, tokenize, store, visualize, and label data. All structured clinical data transformed to OMOP Common Data Model. Clinical notes tokenized using OHNLP toolkit. Imaging converted to DICOM. Waveforms standardized to WFDB format.
collection-005Standardized data transformation
DescriptionIDName
Demographics, medication administration, procedures, nursing flowsheets, and diagnoses acquired from electronic health records and standardized to OMOP Common Data Model. Controlled access with published OMOP schema metadata.
acquisition-001Structured EHR data (OMOP)
Dosing information time-stamped upon each infusion change or dose administration. Stored in OMOP format with controlled access and OMOP schema metadata.
acquisition-002Medication administration records
Procedures and diagnoses documented by healthcare providers. Stored in OMOP format with controlled access and OMOP schema metadata.
acquisition-003Provider documentation
Nursing documentation captured at high frequency intervals. Stored in OMOP format with controlled access and OMOP schema with extensions for high-frequency data.
acquisition-004High-frequency nursing flowsheets
Clinical notes extracted and tokenized using OHNLP (Open Health Natural Language Processing) toolkit. Stored locally except tokens. Controlled access planned with OHNLP open source schema metadata.
acquisition-005Clinical notes (tokenized)
Imaging data from PACS systems in DICOM format. De-identification in process. Planned controlled access with published DICOM schema metadata.
acquisition-006Medical imaging (DICOM)
Bedside monitor waveform data acquired through gateway/middleware systems. Stored in WFDB format with controlled access and published PhysioNet schema (extended) metadata.
acquisition-007Waveform telemetry (WFDB)
EEG recordings from hospital databases in EDF+ and Persyst formats. Extraction in process. Planned controlled access with open source EDF+ and Persyst schema metadata.
acquisition-008EEG waveforms
DescriptionIDNamePreprocessing Details
All structured electronic health record data standardized to the OMOP (Observational Medical Outcomes Partnership) Common Data Model. Ensures interoperability and enables use of OHDSI (Observational Health Data Sciences and Informatics) tool stack for analysis.
preproc-001OMOP Common Data Model transformationTransformation of source EHR data to OMOP CDM, Standardization of terminology and codes, Mapping to OMOP vocabulary standards, Quality assurance of transformed data, Integration with OHDSI tools for analysis
Clinical notes processed using OHNLP (Open Health Natural Language Processing) toolkit for extraction and tokenization. Protects patient privacy while enabling natural language processing and analysis.
preproc-002Clinical note tokenizationText extraction from clinical notes, OHNLP tokenization pipeline, De-identification of sensitive information, Standardization to OHNLP schema, Local storage (except tokens)
Waveform telemetry data from diverse bedside monitoring systems standardized to WFDB (WaveForm DataBase) format following extended PhysioNet schema.
preproc-003Waveform standardizationConversion from proprietary monitor formats, Standardization to WFDB format, Metadata extraction and schema compliance, Quality checks for waveform integrity, Synchronization with clinical events
DICOM imaging data undergoing de-identification process to remove patient identifiable information while preserving clinical utility and metadata.
preproc-004Medical imaging de-identificationDICOM header de-identification, Removal of embedded patient information, Preservation of clinically relevant metadata, DICOM schema compliance verification, Quality assurance of de-identified images
Data transformed using approaches that limit re-identification while maintaining analytical utility. Multiple preprocessing strategies employed to protect patient privacy across all data modalities.
preproc-005Data re-identification limitationApplication of de-identification algorithms, Privacy-preserving transformations, Risk assessment for re-identification, Compliance with ethical and legal requirements, Community ethics focus group input
Cleaning DetailsDescriptionIDName
Semantic mapping validation by clinical experts, Standard operating protocol implementation, Cross-site data quality checks, Resolution of institutional variations, Clinical validation SOP for mappings
Data from 14 acquisition centers harmonized through validated semantic mappings and standard operating protocols (SOPs). Ensures consistency and interoperability across diverse institutional EHR systems and clinical practices.
cleaning-001Multi-center data harmonization
Schema compliance verification, Metadata completeness checks, Validation against published standards, Error detection and correction, Documentation of schema extensions
All data modalities validated against published metadata schemas (OMOP, DICOM, WFDB, OHNLP, EDF+, Persyst) to ensure compliance and data quality.
cleaning-002Metadata schema validation
  1. ID
    labeling-001
    Name
    Visualization and annotation environment
    Description
    Custom visualization and annotation environment developed to label data with targets important for prediction tasks in critical care AI applications.
    Labeling Details
    • Interactive visualization tools
    • Annotation interface for clinical experts
    • Labeling of prediction targets
    • Quality control of annotations
    • Documentation of labeling protocols
DescriptionIDMaintainer DetailsName
Multi-institutional consortium managing dataset maintenance including data acquisition centers, coordinating teams, and infrastructure development.
maintainer-001Massachusetts General Hospital (lead institution), University of Florida (data acquisition and coordination), UT Health Science Center (data acquisition and coordination), Tufts Medicine (data acquisition and coordination), 14 data acquisition centers across United States, Standards team (semantic mappings and validation), Data Acquisition team (extraction and contribution), Tooling team (software development), Project management via GitHub organization (chorus-ai)CHoRUS Consortium
Active GitHub organization (chorus-ai) housing repositories for software, semantic mappings, standard operating protocols, and project management. 28 repositories with comprehensive documentation and community support.
maintainer-002GitHub organization: https://github.com/chorus-ai, Chorus_SOP repository (centralized documentation), chorus-container-apps (Azure deployment), chorus-mapping (semantic mappings), chorus_waveform (waveform tools), privacy_scan_tool (privacy protection), Community discussions and issue trackingCHoRUS GitHub Organization
ID
retention-001
Name
Long-term dataset retention
Description
Digital data maintained according to NIH data sharing policies and institutional requirements. Controlled access model ensures long-term availability for research while protecting patient privacy.
Retention Details
  • NIH data sharing policies govern retention
  • Institutional requirements at participating centers
  • Controlled access model for privacy protection
  • Secure enclave infrastructure for data storage
  • Long-term maintenance through CHoRUS Consortium
DescriptionIDNameSensitive Elements PresentSensitivity Details
Complete electronic health records including demographics, diagnoses, procedures, medications, nursing documentation, and clinical notes for critically ill patients. Contains protected health information subject to HIPAA and institutional privacy requirements.
sensitive-001Clinical and medical dataTrueDemographics and patient identifiers (de-identified), Diagnoses and medical conditions, Medication administration records, Procedures and interventions, Clinical notes (tokenized for privacy), Nursing flowsheet documentation
Continuous waveform telemetry and EEG recordings capturing detailed physiological states of critically ill patients. May reveal sensitive health conditions and treatment responses.
sensitive-002Physiological monitoring dataTrueContinuous cardiac waveforms, Respiratory monitoring data, Hemodynamic measurements, Electroencephalography recordings, High-frequency physiological parameters
Diagnostic imaging studies in DICOM format. Undergoing de-identification to remove embedded patient information while preserving clinical utility.
sensitive-003Medical imagingTrueCT scans and X-rays, MRI and other imaging modalities, DICOM metadata (de-identified), Embedded patient information (removed)
Contextual factors including geographic information (distance to hospital) and social determinants of health data. Collected to support health equity research while maintaining patient privacy.
sensitive-004Social Determinants of HealthTrueGeographic location information, Social determinants of health variables, Contextual factors for equity research, Privacy-preserving transformations applied
RoleNameORCIDAffiliation
ContributorCHoRUS Project Websiteresource-001-
ContributorCHoRUS GitHub Organizationresource-002-
ContributorChorus_SOP Documentationresource-003-
ContributorCHoRUS Developer Documentationresource-004-
ContributorNIH RePORTER Project Detailsresource-005-
ContributorBridge2AI Programresource-006-
ContributorAIM-AHEAD Training Partnershipresource-007-
ContributorOHDSI Communityresource-008-
ContributorPublished Researchresource-009-
ContributorContact for Accessresource-010-
🚀

Uses

What (other) tasks could the dataset be used for?

DescriptionIDName
Generate data for ML/AI applications aimed at characterizing acute and critical care illness patterns, progression, and outcomes across diverse patient populations and hospital settings.
task-001Characterize acute and critical care illness
Enable prediction of complications among patients with acute or critical illness using multi-modal data including structured EHR, waveforms, imaging, and clinical notes.
task-002Predict complications in critically ill patients
Support measurement and analysis of treatment response among critically ill patients through high-frequency documentation, medication administration records, and clinical outcomes data.
task-003Measure treatment response
Provision a holdout test set accessible for model external validation to aid marketplace adoption of AI-developed models for implementation in acute and critical care settings.
task-004External validation for marketplace adoption
Utilize visualization and annotation environment to label data with targets important for prediction tasks in critical care AI applications.
task-005Label data for prediction targets
DescriptionIDName
Primary intended use is development and training of artificial intelligence and machine learning models to characterize acute and critical care illness, predict complications, and measure treatment response in critically ill patients.
use-001AI/ML model development for critical care
Provision of holdout test set accessible for model external validation to aid marketplace adoption of AI-developed models for implementation in acute and critical care settings.
use-002External validation of AI models
Studies examining health equity, social determinants of health, and disparities in critical care outcomes across diverse patient populations and hospital settings.
use-003Research on health equity and disparities
Research aimed at improving clinical care delivery, treatment protocols, and patient outcomes in acute and critical care settings through data-driven insights.
use-004Clinical care improvement research
Training and education of next generation of diverse academic and community AI scientists through hands-on experience with real-world critical care datasets. Integration with AIM-AHEAD training programs.
use-005Educational and training purposes
DescriptionIDName
As data collection continues through November 2026 and quality assurance processes are ongoing, early dataset versions should be used with awareness of completeness limitations and ongoing expansion.
discouraged-001Uses during ongoing data collection
Dataset is for research purposes. Any AI/ML models developed should undergo appropriate clinical validation, regulatory approval, and institutional review before use in patient care or clinical decision-making.
discouraged-002Clinical decision-making without validation
Attempts to re-identify patients from de-identified data violate ethical principles, data use agreements, and legal frameworks established for privacy protection.
discouraged-003Re-identification attempts
ID
license-001
Name
CHoRUS Controlled Access License
Description
Dataset distributed under controlled access requiring institutional email registration and signed licensing agreement. Access granted after review and approval process. Participants must complete registration form with name, email (institutional, not personal), and institution. Once approved, users receive email with access instructions to CHoRUS secure enclave. Contact for access requests: dbold@emory.edu or jared.houghtaling@tuftsmedicine.org.
License Terms
  • Institutional email required for registration
  • Signed licensing agreement required
  • Controlled access through secure enclave
  • Data use agreement specifies permitted uses
  • Prohibition on re-identification attempts
  • Compliance with ethical and legal requirements
📤

Distribution

How will the dataset be distributed?

Controlled Access with Data Use Agreement
🔄

Maintenance

How will the dataset be maintained?

ID
updates-001
Name
Ongoing data collection and expansion
Description
Dataset updated continuously as data collection progresses at 14 acquisition centers. As of November 2024, covers 14 hospitals with 23,400 unique admissions. Target exceeds 100,000 critically ill patients. Project timeline extends through November 30, 2026 (with approved no-cost extension). Regular status updates tracked through GitHub project management system. Sites provide updates via GitHub interface or Google Form submissions.
Frequency
Continuous updates through November 2026
Update Details
  • Ongoing retrospective data collection at 14 sites
  • Current status: 23,400 unique admissions (as of November 2024)
  • Target: 100,000+ critically ill patients
  • Regular site status updates via GitHub/Google Forms
  • GitHub project tracking for deliverables
  • Documentation updates in Chorus_SOP repository
  • Software and tooling continuous development
  • Semantic mapping validation and expansion
👥

Human Subjects

Does the dataset relate to people?

ID
hsr-001
Name
CHoRUS Human Subjects Research
Description
Retrospective data collection from critically ill patients approved through institutional review processes. Community-facing ethics focus groups conducted to determine what data is appropriate for public sharing. Legal framework established for collecting data at scale. Patient-focused efforts determine ethical and legal approaches to manage privacy and bias while accounting for Social Determinants of Health. Project draws expertise from law, ethics, health services, biomedical science, engineering, and scientific journal publications disciplines.
Involves Human Subjects
True
Ethics Review Board
  • Institutional review boards at 14 data acquisition centers
  • Community-facing ethics focus groups
  • Legal and ethical advisory teams
  • Privacy and accountability review processes
Generated on 2026-04-21 15:34:02 using Bridge2AI Data Sheets Schema