CHORUS d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

DescriptionIDName
Answer the grand challenge of improving recovery from acute illness by developing high-resolution multi-center datasets as a critical first step towards actionable and trustworthy AI in critical care. Address the urgent need for infrastructure to support artificial intelligence and machine learning (AI/ML) in critical care settings.
chorus:purpose:1Improve recovery from acute illness
Develop a publicly available, AI-ready critical care dataset from more than 100,000 critically ill patients while ensuring methods promote privacy, accountability, and clinical benefit. Generate the most diverse, high-resolution, ethically sourced dataset for AI/ML applications in acute and critical care, expanding AI and Machine Learning to improve recovery from acute illness.
chorus:purpose:2Create AI-ready critical care dataset
Unify standards to harmonize multi-modal EHR, waveform, imaging, and text data. Develop software and tooling to interact with and extract insight from clinical data in diverse formats. Create validated semantic mappings for connecting clinical data in various source formats to international standards (OMOP Common Data Model, DICOM, WFDB, OHNLP).
chorus:purpose:3Establish data standards and tools
Ensure comprehensive sets of patient conditions and clinical treatment strategies with appropriate contextual factors such as geographic distance to nearest hospital and Social Determinants of Health. Develop the skills and workforce for a next generation of diverse academic and community AI scientists through comprehensive training and education programs in partnership with AIM-AHEAD.
chorus:purpose:4Promote diversity and health equity
  • ID
    chorus:funder:1
    Name
    NIH Common Fund Bridge2AI Program
    Description
    Funded through National Institutes of Health grant OT2OD032701 (project number 1OT2OD032701-01), administered by NIH Office of the Director. Opportunity Number: OTA-21-008. Study Section: Data Coordination, Mapping, and Modeling (DCMM). Fiscal Year 2022. Total funding in 2022: $5,880,300 (all direct costs). Project dates: September 1, 2022 to November 30, 2026 (with approved no-cost extension). Award notice date: September 1, 2022. Assistance Listing Number: 93.310.
📊

Composition

What do the instances represent?

  1. ID
    chorus:instance:1
    Name
    Critically ill patients in ICU settings
    Description
    Individual critically ill patients requiring acute or critical care admitted to intensive care units (ICU), pediatric intensive care units (PICU), or neonatal intensive care units (NICU) across 14 data acquisition centers. As of August 2025, the dataset covers 14 different hospitals with over 45,000 unique admissions, with 50,000 patient admissions from ICU, PICU, and NICU currently released. Target enrollment exceeds 100,000 critically ill patients. Retrospective data collection from patients with acute or critical illness.
    Instance Type
    Human subjects - critically ill patients admitted to participating hospitals (retrospective data collection, ongoing collection through November 2026).
DescriptionIDName
Patients distributed across 14 different hospitals within the CHoRUS network spanning 20 academic centers in the United States, ensuring geographic and institutional diversity in critical care settings.
chorus:subpop:1Critically ill patients by hospital
Patients admitted to intensive care units (ICU), pediatric intensive care units (PICU), and neonatal intensive care units (NICU) across contributing hospitals, with 50,000 patient admissions in the current released dataset.
chorus:subpop:2ICU, PICU, and NICU patients
Access UrlsDescriptionIDName
https://chorus4ai.org/
Structured electronic health record data distributed in OMOP Common Data Model format enabling use of OHDSI tool stack for analysis. Contains 1.6 billion rows. Available in secure enclave with controlled access. Published OMOP schema metadata available.
chorus:format:1OMOP Common Data Model (structured EHR data)
https://chorus4ai.org/
Waveform telemetry data (23 TB) distributed in WFDB (WaveForm DataBase) format following extended PhysioNet schema. Available in secure enclave with controlled access. Published PhysioNet schema (extended) metadata available.
chorus:format:2WFDB waveform format (bedside monitor telemetry)
https://chorus4ai.org/
Medical imaging data (7,642 admissions with radiology; 1,000 images currently available) distributed in DICOM format with comprehensive metadata following DICOM schema. De-identification in process for larger cohort. Planned controlled access.
chorus:format:3DICOM format (medical imaging)
https://chorus4ai.org/Clinical notes distributed as OHNLP-tokenized text following open source OHNLP schema. Stored locally at contributing sites (except tokens). Controlled access planned. chorus:format:4OHNLP tokenized format (clinical notes)
https://chorus4ai.org/EEG waveform data distributed in EDF+ (European Data Format) and Persyst formats following open source schemas. Extraction in process, planned for controlled access. chorus:format:5EDF+ and Persyst formats (EEG waveforms)
🔍

Collection Process

How was the data acquired?

CHoRUS
Patient-Focused Collaborative Hospital Repository Uniting Standards (CHoRUS) for Equitable AI
CHoRUS for Equitable AI is a Bridge2AI data generation project developing the most diverse, high-resolution, ethically sourced, AI-ready critical care dataset to answer the grand challenge of improving recovery from acute illness. The project spans 20 academic centers (14 data acquisition centers) and is building a publicly available dataset targeting over 100,000 critically ill patients with multi-modal data including structured EHR, waveform telemetry, medical imaging, EEG, and clinical notes. All structured data is standardized to the OMOP Common Data Model with additional formats (DICOM, WFDB, OHNLP tokenization) and comprehensive metadata schemas. Patient-focused efforts determine ethical and legal approaches to manage privacy and bias while accounting for Social Determinants of Health. A visualization and annotation environment labels data with targets important for prediction. The project emphasizes skills and workforce development for a next generation of diverse academic and community AI scientists through training programs and partnerships with AIM-AHEAD. As of August 2025, the dataset covers 14 different hospitals with over 45,000 unique admissions and includes 50,000 patient admissions from ICU, PICU, and NICU, 1.6 billion rows of EHR OMOP data, 7,642 admissions with radiology data, and 23 TB of waveform data.
en
  • CHoRUS
  • Bridge2AI
  • critical care
  • acute illness
  • AI-ready dataset
  • OMOP Common Data Model
  • electronic health records
  • EHR
  • waveform telemetry
  • medical imaging
  • DICOM
  • EEG
  • clinical notes
  • OHNLP
  • health equity
  • social determinants of health
  • federated access
  • multi-modal data
  • high-resolution data
  • ethical AI
  • trustworthy AI
  • workforce development
  • data standardization
  • privacy preservation
  • bias mitigation
  • ICU
  • intensive care
DescriptionIDName
Address the absence of large-scale, diverse, high-resolution multi-center datasets for critical care AI/ML by creating a dataset spanning 20 academic centers with 100,000+ critically ill patients and ensuring balanced, diverse cohorts through federated access and sampling methods across 14 data acquisition centers.
chorus:gap:1Lack of diverse multi-center critical care datasets
Overcome the lack of unified standards in critical care data by harmonizing multi-modal EHR, waveform, imaging, and text data to OMOP Common Data Model and other international standards (DICOM, WFDB, OHNLP).
chorus:gap:2Insufficient data standardization in critical care
Address ethical and legal challenges in AI through patient-focused efforts that determine approaches to manage privacy and bias while accounting for Social Determinants of Health and performing community-facing ethics focus groups to determine what data is appropriate for public sharing.
chorus:gap:3Privacy and bias concerns in clinical AI
Bridge the gap in diverse AI/ML workforce through comprehensive educational approaches, training programs (including AIM-AHEAD partnership), and cultivation of expertise in lay and scientific communities to improve AI literacy and utilization.
chorus:gap:4Limited diversity in AI/ML workforce
RoleNameORCIDAffiliation
ContributorEric S. Rosenthalchorus:creator:1-
ContributorAzra Bihoracchorus:creator:2-
ContributorAshley Cordeschorus:creator:3-
ContributorGilles Clermontchorus:creator:4-
ContributorGari David Cliffordchorus:creator:5-
ContributorBarbara J. Evanschorus:creator:6-
ContributorXiao Huchorus:creator:7-
ContributorRishikesan Kamaleswaranchorus:creator:8-
ContributorYulia A. Levites Strekalovachorus:creator:9-
ContributorParisa Rashidichorus:creator:10-
ContributorCynthia Rudinchorus:creator:11-
ContributorIshan Canty Williamschorus:creator:12-
ContributorAndrew Ewing Williamschorus:creator:13-
ContributorXiaoqian Jiangchorus:creator:14-
ContributorMorteza Zabihichorus:creator:15-
ContributorZhenhong Huchorus:creator:16-
ContributorDebora Simmonschorus:creator:17-
ContributorAliyah Geerchorus:creator:18-
ContributorCiera McCrarychorus:creator:19-
DescriptionIDName
Primary dataset with controlled access requiring data use agreement and licensing. Includes OMOP-standardized structured EHR data comprising demographics, medication administration (dosing time-stamped upon each infusion change or dose administration), procedures, nursing flowsheets (high-frequency documentation), and diagnoses. Available in secure enclave. As of August 2025, contains 1.6 billion rows of EHR OMOP data. Registration required with institutional email; all participants must sign licensing agreement.
chorus:subset:1Controlled Access Structured EHR Dataset (OMOP)
Waveform telemetry data from bedside monitors (gateway/middleware) standardized to WFDB (WaveForm DataBase) format following extended PhysioNet schema. Available in secure enclave with controlled access. As of August 2025, contains 23 TB of waveform data. Published metadata schema: PhysioNet schema (extended).
chorus:subset:2Waveform Telemetry Dataset (WFDB)
Imaging data from hospital PACS (Picture Archiving and Communication System) in DICOM format with comprehensive metadata following DICOM schema. As of August 2025, 1,000 images are available with de-identification in process for the larger cohort. 7,642 admissions have radiology data. Controlled access planned.
chorus:subset:3Medical Imaging Dataset (DICOM)
Clinical notes extracted and tokenized using OHNLP (Open Health Natural Language Processing) toolkit following open source OHNLP schema. Stored locally at contributing sites (except tokens). Controlled access planned. De-identification achieved through tokenization approach.
chorus:subset:4Clinical Notes Dataset (OHNLP Tokenized)
Electroencephalography (EEG) waveform data from hospital databases in EDF+ (European Data Format) and Persyst formats. Open source EDF+ and Persyst schema metadata available. Extraction in process, planned for controlled access.
chorus:subset:5EEG Waveform Dataset (EDF+ and Persyst)
Subsets of the dataset currently being used for training activities (AIM-AHEAD Bridge2AI for Clinical Care Training Program) and publications as the dataset undergoes expansion and quality assurance processes.
chorus:subset:6Training and Publication Subset
  1. ID
    chorus:sampling:1
    Name
    Multi-center federated sampling
    Description
    Federated access enables sampling methods to ensure a balanced and diverse cohort across 20 academic centers (14 data acquisition centers). Legal framework established for collecting data at scale with sampling to ensure comprehensive sets of patient conditions and clinical treatment strategies. Community-facing ethics focus groups conducted to determine what data is appropriate for public sharing.
    Sample
    True
    Random Sampling
    False
    Representative Sample
    True
    Source Data
    • Critically ill patients in ICU, PICU, and NICU settings at 14 data acquisition centers across 20 academic centers in the United States
    Representative Verification
    • Federated multi-center data collection across 14 hospitals ensures geographic and institutional diversity; sampling for balanced and diverse patient populations; inclusion of Social Determinants of Health contextual factors
DescriptionIDNameSensitive Elements PresentSensitivity Details
Complete electronic health records including demographics, diagnoses, procedures, medications, nursing documentation, and clinical notes for critically ill patients. Contains protected health information subject to HIPAA and institutional privacy requirements. De-identified before release through OMOP transformation and privacy scanning tools.
chorus:sensitive:1Protected health information - clinical EHR dataTrueDemographics and patient identifiers (de-identified before controlled access release), Diagnoses and medical conditions (OMOP standardized), Medication administration records with dosing timestamps (OMOP standardized), Procedures and clinical interventions documented by providers (OMOP standardized), Clinical notes (tokenized using OHNLP toolkit for privacy protection), Nursing flowsheet documentation at high frequency (OMOP with extensions)
Continuous waveform telemetry from bedside monitors and EEG recordings capturing detailed physiological states of critically ill patients. May reveal sensitive health conditions and treatment responses. Stored in WFDB format (telemetry) and EDF+/Persyst formats (EEG) with controlled access.
chorus:sensitive:2Physiological monitoring and EEG waveform dataTrueContinuous cardiac waveforms from bedside monitors (WFDB format, 23 TB), Respiratory monitoring and hemodynamic measurement data, Electroencephalography (EEG) recordings from hospital databases, High-frequency physiological parameters
Diagnostic imaging studies in DICOM format from hospital PACS systems. De-identification in process to remove embedded patient information while preserving clinical utility. As of August 2025, 1,000 images available with de-id in process for larger cohort (7,642 admissions with radiology data total).
chorus:sensitive:3Medical imaging dataTrueRadiology images (CT scans, X-rays, MRI and other modalities) in DICOM format, DICOM metadata (de-identification in process), Embedded patient information being removed through de-identification pipeline
Contextual factors including geographic information (distance to hospital) and social determinants of health data. Collected to support health equity research while maintaining patient privacy through privacy-preserving transformations.
chorus:sensitive:4Social Determinants of Health dataTrueGeographic location information (distance to nearest hospital), Social determinants of health variables, Contextual equity factors with privacy-preserving transformations applied
DescriptionIDName
Retrospective data collection from electronic health record systems at 14 data acquisition centers. Data extracted includes demographics, medication administration (dosing time-stamped upon each infusion change or dose administration), procedures, nursing flowsheets (high-frequency documentation), diagnoses, and clinical notes. Standardized to OMOP Common Data Model.
chorus:collection:1Retrospective EHR extraction
Continuous waveform telemetry data captured from bedside monitors through gateway and middleware systems at each contributing hospital. Stored in WFDB (WaveForm DataBase) format following extended PhysioNet schema. Contains 23 TB of waveform data as of August 2025.
chorus:collection:2Waveform telemetry capture
Medical imaging data acquired from hospital Picture Archiving and Communication Systems (PACS) and stored in DICOM format with comprehensive metadata following DICOM schema. 7,642 admissions have radiology data; 1,000 images currently available.
chorus:collection:3Medical imaging acquisition from PACS
Electroencephalography recordings extracted from hospital EEG databases in EDF+ (European Data Format) and Persyst formats with metadata following open source schemas. Extraction in process as of August 2025.
chorus:collection:4EEG recording extraction
DescriptionIDName
Demographics, medication administration (dosing time-stamped upon each infusion change or dose administration), procedures, nursing flowsheets (high-frequency documentation), and diagnoses acquired from electronic health records and standardized to OMOP Common Data Model. Controlled access with published OMOP schema metadata. Contains 1.6 billion rows of OMOP data as of August 2025.
chorus:acquisition:1Structured EHR data via OMOP transformation
Clinical notes extracted from EHR systems and tokenized using OHNLP (Open Health Natural Language Processing) toolkit. Stored locally at sites except for tokens. Controlled access planned with OHNLP open source schema metadata.
chorus:acquisition:2Clinical notes via OHNLP tokenization
Imaging data acquired from hospital PACS systems in DICOM format. De-identification in process. Planned controlled access with published DICOM schema metadata. chorus:acquisition:3Medical imaging via DICOM from PACS
Bedside monitor waveform data acquired through gateway and middleware systems. Stored in WFDB format with controlled access and published PhysioNet schema (extended) metadata. 23 TB available as of August 2025.
chorus:acquisition:4Waveform telemetry via WFDB from bedside monitors
EEG recordings from hospital databases in EDF+ and Persyst formats. Extraction in process. Planned controlled access with open source EDF+ and Persyst schema metadata. chorus:acquisition:5EEG waveforms from hospital databases
DescriptionIDNamePreprocessing Details
All structured electronic health record data standardized to the OMOP (Observational Medical Outcomes Partnership) Common Data Model. Ensures interoperability and enables use of OHDSI (Observational Health Data Sciences and Informatics) tool stack for analysis. Data from all 14 acquisition centers unified through this transformation.
chorus:preproc:1OMOP Common Data Model transformationTransformation of source EHR data from diverse institutional formats to OMOP CDM, Standardization of terminology and clinical codes to OMOP vocabulary standards, Mapping to OMOP vocabulary standards using validated semantic mappings, Quality assurance of transformed data against OMOP schema specifications, Integration with OHDSI tools for downstream analysis and characterization, Generation of characterization reports returned to contributing sites via CHoRUSReports
Clinical notes processed using OHNLP (Open Health Natural Language Processing) toolkit for extraction and tokenization. Protects patient privacy while enabling natural language processing and analysis of clinical text data.
chorus:preproc:2Clinical note tokenization via OHNLPText extraction from clinical notes in source EHR systems, OHNLP tokenization pipeline applied to free-text clinical documentation, De-identification of sensitive information through tokenization approach, Standardization to OHNLP open source schema, Local storage of full notes at sites; only tokens available in enclave
Waveform telemetry data from diverse bedside monitoring systems standardized to WFDB (WaveForm DataBase) format following extended PhysioNet schema. Scripts available in chorus_waveform repository. chorus:preproc:3Waveform standardization to WFDB formatConversion from proprietary bedside monitor formats using gateway/middleware systems, Standardization to WFDB format per extended PhysioNet schema, Metadata extraction and schema compliance verification, Quality checks for waveform integrity and completeness, Synchronization with clinical events from OMOP EHR data
DICOM imaging data undergoing de-identification process to remove patient identifiable information from image metadata while preserving clinical utility and image quality. chorus:preproc:4Medical imaging de-identificationDICOM header de-identification to remove embedded patient information, Preservation of clinically relevant imaging metadata, DICOM schema compliance verification after de-identification, Quality assurance of de-identified images for clinical utility, Privacy scan tool (privacy_scan_tool) used for medical records privacy scanning
Data transformed using approaches that limit re-identification while maintaining analytical utility. Multiple preprocessing strategies employed to protect patient privacy across all data modalities per HIPAA and institutional requirements.
chorus:preproc:5Re-identification limitation transformationsApplication of de-identification algorithms across all data modalities, Privacy-preserving transformations for EHR and imaging data, Geocoding via DeGauss (UF-Geocoding tool) for OMOP Location entities, Risk assessment and compliance with ethical and legal requirements, Community ethics focus group input on appropriate data for public sharing
Cleaning DetailsDescriptionIDName
Semantic mapping validation by clinical experts using chorus-mapping repository, Standard operating protocol (SOP) implementation per Chorus_SOP documentation, Cross-site data quality checks through CHoRUSReports characterization reports, Resolution of institutional variations in coding and clinical terminology, Clinical validation SOP for contributing to mapping efforts, Site status tracking via GitHub interface and Google Form submissions
Data from 14 acquisition centers harmonized through validated semantic mappings and standard operating protocols (SOPs). Ensures consistency and interoperability across diverse institutional EHR systems and clinical practices. Mappings maintained in chorus-mapping repository with clinical validation SOP.
chorus:cleaning:1Multi-center data harmonization with validated semantic mappings
Schema compliance verification against OMOP CDM specifications, DICOM schema metadata validation for imaging data, WFDB/PhysioNet schema validation for waveform data, OHNLP open source schema validation for tokenized clinical notes, EDF+ and Persyst schema validation for EEG data, Documentation of schema extensions for OMOP high-frequency nursing flowsheetsAll data modalities validated against published metadata schemas (OMOP, DICOM, WFDB, OHNLP, EDF+, Persyst) to ensure compliance and data quality across the multi-modal dataset. chorus:cleaning:2Schema compliance validation across all modalities
  1. ID
    chorus:labeling:1
    Name
    Visualization and annotation environment for prediction targets
    Description
    Custom visualization and annotation environment developed to label data with targets important for prediction tasks in critical care AI applications. Supports labeling for characterizing acute illness, predicting complications, and measuring treatment response.
    Data Annotation Protocol
    • Interactive visualization tools for clinical data exploration
    • Annotation interface for clinical expert labeling of prediction targets
    • Labeling of targets for characterization, prediction, and treatment response tasks
    • Quality control of annotations through review processes
    • Documentation of labeling protocols per Chorus_SOP standard operating procedures
DescriptionIDMaintainer DetailsName
Multi-institutional consortium managing dataset maintenance including 14 data acquisition centers, coordinating teams at Massachusetts General Hospital (lead), University of Florida, UT Health Science Center, and Tufts Medicine, plus infrastructure development team. Standards, Data Acquisition, and Tooling sub-teams manage ongoing data delivery and quality. Contact: cmccrary@mgh.harvard.edu (Ciera McCrary, Program Manager, MGH).
chorus:maintainer:1Massachusetts General Hospital (lead institution, Contact PI Eric S. Rosenthal), University of Florida (data acquisition and coordination, Azra Bihorac, Parisa Rashidi, Yulia Strekalova), UT Health Science Center / UTHealth Houston (Xiaoqian Jiang), Tufts Medicine (Andrew Ewing Williams, Manlik Kwong, Jared Houghtaling), 14 data acquisition centers across United States contributing clinical data extracts, Standards team (semantic mappings and validation via chorus-mapping repository), Data Acquisition team (extraction and contribution per Chorus_SOP), Tooling team (software development across chorus-ai GitHub organization), Project management via GitHub organization (chorus-ai) with 28 active repositoriesCHoRUS Consortium
Active GitHub organization housing repositories for software, semantic mappings, standard operating protocols, and project management. 28 repositories with comprehensive documentation and community support. Licensed under MIT License (GitHub organization).
chorus:maintainer:2GitHub organization at https://github.com/chorus-ai, Chorus_SOP repository - centralized SOP documentation site (Apache-2.0 license), chorus-container-apps - Azure deployment infrastructure (JavaScript), chorus-mapping - semantic mappings repository, chorus_waveform - waveform documentation and conversion scripts (MIT license), privacy_scan_tool - privacy scan tool for medical records (Python), CHoRUSReports - characterization reports returned to contributing sites (R), chorus-extract-upload - tools to create and upload CHoRUS data extract (MIT license), UF-Geocoding - open source code to geocode OMOP Location entities via DeGauss, Community discussions and issue tracking for Standards and Data Acquisition teamsCHoRUS GitHub Organization (chorus-ai)
ID
chorus:retention:1
Name
Long-term dataset retention per NIH data sharing policies
Description
Digital data maintained according to NIH data sharing policies and institutional requirements at participating centers. Controlled access model ensures long-term availability for research while protecting patient privacy. Funded under NIH award OT2OD032701 with project end date of November 30, 2026.
Retention Details
  • NIH data sharing policies govern long-term retention requirements
  • Institutional requirements at 14 participating data acquisition centers apply
  • Controlled access model via secure enclave for ongoing privacy protection
  • Long-term maintenance through CHoRUS Consortium and associated institutions
DescriptionExternal ResourcesIDName
Official project website with dataset overview, team information, project components, and access instructionshttps://chorus4ai.org/chorus:resource:1CHoRUS Project Website
Comprehensive GitHub organization with 28 repositories including software, documentation, SOPs, and toolinghttps://github.com/chorus-aichorus:resource:2CHoRUS GitHub Organization
Centralized standard operating protocol documentation with interactive workflow diagrams for data extraction and contributionhttps://github.com/chorus-ai/Chorus_SOPchorus:resource:3Chorus_SOP Documentation Site
Federal grant information and project details from NIH Research Portfolio Online Reporting Tools for grant 1OT2OD032701-01https://reporter.nih.gov/project-details/10472824chorus:resource:4NIH RePORTER Project Details
Parent NIH Common Fund program supporting AI-ready biomedical datasets across four data generation projectshttps://bridge2ai.org/, https://bridge2ai.org/choruschorus:resource:5Bridge2AI Program
Partnership with AIM-AHEAD for Bridge2AI Clinical Care Training Program (Cohort 1 and Cohort 2) providing AI/ML training for underrepresented traineeshttps://aim-ahead.net/chorus:resource:6AIM-AHEAD Bridge2AI Training Program
Observational Health Data Sciences and Informatics community supporting OMOP Common Data Model used for CHoRUS structured EHR datahttps://www.ohdsi.org/chorus:resource:7OHDSI Community and OMOP CDM
Peer-reviewed publication documenting CHoRUS dataset methodology and design in Neurocritical Care journalhttps://doi.org/10.1007/s12028-024-02007chorus:resource:8Published Research (Neurocritical Care)
🚀

Uses

What (other) tasks could the dataset be used for?

DescriptionIDName
Generate data for ML/AI applications aimed at characterizing acute and critical care illness patterns, progression, and outcomes across diverse patient populations and hospital settings. chorus:task:1Characterize acute and critical care illness
Enable prediction of complications among patients with acute or critical illness using multi-modal data including structured EHR, waveforms, imaging, and clinical notes. chorus:task:2Predict complications in critically ill patients
Support measurement and analysis of treatment response among critically ill patients through high-frequency documentation, medication administration records, and clinical outcomes data. chorus:task:3Measure treatment response
Provision a holdout test set accessible for model external validation to aid marketplace adoption of AI-developed models for implementation in acute and critical care settings. chorus:task:4External validation for AI model marketplace adoption
Utilize visualization and annotation environment to label data with targets important for prediction tasks in critical care AI applications. chorus:task:5Label data for prediction targets
DescriptionExamplesIDName
Primary intended use is development and training of artificial intelligence and machine learning models to characterize acute and critical care illness, predict complications, and measure treatment response in critically ill patients across diverse hospital settings.
Characterizing acute and critical care illness patterns using multi-modal data, Predicting complications (e.g., sepsis, respiratory failure) in critically ill patients, Measuring treatment response in ICU patients using medication and waveform data, Developing clinical deep learning models for critical care AI applicationschorus:use:1AI/ML model development for critical care
Provision of holdout test set accessible for model external validation to aid marketplace adoption of AI-developed models for implementation in acute and critical care settings. External validation of sepsis prediction models developed at other institutions, Benchmarking AI algorithms for critical care across diverse patient populationschorus:use:2External validation of AI models
Studies examining health equity, social determinants of health, and disparities in critical care outcomes across diverse patient populations and hospital settings. Dataset includes contextual factors such as geographic distance to nearest hospital.
Analysis of disparities in critical care outcomes by race, ethnicity, and geography, Research on social determinants of health in ICU patient populationschorus:use:3Health equity and disparities research
Training and education of next generation of diverse academic and community AI scientists through hands-on experience with real-world critical care datasets. Integrated with AIM-AHEAD Bridge2AI for Clinical Care Training Program (Cohorts 1 and 2).
AIM-AHEAD training program for underrepresented trainees (Cohort 1: 2024-2025, Cohort 2: 2025-2026), Foundational hands-on training using Jupyter Notebooks with Bridge2AI CHoRUS ecosystem, Workshops on OHDSI/OMOP common data model and clinical AI, Development of practical use cases for AI/ML in clinical carechorus:use:4Educational and training purposes for AI scientists
DescriptionDiscouragement DetailsIDName
Dataset is for research purposes only. AI/ML models developed should undergo appropriate clinical validation, regulatory approval, and institutional review before use in patient care or clinical decision-making.
Models trained on CHoRUS data must undergo independent clinical validation, Regulatory approval processes (e.g., FDA clearance) required for clinical use, Institutional review board approval needed for clinical deploymentchorus:discouraged:1Clinical decision-making without proper validation and regulatory approval
Attempts to re-identify patients from de-identified data violate ethical principles, data use agreements, and legal frameworks established for privacy protection under HIPAA and institutional requirements.
Re-identification attempts violate the signed data use agreement, Prohibited under HIPAA and applicable institutional data privacy regulations, Data use agreement explicitly prohibits re-identification effortschorus:discouraged:2Re-identification attempts
As data collection continues through November 2026 and quality assurance processes are ongoing, early dataset versions should be used with awareness of completeness limitations and ongoing expansion (from 45K to target 100K+ admissions).
Dataset is actively growing; cohort coverage varies by data modality, EEG extraction and full imaging de-identification still in process as of 2025, Clinical notes stored locally at sites; only tokens available in enclavechorus:discouraged:3Use without awareness of ongoing data collection limitations
ID
chorus:license:1
Name
CHoRUS Controlled Access License with Data Use Agreement
Description
Dataset distributed under controlled access requiring institutional email registration and signed licensing agreement. Access granted after review and approval process. Participants must complete registration form with name, institutional email (not personal), and institution. Once approved, users receive email with access instructions to CHoRUS secure enclave. For training program access, program administrators assist with licensing.
License Terms
  • Institutional (.edu) email required for registration
  • All participants must sign a licensing agreement before gaining access to the dataset
  • Controlled access through secure enclave (Azure-based infrastructure)
  • Data use agreement specifies permitted research uses and prohibits re-identification
  • Access request contacts - dbold@emory.edu or jared.houghtaling@tuftsmedicine.org
  • Funded under NIH award OT2OD032701; content is solely responsibility of authors
📤

Distribution

How will the dataset be distributed?

Controlled Access - Data Use Agreement Required (OT2OD032701)
🔄

Maintenance

How will the dataset be maintained?

ID
chorus:updates:1
Name
Ongoing data collection and continuous expansion
Description
Dataset updated continuously as data collection progresses at 14 acquisition centers. As of August 2025, covers 14 hospitals with over 45,000 unique admissions; current released dataset includes 50,000 patient admissions (ICU, PICU, NICU) and 1.6 billion rows of EHR OMOP data. Target exceeds 100,000 critically ill patients. Project timeline extends through November 30, 2026 (approved no-cost extension). Regular status updates tracked through GitHub project management system via GitHub interface or Google Form submissions. Sites statuses tracked in Standards Project and Data Acquisition Project.
Frequency
Continuous updates through November 30, 2026
Update Details
  • Ongoing retrospective data collection at 14 sites through project end date
  • Current status (August 2025) - 45K+ unique admissions; 50K released (ICU, PICU, NICU)
  • Target - 100,000+ critically ill patients across 9 data modalities
  • Regular site status updates via GitHub interface or Google Form submissions
  • GitHub project tracking for deliverables and task dependencies
  • Documentation updates maintained in Chorus_SOP repository
  • Software and tooling continuous development across chorus-ai GitHub organization
  • Semantic mapping validation and expansion via chorus-mapping repository
  • EEG extraction and full imaging de-identification in progress
👥

Human Subjects

Does the dataset relate to people?

ID
chorus:hsr:1
Name
CHoRUS Human Subjects Research
Description
Retrospective data collection from critically ill patients (ICU, PICU, NICU) approved through institutional review processes at 14 data acquisition centers. Community-facing ethics focus groups conducted to determine what data is appropriate for public sharing. Legal framework established for collecting data at scale. Patient-focused efforts determine ethical and legal approaches to manage privacy and bias while accounting for Social Determinants of Health. Project includes expertise from law, ethics, health services, biomedical science, engineering, and scientific journal publications disciplines. Ethics of AI component addressed through AIM-AHEAD training curriculum (safety, risk, and legal considerations; IRB, HIPAA/GDPR compliance for OMOP/FHIR data).
Involves Human Subjects
True
Ethics Review Board
  • Institutional review boards at 14 data acquisition centers across United States
  • Community-facing ethics focus groups determining appropriate data for public sharing
  • Legal and ethical advisory teams including law and ethics discipline experts
  • Privacy and accountability review processes per project three-pillar structure (Data, Ethics, People)
Regulatory Compliance
  • HIPAA (Health Insurance Portability and Accountability Act) compliance for protected health information
  • 45 CFR 46 (Common Rule) for human subjects research protections
  • Institutional data privacy regulations at each of the 14 contributing sites
  • NIH Common Fund Bridge2AI program ethical and trustworthy AI requirements
  • IRB protocol drafting and HIPAA/GDPR compliance guidance provided in training curriculum
Generated on 2026-04-15 17:53:09 using Bridge2AI Data Sheets Schema