id: https://chorus4ai.org/
name: CHoRUS
title: Patient-Focused Collaborative Hospital Repository Uniting Standards (CHoRUS) for Equitable AI
description: 'CHoRUS for Equitable AI is a Bridge2AI data generation project developing the most diverse, high-resolution, ethically sourced, AI-ready critical care dataset to answer the grand challenge of improving recovery from acute illness. The project spans 20 academic centers (14 data acquisition centers) and is building a publicly available dataset targeting over 100,000 critically ill patients with multi-modal data including structured EHR, waveform telemetry, medical imaging, EEG, and clinical notes. All structured data is standardized to the OMOP Common Data Model with additional formats (DICOM, WFDB, OHNLP tokenization) and comprehensive metadata schemas. Patient-focused efforts determine ethical and legal approaches to manage privacy and bias while accounting for Social Determinants of Health. A visualization and annotation environment labels data with targets important for prediction. The project emphasizes skills and workforce development for a next generation of diverse academic and community AI scientists through training programs and partnerships with AIM-AHEAD. As of August 2025, the dataset covers 14 different hospitals with over 45,000 unique admissions and includes 50,000 patient admissions from ICU, PICU, and NICU, 1.6 billion rows of EHR OMOP data, 7,642 admissions with radiology data, and 23 TB of waveform data. '
page: https://chorus4ai.org/
language: en
license: Controlled Access - Data Use Agreement Required (OT2OD032701)
keywords: - CHoRUS - Bridge2AI - critical care - acute illness - AI-ready dataset - OMOP Common Data Model - electronic health records - EHR - waveform telemetry - medical imaging - DICOM - EEG - clinical notes - OHNLP - health equity - social determinants of health - federated access - multi-modal data - high-resolution data - ethical AI - trustworthy AI - workforce development - data standardization - privacy preservation - bias mitigation - ICU - intensive care
purposes:
- id: chorus:purpose:1
name: Improve recovery from acute illness
description: 'Answer the grand challenge of improving recovery from acute illness by developing high-resolution
multi-center datasets as a critical first step towards actionable and trustworthy AI in critical care.
Address the urgent need for infrastructure to support artificial intelligence and machine learning
(AI/ML) in critical care settings.
'
- id: chorus:purpose:2
name: Create AI-ready critical care dataset
description: 'Develop a publicly available, AI-ready critical care dataset from more than 100,000 critically
ill patients while ensuring methods promote privacy, accountability, and clinical benefit. Generate
the most diverse, high-resolution, ethically sourced dataset for AI/ML applications in acute and critical
care, expanding AI and Machine Learning to improve recovery from acute illness.
'
- id: chorus:purpose:3
name: Establish data standards and tools
description: 'Unify standards to harmonize multi-modal EHR, waveform, imaging, and text data. Develop
software and tooling to interact with and extract insight from clinical data in diverse formats. Create
validated semantic mappings for connecting clinical data in various source formats to international
standards (OMOP Common Data Model, DICOM, WFDB, OHNLP).
'
- id: chorus:purpose:4
name: Promote diversity and health equity
description: 'Ensure comprehensive sets of patient conditions and clinical treatment strategies with
appropriate contextual factors such as geographic distance to nearest hospital and Social Determinants
of Health. Develop the skills and workforce for a next generation of diverse academic and community
AI scientists through comprehensive training and education programs in partnership with AIM-AHEAD.
'
tasks:
- id: chorus:task:1
name: Characterize acute and critical care illness
description: 'Generate data for ML/AI applications aimed at characterizing acute and critical care illness
patterns, progression, and outcomes across diverse patient populations and hospital settings.
'
- id: chorus:task:2
name: Predict complications in critically ill patients
description: 'Enable prediction of complications among patients with acute or critical illness using
multi-modal data including structured EHR, waveforms, imaging, and clinical notes.
'
- id: chorus:task:3
name: Measure treatment response
description: 'Support measurement and analysis of treatment response among critically ill patients through
high-frequency documentation, medication administration records, and clinical outcomes data.
'
- id: chorus:task:4
name: External validation for AI model marketplace adoption
description: 'Provision a holdout test set accessible for model external validation to aid marketplace
adoption of AI-developed models for implementation in acute and critical care settings.
'
- id: chorus:task:5
name: Label data for prediction targets
description: 'Utilize visualization and annotation environment to label data with targets important
for prediction tasks in critical care AI applications.
'
addressing_gaps:
- id: chorus:gap:1
name: Lack of diverse multi-center critical care datasets
description: 'Address the absence of large-scale, diverse, high-resolution multi-center datasets for
critical care AI/ML by creating a dataset spanning 20 academic centers with 100,000+ critically ill
patients and ensuring balanced, diverse cohorts through federated access and sampling methods across
14 data acquisition centers.
'
- id: chorus:gap:2
name: Insufficient data standardization in critical care
description: 'Overcome the lack of unified standards in critical care data by harmonizing multi-modal
EHR, waveform, imaging, and text data to OMOP Common Data Model and other international standards
(DICOM, WFDB, OHNLP).
'
- id: chorus:gap:3
name: Privacy and bias concerns in clinical AI
description: 'Address ethical and legal challenges in AI through patient-focused efforts that determine
approaches to manage privacy and bias while accounting for Social Determinants of Health and performing
community-facing ethics focus groups to determine what data is appropriate for public sharing.
'
- id: chorus:gap:4
name: Limited diversity in AI/ML workforce
description: 'Bridge the gap in diverse AI/ML workforce through comprehensive educational approaches,
training programs (including AIM-AHEAD partnership), and cultivation of expertise in lay and scientific
communities to improve AI literacy and utilization.
'
creators:
- id: chorus:creator:1
name: Eric S. Rosenthal
description: Contact PI/Project Leader, Massachusetts General Hospital (MGH), Director of MGH Neurosciences
ICU
- id: chorus:creator:2
name: Azra Bihorac
description: Principal Investigator, University of Florida (UF)
- id: chorus:creator:3
name: Ashley Cordes
description: Principal Investigator, CHoRUS Consortium
- id: chorus:creator:4
name: Gilles Clermont
description: Principal Investigator, CHoRUS Consortium
- id: chorus:creator:5
name: Gari David Clifford
description: Principal Investigator, CHoRUS Consortium
- id: chorus:creator:6
name: Barbara J. Evans
description: Principal Investigator, CHoRUS Consortium
- id: chorus:creator:7
name: Xiao Hu
description: Principal Investigator, CHoRUS Consortium
- id: chorus:creator:8
name: Rishikesan Kamaleswaran
description: Principal Investigator, CHoRUS Consortium
- id: chorus:creator:9
name: Yulia A. Levites Strekalova
description: Principal Investigator, University of Florida (UF)
- id: chorus:creator:10
name: Parisa Rashidi
description: Principal Investigator, University of Florida (UF)
- id: chorus:creator:11
name: Cynthia Rudin
description: Principal Investigator, CHoRUS Consortium
- id: chorus:creator:12
name: Ishan Canty Williams
description: Principal Investigator, CHoRUS Consortium
- id: chorus:creator:13
name: Andrew Ewing Williams
description: Principal Investigator, Tufts Medicine
- id: chorus:creator:14
name: Xiaoqian Jiang
description: Program Lead, UT Health Science Center (UTHealth Houston)
- id: chorus:creator:15
name: Morteza Zabihi
description: Lecturer (Machine Learning Basics), Massachusetts General Hospital
- id: chorus:creator:16
name: Zhenhong Hu
description: Instructor (Python and Version Control), University of Florida
- id: chorus:creator:17
name: Debora Simmons
description: Lecturer (Ethics of AI in Clinical Practice), UT Health Science Center
- id: chorus:creator:18
name: Aliyah Geer
description: Workshop Lead (Data Schemas in Clinical Cloud), Massachusetts General Hospital
- id: chorus:creator:19
name: Ciera McCrary
description: Program Manager, Massachusetts General Hospital (MGH)
funders:
- id: chorus:funder:1
name: NIH Common Fund Bridge2AI Program
description: 'Funded through National Institutes of Health grant OT2OD032701 (project number 1OT2OD032701-01),
administered by NIH Office of the Director. Opportunity Number: OTA-21-008. Study Section: Data Coordination,
Mapping, and Modeling (DCMM). Fiscal Year 2022. Total funding in 2022: $5,880,300 (all direct costs).
Project dates: September 1, 2022 to November 30, 2026 (with approved no-cost extension). Award notice
date: September 1, 2022. Assistance Listing Number: 93.310.
'
instances:
- id: chorus:instance:1
name: Critically ill patients in ICU settings
description: 'Individual critically ill patients requiring acute or critical care admitted to intensive
care units (ICU), pediatric intensive care units (PICU), or neonatal intensive care units (NICU) across
14 data acquisition centers. As of August 2025, the dataset covers 14 different hospitals with over
45,000 unique admissions, with 50,000 patient admissions from ICU, PICU, and NICU currently released.
Target enrollment exceeds 100,000 critically ill patients. Retrospective data collection from patients
with acute or critical illness.
'
instance_type: 'Human subjects - critically ill patients admitted to participating hospitals (retrospective
data collection, ongoing collection through November 2026).
'
subsets:
- id: chorus:subset:1
name: Controlled Access Structured EHR Dataset (OMOP)
description: 'Primary dataset with controlled access requiring data use agreement and licensing. Includes
OMOP-standardized structured EHR data comprising demographics, medication administration (dosing time-stamped
upon each infusion change or dose administration), procedures, nursing flowsheets (high-frequency
documentation), and diagnoses. Available in secure enclave. As of August 2025, contains 1.6 billion
rows of EHR OMOP data. Registration required with institutional email; all participants must sign
licensing agreement.
'
- id: chorus:subset:2
name: Waveform Telemetry Dataset (WFDB)
description: 'Waveform telemetry data from bedside monitors (gateway/middleware) standardized to WFDB
(WaveForm DataBase) format following extended PhysioNet schema. Available in secure enclave with controlled
access. As of August 2025, contains 23 TB of waveform data. Published metadata schema: PhysioNet schema
(extended).
'
- id: chorus:subset:3
name: Medical Imaging Dataset (DICOM)
description: 'Imaging data from hospital PACS (Picture Archiving and Communication System) in DICOM
format with comprehensive metadata following DICOM schema. As of August 2025, 1,000 images are available
with de-identification in process for the larger cohort. 7,642 admissions have radiology data. Controlled
access planned.
'
- id: chorus:subset:4
name: Clinical Notes Dataset (OHNLP Tokenized)
description: 'Clinical notes extracted and tokenized using OHNLP (Open Health Natural Language Processing)
toolkit following open source OHNLP schema. Stored locally at contributing sites (except tokens).
Controlled access planned. De-identification achieved through tokenization approach.
'
- id: chorus:subset:5
name: EEG Waveform Dataset (EDF+ and Persyst)
description: 'Electroencephalography (EEG) waveform data from hospital databases in EDF+ (European Data
Format) and Persyst formats. Open source EDF+ and Persyst schema metadata available. Extraction in
process, planned for controlled access.
'
- id: chorus:subset:6
name: Training and Publication Subset
description: 'Subsets of the dataset currently being used for training activities (AIM-AHEAD Bridge2AI
for Clinical Care Training Program) and publications as the dataset undergoes expansion and quality
assurance processes.
'
sampling_strategies:
- id: chorus:sampling:1
name: Multi-center federated sampling
description: 'Federated access enables sampling methods to ensure a balanced and diverse cohort across
20 academic centers (14 data acquisition centers). Legal framework established for collecting data
at scale with sampling to ensure comprehensive sets of patient conditions and clinical treatment strategies.
Community-facing ethics focus groups conducted to determine what data is appropriate for public sharing.
'
is_sample: true
is_random: false
is_representative: true
source_data:
- Critically ill patients in ICU, PICU, and NICU settings at 14 data acquisition centers across 20 academic
centers in the United States
representative_verification:
- Federated multi-center data collection across 14 hospitals ensures geographic and institutional diversity;
sampling for balanced and diverse patient populations; inclusion of Social Determinants of Health
contextual factors
subpopulations:
- id: chorus:subpop:1
name: Critically ill patients by hospital
description: 'Patients distributed across 14 different hospitals within the CHoRUS network spanning
20 academic centers in the United States, ensuring geographic and institutional diversity in critical
care settings.
'
- id: chorus:subpop:2
name: ICU, PICU, and NICU patients
description: 'Patients admitted to intensive care units (ICU), pediatric intensive care units (PICU),
and neonatal intensive care units (NICU) across contributing hospitals, with 50,000 patient admissions
in the current released dataset.
'
sensitive_elements:
- id: chorus:sensitive:1
name: Protected health information - clinical EHR data
description: 'Complete electronic health records including demographics, diagnoses, procedures, medications,
nursing documentation, and clinical notes for critically ill patients. Contains protected health information
subject to HIPAA and institutional privacy requirements. De-identified before release through OMOP
transformation and privacy scanning tools.
'
sensitive_elements_present: true
sensitivity_details:
- Demographics and patient identifiers (de-identified before controlled access release)
- Diagnoses and medical conditions (OMOP standardized)
- Medication administration records with dosing timestamps (OMOP standardized)
- Procedures and clinical interventions documented by providers (OMOP standardized)
- Clinical notes (tokenized using OHNLP toolkit for privacy protection)
- Nursing flowsheet documentation at high frequency (OMOP with extensions)
- id: chorus:sensitive:2
name: Physiological monitoring and EEG waveform data
description: 'Continuous waveform telemetry from bedside monitors and EEG recordings capturing detailed
physiological states of critically ill patients. May reveal sensitive health conditions and treatment
responses. Stored in WFDB format (telemetry) and EDF+/Persyst formats (EEG) with controlled access.
'
sensitive_elements_present: true
sensitivity_details:
- Continuous cardiac waveforms from bedside monitors (WFDB format, 23 TB)
- Respiratory monitoring and hemodynamic measurement data
- Electroencephalography (EEG) recordings from hospital databases
- High-frequency physiological parameters
- id: chorus:sensitive:3
name: Medical imaging data
description: 'Diagnostic imaging studies in DICOM format from hospital PACS systems. De-identification
in process to remove embedded patient information while preserving clinical utility. As of August
2025, 1,000 images available with de-id in process for larger cohort (7,642 admissions with radiology
data total).
'
sensitive_elements_present: true
sensitivity_details:
- Radiology images (CT scans, X-rays, MRI and other modalities) in DICOM format
- DICOM metadata (de-identification in process)
- Embedded patient information being removed through de-identification pipeline
- id: chorus:sensitive:4
name: Social Determinants of Health data
description: 'Contextual factors including geographic information (distance to hospital) and social
determinants of health data. Collected to support health equity research while maintaining patient
privacy through privacy-preserving transformations.
'
sensitive_elements_present: true
sensitivity_details:
- Geographic location information (distance to nearest hospital)
- Social determinants of health variables
- Contextual equity factors with privacy-preserving transformations applied
collection_mechanisms:
- id: chorus:collection:1
name: Retrospective EHR extraction
description: 'Retrospective data collection from electronic health record systems at 14 data acquisition
centers. Data extracted includes demographics, medication administration (dosing time-stamped upon
each infusion change or dose administration), procedures, nursing flowsheets (high-frequency documentation),
diagnoses, and clinical notes. Standardized to OMOP Common Data Model.
'
- id: chorus:collection:2
name: Waveform telemetry capture
description: 'Continuous waveform telemetry data captured from bedside monitors through gateway and
middleware systems at each contributing hospital. Stored in WFDB (WaveForm DataBase) format following
extended PhysioNet schema. Contains 23 TB of waveform data as of August 2025.
'
- id: chorus:collection:3
name: Medical imaging acquisition from PACS
description: 'Medical imaging data acquired from hospital Picture Archiving and Communication Systems
(PACS) and stored in DICOM format with comprehensive metadata following DICOM schema. 7,642 admissions
have radiology data; 1,000 images currently available.
'
- id: chorus:collection:4
name: EEG recording extraction
description: 'Electroencephalography recordings extracted from hospital EEG databases in EDF+ (European
Data Format) and Persyst formats with metadata following open source schemas. Extraction in process
as of August 2025.
'
acquisition_methods:
- id: chorus:acquisition:1
name: Structured EHR data via OMOP transformation
description: 'Demographics, medication administration (dosing time-stamped upon each infusion change
or dose administration), procedures, nursing flowsheets (high-frequency documentation), and diagnoses
acquired from electronic health records and standardized to OMOP Common Data Model. Controlled access
with published OMOP schema metadata. Contains 1.6 billion rows of OMOP data as of August 2025.
'
- id: chorus:acquisition:2
name: Clinical notes via OHNLP tokenization
description: 'Clinical notes extracted from EHR systems and tokenized using OHNLP (Open Health Natural
Language Processing) toolkit. Stored locally at sites except for tokens. Controlled access planned
with OHNLP open source schema metadata.
'
- id: chorus:acquisition:3
name: Medical imaging via DICOM from PACS
description: 'Imaging data acquired from hospital PACS systems in DICOM format. De-identification in
process. Planned controlled access with published DICOM schema metadata.
'
- id: chorus:acquisition:4
name: Waveform telemetry via WFDB from bedside monitors
description: 'Bedside monitor waveform data acquired through gateway and middleware systems. Stored
in WFDB format with controlled access and published PhysioNet schema (extended) metadata. 23 TB available
as of August 2025.
'
- id: chorus:acquisition:5
name: EEG waveforms from hospital databases
description: 'EEG recordings from hospital databases in EDF+ and Persyst formats. Extraction in process.
Planned controlled access with open source EDF+ and Persyst schema metadata.
'
preprocessing_strategies:
- id: chorus:preproc:1
name: OMOP Common Data Model transformation
description: 'All structured electronic health record data standardized to the OMOP (Observational Medical
Outcomes Partnership) Common Data Model. Ensures interoperability and enables use of OHDSI (Observational
Health Data Sciences and Informatics) tool stack for analysis. Data from all 14 acquisition centers
unified through this transformation.
'
preprocessing_details:
- Transformation of source EHR data from diverse institutional formats to OMOP CDM
- Standardization of terminology and clinical codes to OMOP vocabulary standards
- Mapping to OMOP vocabulary standards using validated semantic mappings
- Quality assurance of transformed data against OMOP schema specifications
- Integration with OHDSI tools for downstream analysis and characterization
- Generation of characterization reports returned to contributing sites via CHoRUSReports
- id: chorus:preproc:2
name: Clinical note tokenization via OHNLP
description: 'Clinical notes processed using OHNLP (Open Health Natural Language Processing) toolkit
for extraction and tokenization. Protects patient privacy while enabling natural language processing
and analysis of clinical text data.
'
preprocessing_details:
- Text extraction from clinical notes in source EHR systems
- OHNLP tokenization pipeline applied to free-text clinical documentation
- De-identification of sensitive information through tokenization approach
- Standardization to OHNLP open source schema
- Local storage of full notes at sites; only tokens available in enclave
- id: chorus:preproc:3
name: Waveform standardization to WFDB format
description: 'Waveform telemetry data from diverse bedside monitoring systems standardized to WFDB (WaveForm
DataBase) format following extended PhysioNet schema. Scripts available in chorus_waveform repository.
'
preprocessing_details:
- Conversion from proprietary bedside monitor formats using gateway/middleware systems
- Standardization to WFDB format per extended PhysioNet schema
- Metadata extraction and schema compliance verification
- Quality checks for waveform integrity and completeness
- Synchronization with clinical events from OMOP EHR data
- id: chorus:preproc:4
name: Medical imaging de-identification
description: 'DICOM imaging data undergoing de-identification process to remove patient identifiable
information from image metadata while preserving clinical utility and image quality.
'
preprocessing_details:
- DICOM header de-identification to remove embedded patient information
- Preservation of clinically relevant imaging metadata
- DICOM schema compliance verification after de-identification
- Quality assurance of de-identified images for clinical utility
- Privacy scan tool (privacy_scan_tool) used for medical records privacy scanning
- id: chorus:preproc:5
name: Re-identification limitation transformations
description: 'Data transformed using approaches that limit re-identification while maintaining analytical
utility. Multiple preprocessing strategies employed to protect patient privacy across all data modalities
per HIPAA and institutional requirements.
'
preprocessing_details:
- Application of de-identification algorithms across all data modalities
- Privacy-preserving transformations for EHR and imaging data
- Geocoding via DeGauss (UF-Geocoding tool) for OMOP Location entities
- Risk assessment and compliance with ethical and legal requirements
- Community ethics focus group input on appropriate data for public sharing
cleaning_strategies:
- id: chorus:cleaning:1
name: Multi-center data harmonization with validated semantic mappings
description: 'Data from 14 acquisition centers harmonized through validated semantic mappings and standard
operating protocols (SOPs). Ensures consistency and interoperability across diverse institutional
EHR systems and clinical practices. Mappings maintained in chorus-mapping repository with clinical
validation SOP.
'
cleaning_details:
- Semantic mapping validation by clinical experts using chorus-mapping repository
- Standard operating protocol (SOP) implementation per Chorus_SOP documentation
- Cross-site data quality checks through CHoRUSReports characterization reports
- Resolution of institutional variations in coding and clinical terminology
- Clinical validation SOP for contributing to mapping efforts
- Site status tracking via GitHub interface and Google Form submissions
- id: chorus:cleaning:2
name: Schema compliance validation across all modalities
description: 'All data modalities validated against published metadata schemas (OMOP, DICOM, WFDB, OHNLP,
EDF+, Persyst) to ensure compliance and data quality across the multi-modal dataset.
'
cleaning_details:
- Schema compliance verification against OMOP CDM specifications
- DICOM schema metadata validation for imaging data
- WFDB/PhysioNet schema validation for waveform data
- OHNLP open source schema validation for tokenized clinical notes
- EDF+ and Persyst schema validation for EEG data
- Documentation of schema extensions for OMOP high-frequency nursing flowsheets
labeling_strategies:
- id: chorus:labeling:1
name: Visualization and annotation environment for prediction targets
description: 'Custom visualization and annotation environment developed to label data with targets important
for prediction tasks in critical care AI applications. Supports labeling for characterizing acute
illness, predicting complications, and measuring treatment response.
'
data_annotation_protocol:
- Interactive visualization tools for clinical data exploration
- Annotation interface for clinical expert labeling of prediction targets
- Labeling of targets for characterization, prediction, and treatment response tasks
- Quality control of annotations through review processes
- Documentation of labeling protocols per Chorus_SOP standard operating procedures
intended_uses:
- id: chorus:use:1
name: AI/ML model development for critical care
description: 'Primary intended use is development and training of artificial intelligence and machine
learning models to characterize acute and critical care illness, predict complications, and measure
treatment response in critically ill patients across diverse hospital settings.
'
examples:
- Characterizing acute and critical care illness patterns using multi-modal data
- Predicting complications (e.g., sepsis, respiratory failure) in critically ill patients
- Measuring treatment response in ICU patients using medication and waveform data
- Developing clinical deep learning models for critical care AI applications
- id: chorus:use:2
name: External validation of AI models
description: 'Provision of holdout test set accessible for model external validation to aid marketplace
adoption of AI-developed models for implementation in acute and critical care settings.
'
examples:
- External validation of sepsis prediction models developed at other institutions
- Benchmarking AI algorithms for critical care across diverse patient populations
- id: chorus:use:3
name: Health equity and disparities research
description: 'Studies examining health equity, social determinants of health, and disparities in critical
care outcomes across diverse patient populations and hospital settings. Dataset includes contextual
factors such as geographic distance to nearest hospital.
'
examples:
- Analysis of disparities in critical care outcomes by race, ethnicity, and geography
- Research on social determinants of health in ICU patient populations
- id: chorus:use:4
name: Educational and training purposes for AI scientists
description: 'Training and education of next generation of diverse academic and community AI scientists
through hands-on experience with real-world critical care datasets. Integrated with AIM-AHEAD Bridge2AI
for Clinical Care Training Program (Cohorts 1 and 2).
'
examples:
- 'AIM-AHEAD training program for underrepresented trainees (Cohort 1: 2024-2025, Cohort 2: 2025-2026)'
- Foundational hands-on training using Jupyter Notebooks with Bridge2AI CHoRUS ecosystem
- Workshops on OHDSI/OMOP common data model and clinical AI
- Development of practical use cases for AI/ML in clinical care
discouraged_uses:
- id: chorus:discouraged:1
name: Clinical decision-making without proper validation and regulatory approval
description: 'Dataset is for research purposes only. AI/ML models developed should undergo appropriate
clinical validation, regulatory approval, and institutional review before use in patient care or clinical
decision-making.
'
discouragement_details:
- Models trained on CHoRUS data must undergo independent clinical validation
- Regulatory approval processes (e.g., FDA clearance) required for clinical use
- Institutional review board approval needed for clinical deployment
- id: chorus:discouraged:2
name: Re-identification attempts
description: 'Attempts to re-identify patients from de-identified data violate ethical principles, data
use agreements, and legal frameworks established for privacy protection under HIPAA and institutional
requirements.
'
discouragement_details:
- Re-identification attempts violate the signed data use agreement
- Prohibited under HIPAA and applicable institutional data privacy regulations
- Data use agreement explicitly prohibits re-identification efforts
- id: chorus:discouraged:3
name: Use without awareness of ongoing data collection limitations
description: 'As data collection continues through November 2026 and quality assurance processes are
ongoing, early dataset versions should be used with awareness of completeness limitations and ongoing
expansion (from 45K to target 100K+ admissions).
'
discouragement_details:
- Dataset is actively growing; cohort coverage varies by data modality
- EEG extraction and full imaging de-identification still in process as of 2025
- Clinical notes stored locally at sites; only tokens available in enclave
license_and_use_terms:
id: chorus:license:1
name: CHoRUS Controlled Access License with Data Use Agreement
description: 'Dataset distributed under controlled access requiring institutional email registration
and signed licensing agreement. Access granted after review and approval process. Participants must
complete registration form with name, institutional email (not personal), and institution. Once approved,
users receive email with access instructions to CHoRUS secure enclave. For training program access,
program administrators assist with licensing.
'
license_terms:
- Institutional (.edu) email required for registration
- All participants must sign a licensing agreement before gaining access to the dataset
- Controlled access through secure enclave (Azure-based infrastructure)
- Data use agreement specifies permitted research uses and prohibits re-identification
- Access request contacts - dbold@emory.edu or jared.houghtaling@tuftsmedicine.org
- Funded under NIH award OT2OD032701; content is solely responsibility of authors
distribution_formats:
- id: chorus:format:1
name: OMOP Common Data Model (structured EHR data)
description: 'Structured electronic health record data distributed in OMOP Common Data Model format
enabling use of OHDSI tool stack for analysis. Contains 1.6 billion rows. Available in secure enclave
with controlled access. Published OMOP schema metadata available.
'
access_urls:
- https://chorus4ai.org/
- id: chorus:format:2
name: WFDB waveform format (bedside monitor telemetry)
description: 'Waveform telemetry data (23 TB) distributed in WFDB (WaveForm DataBase) format following
extended PhysioNet schema. Available in secure enclave with controlled access. Published PhysioNet
schema (extended) metadata available.
'
access_urls:
- https://chorus4ai.org/
- id: chorus:format:3
name: DICOM format (medical imaging)
description: 'Medical imaging data (7,642 admissions with radiology; 1,000 images currently available)
distributed in DICOM format with comprehensive metadata following DICOM schema. De-identification
in process for larger cohort. Planned controlled access.
'
access_urls:
- https://chorus4ai.org/
- id: chorus:format:4
name: OHNLP tokenized format (clinical notes)
description: 'Clinical notes distributed as OHNLP-tokenized text following open source OHNLP schema.
Stored locally at contributing sites (except tokens). Controlled access planned.
'
access_urls:
- https://chorus4ai.org/
- id: chorus:format:5
name: EDF+ and Persyst formats (EEG waveforms)
description: 'EEG waveform data distributed in EDF+ (European Data Format) and Persyst formats following
open source schemas. Extraction in process, planned for controlled access.
'
access_urls:
- https://chorus4ai.org/
maintainers:
- id: chorus:maintainer:1
name: CHoRUS Consortium
description: 'Multi-institutional consortium managing dataset maintenance including 14 data acquisition
centers, coordinating teams at Massachusetts General Hospital (lead), University of Florida, UT Health
Science Center, and Tufts Medicine, plus infrastructure development team. Standards, Data Acquisition,
and Tooling sub-teams manage ongoing data delivery and quality. Contact: cmccrary@mgh.harvard.edu
(Ciera McCrary, Program Manager, MGH).
'
maintainer_details:
- Massachusetts General Hospital (lead institution, Contact PI Eric S. Rosenthal)
- University of Florida (data acquisition and coordination, Azra Bihorac, Parisa Rashidi, Yulia Strekalova)
- UT Health Science Center / UTHealth Houston (Xiaoqian Jiang)
- Tufts Medicine (Andrew Ewing Williams, Manlik Kwong, Jared Houghtaling)
- 14 data acquisition centers across United States contributing clinical data extracts
- Standards team (semantic mappings and validation via chorus-mapping repository)
- Data Acquisition team (extraction and contribution per Chorus_SOP)
- Tooling team (software development across chorus-ai GitHub organization)
- Project management via GitHub organization (chorus-ai) with 28 active repositories
- id: chorus:maintainer:2
name: CHoRUS GitHub Organization (chorus-ai)
description: 'Active GitHub organization housing repositories for software, semantic mappings, standard
operating protocols, and project management. 28 repositories with comprehensive documentation and
community support. Licensed under MIT License (GitHub organization).
'
maintainer_details:
- GitHub organization at https://github.com/chorus-ai
- Chorus_SOP repository - centralized SOP documentation site (Apache-2.0 license)
- chorus-container-apps - Azure deployment infrastructure (JavaScript)
- chorus-mapping - semantic mappings repository
- chorus_waveform - waveform documentation and conversion scripts (MIT license)
- privacy_scan_tool - privacy scan tool for medical records (Python)
- CHoRUSReports - characterization reports returned to contributing sites (R)
- chorus-extract-upload - tools to create and upload CHoRUS data extract (MIT license)
- UF-Geocoding - open source code to geocode OMOP Location entities via DeGauss
- Community discussions and issue tracking for Standards and Data Acquisition teams
updates:
id: chorus:updates:1
name: Ongoing data collection and continuous expansion
description: 'Dataset updated continuously as data collection progresses at 14 acquisition centers.
As of August 2025, covers 14 hospitals with over 45,000 unique admissions; current released dataset
includes 50,000 patient admissions (ICU, PICU, NICU) and 1.6 billion rows of EHR OMOP data. Target
exceeds 100,000 critically ill patients. Project timeline extends through November 30, 2026 (approved
no-cost extension). Regular status updates tracked through GitHub project management system via GitHub
interface or Google Form submissions. Sites statuses tracked in Standards Project and Data Acquisition
Project.
'
frequency: Continuous updates through November 30, 2026
update_details:
- Ongoing retrospective data collection at 14 sites through project end date
- Current status (August 2025) - 45K+ unique admissions; 50K released (ICU, PICU, NICU)
- Target - 100,000+ critically ill patients across 9 data modalities
- Regular site status updates via GitHub interface or Google Form submissions
- GitHub project tracking for deliverables and task dependencies
- Documentation updates maintained in Chorus_SOP repository
- Software and tooling continuous development across chorus-ai GitHub organization
- Semantic mapping validation and expansion via chorus-mapping repository
- EEG extraction and full imaging de-identification in progress
retention_limit:
id: chorus:retention:1
name: Long-term dataset retention per NIH data sharing policies
description: 'Digital data maintained according to NIH data sharing policies and institutional requirements
at participating centers. Controlled access model ensures long-term availability for research while
protecting patient privacy. Funded under NIH award OT2OD032701 with project end date of November 30,
2026.
'
retention_details:
- NIH data sharing policies govern long-term retention requirements
- Institutional requirements at 14 participating data acquisition centers apply
- Controlled access model via secure enclave for ongoing privacy protection
- Long-term maintenance through CHoRUS Consortium and associated institutions
human_subject_research:
id: chorus:hsr:1
name: CHoRUS Human Subjects Research
description: 'Retrospective data collection from critically ill patients (ICU, PICU, NICU) approved
through institutional review processes at 14 data acquisition centers. Community-facing ethics focus
groups conducted to determine what data is appropriate for public sharing. Legal framework established
for collecting data at scale. Patient-focused efforts determine ethical and legal approaches to manage
privacy and bias while accounting for Social Determinants of Health. Project includes expertise from
law, ethics, health services, biomedical science, engineering, and scientific journal publications
disciplines. Ethics of AI component addressed through AIM-AHEAD training curriculum (safety, risk,
and legal considerations; IRB, HIPAA/GDPR compliance for OMOP/FHIR data).
'
involves_human_subjects: true
ethics_review_board:
- Institutional review boards at 14 data acquisition centers across United States
- Community-facing ethics focus groups determining appropriate data for public sharing
- Legal and ethical advisory teams including law and ethics discipline experts
- Privacy and accountability review processes per project three-pillar structure (Data, Ethics, People)
regulatory_compliance:
- HIPAA (Health Insurance Portability and Accountability Act) compliance for protected health information
- 45 CFR 46 (Common Rule) for human subjects research protections
- Institutional data privacy regulations at each of the 14 contributing sites
- NIH Common Fund Bridge2AI program ethical and trustworthy AI requirements
- IRB protocol drafting and HIPAA/GDPR compliance guidance provided in training curriculum
external_resources:
- id: chorus:resource:1
name: CHoRUS Project Website
description: Official project website with dataset overview, team information, project components, and
access instructions
external_resources:
- https://chorus4ai.org/
- id: chorus:resource:2
name: CHoRUS GitHub Organization
description: Comprehensive GitHub organization with 28 repositories including software, documentation,
SOPs, and tooling
external_resources:
- https://github.com/chorus-ai
- id: chorus:resource:3
name: Chorus_SOP Documentation Site
description: Centralized standard operating protocol documentation with interactive workflow diagrams
for data extraction and contribution
external_resources:
- https://github.com/chorus-ai/Chorus_SOP
- id: chorus:resource:4
name: NIH RePORTER Project Details
description: Federal grant information and project details from NIH Research Portfolio Online Reporting
Tools for grant 1OT2OD032701-01
external_resources:
- https://reporter.nih.gov/project-details/10472824
- id: chorus:resource:5
name: Bridge2AI Program
description: Parent NIH Common Fund program supporting AI-ready biomedical datasets across four data
generation projects
external_resources:
- https://bridge2ai.org/
- https://bridge2ai.org/chorus
- id: chorus:resource:6
name: AIM-AHEAD Bridge2AI Training Program
description: Partnership with AIM-AHEAD for Bridge2AI Clinical Care Training Program (Cohort 1 and Cohort
2) providing AI/ML training for underrepresented trainees
external_resources:
- https://aim-ahead.net/
- id: chorus:resource:7
name: OHDSI Community and OMOP CDM
description: Observational Health Data Sciences and Informatics community supporting OMOP Common Data
Model used for CHoRUS structured EHR data
external_resources:
- https://www.ohdsi.org/
- id: chorus:resource:8
name: Published Research (Neurocritical Care)
description: Peer-reviewed publication documenting CHoRUS dataset methodology and design in Neurocritical
Care journal
external_resources:
- https://doi.org/10.1007/s12028-024-02007