Create an ethically sourced flagship dataset to enable AI research on voice as a biomarker of health and support insights across multiple clinical domains.
Grantor
Grant Name
Grant Number
National Institutes of Health (NIH)
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
Counts
Data Type
Instance Type
Representation
12523
Derived audio data: spectrograms (513 x N), acoustic features (openSMILE), phonetic and prosodic fea...
Audio-derived features and spectrograms per recording
Voice recordings (derived)
306
Demographics, clinical information, and validated questionnaire responses
Human subjects enrolled across five North American sites
Participants
Description
Clinical voice and health data acquired at point of care via standardized tasks and questionnaires; derived features computed from raw audio.
Was Directly Observed
yes
Was Reported By Subjects
yes
Was Inferred Derived
yes
Was Validated Verified
Validated questionnaires were used; derived signals followed standardized preprocessing.
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically sourced dataset linking derived voice recordings to clinical and demographic information to advance research on voice as a biomarker of health. Version 1.0 (released Nov 27, 2024) includes 12,523 recordings from 306 adult participants across five North American sites. This initial release contains low-risk derived data (e.g., spectrograms and acoustic/phonetic features) and detailed demographics, clinical data, and validated questionnaire responses. Original audio waveforms are omitted in this release; transcripts of free speech are removed. The dataset supports AI research across cohorts including voice disorders, neurological disorders, mood/psychiatric disorders, and respiratory disorders.
Lack of large, high-quality, multi-institutional, demographically diverse voice datasets linked to health information and collected under standardized, ethically grounded protocols.
ID
bridge2ai-voice-adult-v1.0
Name
Adult cohort v1.0
Title
Adult cohort (derived data only)
Description
Adult participants only for the initial release; derived data provided, original audio omitted.
Is Data Split
no
Is Subpopulation
yes
Is Sample
yes
Is Random
no
Source Data
Patients at specialty clinics enrolled into predefined disorder cohorts (adult cohort in v1.0)
Is Representative
no
Why Not Representative
Disorder-focused recruitment selected for specific conditions; not a general population sample
Strategies
Deterministic cohort assignment based on inclusion/exclusion criteria within five sites
Description
Project investigators at five specialty clinical sites in North America collected data under a standardized protocol.
Description
Standardized protocol using a custom tablet application; headset used for data collection when possible; REDCap used for source data capture and export.
Description
Data collection and sharing approved by the University of South Florida IRB; submitted to the University of Toronto Research Ethics Board.
Description
Raw audio converted to monaural, resampled to 16 kHz with a Butterworth anti-aliasing filter; spectrograms computed via short-time FFT (25 ms window, 10 ms hop, 512-point FFT); acoustic features extracted with openSMILE; phonetic/prosodic features via Parselmouth/Praat; automatic transcriptions via Whisper Large; integration and parquet generation via b2aiprep.
Used Software
Name
openSMILE
Parselmouth
Praat
torchaudio
Whisper Large
b2aiprep
Description
Standardized signal processing (resampling, channel conversion, anti-alias filtering); harmonized integration of sources (REDCap exports) into phenotype and feature tables.
Description
Automatic transcriptions generated using OpenAI Whisper Large; transcripts of free-speech audio removed in this release.
Description
Original audio waveforms were collected but are not included in v1.0 distribution; only derived data (e.g., spectrograms, features) are released.
Develop and evaluate AI/ML methods for health-related inference from voice, including detection, classification, and risk stratification across specified disorder cohorts.
Description
License
Bridge2AI Voice Registered Access License
Data Use Agreement
Bridge2AI Voice Registered Access Agreement
Access Policy
Only credentialed users who sign the DUA can access the files
Required Training
TCPS 2: CORE 2022
Versioned Dois
version-specific (https://doi.org/10.57764/qb6h-em84) and latest (https://doi.org/10.57764/3sg0-7440)