Create an ethically sourced, diverse, multi-institutional voice dataset linked to clinical information to enable AI research on voice as a biomarker of health.
Grantor
Grant Name
Grant Number
National Institutes of Health
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before.
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
Name
Instance description
Representation
Voice recordings with derived spectrograms/features and participant-level phenotype/clinical data.
Instance Type
participants, sessions, and recordings (participant_id, session_id, task_name).
Data Type
Derived spectrogram arrays; acoustic, phonetic, and prosodic features; limited transcriptions; tabular phenotype and feature files.
Counts
12,523
Label
Cohort/disease group membership and clinical/phenotype variables; no raw audio included in v1.0.
Sampling Strategies
Name
Targeted clinical cohort sampling
Is Sample
sample from patients presenting at specialty clinics at five North American sites
Is Random
False
Source Data
Adults with conditions affecting voice (voice, neurological/neurodegenerative, mood/psychiatric, respiratory disorders)
Is Representative
not stated
Why Not Representative
Targeted sampling of specific disorders; pediatric cohort not included in v1.0
Strategies
targeted clinical cohort sampling at participating sites
Name
Cohort categories
Identification
Disease Cohorts Identified At Enrollment
Voice disorders; Neurological and Neurodegenerative; Mood and Psychiatric; Respiratory; Pediatric (planned but not included in v1.0).
Distribution
Adult cohort only in v1.0; 306 participants across five sites in North America.
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
The Bridge2AI-Voice project presents a comprehensive, ethically sourced dataset enabling research on voice as a biomarker of health. Bridge2AI-Voice v1.0 provides 12,523 recordings for 306 adult participants collected across five sites in North America, with corresponding demographic, clinical, and validated questionnaire information. The initial release is considered low risk and includes derived data (e.g., spectrograms, acoustic/phonetic/prosodic features) and data dictionaries; original voice audio waveforms are not included. Participants were enrolled in predetermined groups reflecting conditions that manifest in voice (voice disorders, neurological and neurodegenerative disorders, mood and psychiatric disorders, respiratory disorders), with pediatric cohorts planned but not included in v1.0. Data were collected via a standardized protocol and subsequently de-identified using HIPAA Safe Harbor principles.
Standardized data collection protocol; IRB/REB oversight.
Name
Collection mechanisms
Description
Custom tablet application for standardized data capture; headset microphone when feasible.
Export and conversion from REDCap using an open-source b2aiprep library.
Used Software
Name
URL
Version
REDCap
3.20.0
b2aiprep
https://github.com/sensein/b2aiprep
Name
Data collection team
Description
Project investigators at five North American clinical sites.
Name
Ethics and review
Description
Data collection and sharing approved by the University of South Florida Institutional Review Board.
Submitted for review to the University of Toronto Research Ethics Board.
Name
Audio preprocessing and feature extraction
Description
Raw audio converted to mono and resampled to 16 kHz with a Butterworth anti-aliasing filter.
Spectrograms computed via STFT with 25 ms window, 10 ms hop, 512-point FFT.
Acoustic features via OpenSMILE; phonetic/prosodic features via Parselmouth and Praat.
Transcriptions generated with OpenAI's Whisper Large model (free-speech transcripts later removed from release).
Used Software
Name
Version
OpenSMILE
Parselmouth
Praat
Torchaudio
2.1
OpenAI Whisper Large
Name
De-identification and release filtering
Description
HIPAA Safe Harbor removal; state/province removed; country retained.
Free-speech transcripts removed; only derived data released (no raw audio).
Name
Transcription
Description
Automatic transcriptions generated using OpenAI's Whisper Large model; free-speech transcripts removed prior to release.
Used Software
Name
OpenAI Whisper Large
Name
Raw audio availability
Description
Raw audio waveforms are not included in v1.0; only spectrograms and derived features are provided. Future releases aim to include voice waveforms with additional security precautions.