Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before
3OT2OD032720-01S1
Response
Create an ethically sourced, diverse, multi-institutional voice dataset linked to health information to enable AI research on voice as a biomarker of health.
📊
Composition
What do the instances represent?
Representation
Voice-derived data (spectrograms and acoustic/phonetic/prosodic features) and associated phenotype/clinical data.
Instance Type
Participants, sessions, and recordings; derived data per recording; phenotype per participant.
Data Type
Derived features from audio (spectrograms 513xN; static acoustic features), plus tabular phenotype data and accompanying data dictionaries.
Counts
12,523
Label
Not applicable; dataset includes derived features and task labels; free speech transcripts removed.
Sampling Strategies
Is Sample
True
Is Random
False
Source Data
Patients at specialty clinics across five North American sites.
Is Representative
No (targeted disease cohorts rather than a general population sample).
Representative Verification
Not applicable for targeted cohort recruitment.
Why Not Representative
Participants selected based on membership in predefined disease cohorts (respiratory, voice, neurological, mood); adult cohort only in v1.0.
Strategies
Targeted recruitment at specialty clinics using inclusion/exclusion criteria and standardized protocols.
Missing Information
Missing
Original audio waveforms (omitted in v1.0).
Transcripts of free speech audio (removed).
Why Missing
Privacy protection and de-identification for low-risk release.
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information v1.0
Bridge2AI-Voice is an ethically sourced, multi-institutional dataset linking voice-derived data with clinical and demographic information to enable research on voice as a biomarker of health. The v1.0 release includes 12,523 recordings from 306 adult participants collected across five sites in North America. Participants were selected from disease cohorts where voice and speech changes are clinically relevant (voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, and respiratory disorders). This initial release provides low-risk derived data (e.g., spectrograms and acoustic/phonetic/prosodic features) and detailed phenotype data; original audio waveforms and free speech transcripts are not included to protect privacy. Data collection followed a standardized protocol with IRB/REB oversight, and preprocessing used established audio analysis tools. Access is credentialed and governed by a registered-access license, data use agreement, and required training.
Addresses the lack of large, high-quality, diverse, multi-institutional voice datasets with standardized protocols and linked health information needed for clinically relevant AI research.
Description
HIPAA Safe Harbor identifiers removed (e.g., names, contact details, precise dates, device IDs, medical identifiers, and other unique identifiers).
State and province removed; country of data collection retained.
Transcripts of free speech audio removed.
Original audio waveforms omitted in v1.0; only spectrograms and derived features are provided.
Description
Health-related demographic, clinical, and validated questionnaire data.
Description
Clinical information and questionnaire responses distributed under registered access with DUA.
Description
Voice recordings collected in clinic using a standardized protocol via a custom tablet application; headset microphone used when possible.
Demographics, clinical questionnaires, and confounders collected via the same application; exported from REDCap.
Standardized multi-site protocol with IRB/REB oversight; derived features computed using established tools.
Description
Custom tablet application for data capture with headset when possible.
Data export and conversion from REDCap using the open-source b2aiprep library.
Description
Project investigators at specialty clinics across five North American sites; recruitment based on inclusion/exclusion criteria prior to clinic visits.
Description
Collected during clinic visits across five North American sites; typically single-session per participant with some multi-session participants.
Description
Approved by the University of South Florida Institutional Review Board (IRB).
Submitted for review to the University of Toronto Research Ethics Board (REB).
Description
Raw audio converted to mono and resampled to 16 kHz with a Butterworth anti-aliasing filter.
Spectrograms computed via STFT using 25 ms window, 10 ms hop, 512-point FFT (yielding 513xN spectrograms).
Acoustic features extracted with OpenSMILE; phonetic/prosodic features computed with Parselmouth and Praat.
Transcriptions generated using OpenAI's Whisper Large model (free speech transcripts removed in release).
Used Software
Name
URL
OpenSMILE
https://audeering.github.io/opensmile/
Parselmouth
https://parselmouth.readthedocs.io/
Praat
http://www.fon.hum.uva.nl/praat/
Torchaudio
https://pytorch.org/audio
OpenAI Whisper Large
https://openai.com/research/whisper
b2aiprep
https://github.com/sensein/b2aiprep
Description
De-identification per HIPAA Safe Harbor; removal of state/province; removal of free speech transcripts; omission of original audio waveforms in v1.0.
Description
Automatic speech transcription using OpenAI Whisper Large (with free speech transcripts excluded from release).
Validated questionnaires for demographic and clinical variables.
Used Software
Name
OpenAI Whisper Large
Custom tablet data collection application
Description
Raw audio recordings were collected but are not released in v1.0; only derived spectrograms and features are provided. Future releases may include voice data with additional safeguards.