Create an ethically sourced flagship dataset to enable AI research on voice as a biomarker of health and support insights into links between acoustic markers and health conditions.
Grantor
Grant Name
Grant Number
National Institutes of Health (NIH)
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
ID
instances-v1-0
Name
Instance definition
Representation
Voice recordings (not released in v1.0) and derived data (spectrograms, acoustic/phonetic/prosodic features), with participant-level phenotype data (demographics, clinical, validated questionnaires).
Instance Type
Participants and recording sessions/recordings
Data Type
Derived spectrograms and features (from standardized audio), plus participant phenotype data; no raw audio in v1.0.
Counts
12,523
Label
Not specified; dataset includes clinical and questionnaire variables that can be used as targets.
Sampling Strategies
ID
sampling-v1-0
Name
Sampling strategy
Is Sample
yes
Is Random
no
Source Data
Specialty clinics at five North American sites
Is Representative
Not explicitly validated as representative of the general population
Why Not Representative
Clinic-based, condition-targeted sampling across predefined cohorts
ID
subpops-cohorts
Name
Cohorts
Identification
Predefined disease cohorts: voice disorders, neurological disorders, mood disorders, respiratory disorders, and pediatric (pediatric not included in v1.0).
Distribution
v1.0 includes adult cohort only; 306 participants; 12,523 recordings.
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically sourced dataset derived from voice recordings and linked to health information to enable research on voice as a biomarker of health. Version 1.0 provides 12,523 recordings for 306 adult participants collected across five sites in North America. Participants were selected from cohorts where conditions manifest in the voice waveform, including voice disorders, neurological disorders, mood disorders, respiratory disorders, and a pediatric cohort (pediatric data not included in v1.0). The initial release contains low-risk, de-identified derived data (e.g., spectrograms and acoustic/phonetic/prosodic features) and detailed demographic, clinical, and validated questionnaire data; original audio waveforms and free-speech transcripts are not included in v1.0.
Addresses the lack of large, high-quality, multi-institutional, diverse, ethically sourced voice datasets linked to health information.
ID
mech-protocol
Name
Collection mechanisms
Description
Standardized multi-site protocol with a custom tablet application; headset used for data collection when possible.
ID
collectors-sites
Name
Data collectors
Description
Project investigators and site personnel at five North American sites
ID
irb-approvals
Name
Ethical review
Description
Data collection and sharing approved by the University of South Florida Institutional Review Board; submitted for review to the University of Toronto Research Ethics Board.
ID
acquisition-methods
Name
Instance acquisition
Description
Directly observed voice tasks (e.g., sustained phonation) recorded under a standardized protocol; demographic, clinical, and validated questionnaires collected via app; derived features computed from standardized audio.
Was Directly Observed
yes
Was Reported By Subjects
yes
Was Inferred Derived
yes
Was Validated Verified
yes
ID
preprocessing-audio
Name
Audio preprocessing and feature derivation
Description
Raw audio standardized to mono and 16 kHz with a Butterworth anti-aliasing filter; STFT spectrograms with 25 ms window, 10 ms hop, 512-point FFT.
Acoustic features via OpenSMILE; phonetic/prosodic features via Parselmouth/Praat; additional audio processing via torchaudio; transcriptions generated using OpenAI Whisper Large (transcripts of free speech removed from release).
Used Software
ID
Name
URL
Version
software-opensmile
openSMILE
https://audeering.github.io/opensmile/
software-praat
Praat
https://www.fon.hum.uva.nl/praat/
software-parselmouth
Parselmouth (Python interface to Praat)
https://parselmouth.readthedocs.io/
software-torchaudio
torchaudio
https://pytorch.org/audio/
2.1
software-whisper
OpenAI Whisper Large
https://github.com/openai/whisper
ID
labeling-transcription
Name
Transcription (removed in v1.0 release)
Description
Automatic transcriptions were generated using OpenAI Whisper Large as part of derivation; free-speech transcripts were removed from the public release for de-identification.
Parquet file storing dense spectrogram data derived from standardized audio; includes participant_id...
file-spectrograms-parquet
application/x-parquet
spectrograms.parquet
spectrograms.parquet
Tab-delimited participant-level data (demographics, acoustic confounders, validated questionnaires);...
file-phenotype-tsv
text/tab-separated-values
phenotype.tsv
phenotype.tsv
Data dictionary describing columns in phenotype.tsv.
JSON
file-phenotype-json
application/json
phenotype.json
phenotype.json
Derived acoustic/phonetic/prosodic features with one row per recording.
file-static-features-tsv
text/tab-separated-values
static_features.tsv
static_features.tsv
Data dictionary describing columns in static_features.tsv.
JSON
file-static-features-json
application/json
static_features.json
static_features.json
ID
ip-restrictions
Name
IP restrictions
Description
Not specified.
ID
export-regulatory
Name
Export/regulatory restrictions
Description
Not specified.
Source clinical/phenotype data collected via custom app and exported from REDCap (Bridge2AI Voice REDCap v3.20.0; https://doi.org/10.5281/zenodo.14148755). Project documentation: "https://docs.b2ai-voice.org/"
🚀
Uses
What (other) tasks could the dataset be used for?
ID
task-voice-ai
Name
Target tasks
Response
AI/ML research using derived voice features and clinical/phenotype data, such as disorder detection, stratification, and exploration of voice–health associations.
ID
license-registered-access
Name
License and use terms
Description
Access policy: Only credentialed users who sign the DUA can access the files.
Data Use Agreement: Bridge2AI Voice Registered Access Agreement.
Required training: TCPS 2: CORE 2022.
ID
other-tasks
Name
Other potential tasks
Description
Exploratory analyses of voice–health associations; development/evaluation of AI models for condition screening and monitoring using derived voice features.