Create an ethically sourced, diverse, multi-institutional voice dataset linked to health information to enable future AI research and insights into voice as a biomarker of health.
Grantor
Grant Name
Grant Number
National Institutes of Health (NIH)
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
Name
Instance structure
Representation
Voice recordings linked to clinical and questionnaire information
Instance Type
Participants, recording sessions, and derived per-recording data
Data Type
Derived data only: STFT spectrograms, acoustic/phonetic/prosodic features, and (non-free-speech) transcriptions; tabular phenotype data per participant.
Counts
12,523
Label
Dataset includes cohort membership and clinical/questionnaire variables; no explicit task labels provided in this release.
Sampling Strategies
Name
Sampling
Is Sample
True
Is Random
False
Source Data
Patients at specialty clinics across five sites in North America
Is Representative
False
Why Not Representative
Targeted enrollment of disease cohorts with known voice manifestations
Strategies
Purposive/clinical cohort-based sampling
Missing Information
Name
Disease cohorts
Identification
Voice disorders
Neurological and neurodegenerative disorders
Mood and psychiatric disorders
Respiratory disorders
Distribution
Adult cohort only in v1.0; detailed cohort counts not provided.
Name
Release files
Description
spectrograms.parquet (derived spectrogram data)
static_features.tsv and static_features.json (per-recording features and data dictionary)
phenotype.tsv and phenotype.json (per-participant phenotype data and data dictionary)
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice v1.0 is a restricted-access dataset enabling research into voice as a biomarker of health. The initial release provides 12,523 recordings for 306 adult participants collected across five sites in North America, with accompanying demographic, clinical, and validated questionnaire data. To reduce re-identification risk, only derived data are released (e.g., spectrograms and extracted acoustic/phonetic features); original audio waveforms and free-speech transcripts are not included. Participants were selected from cohorts with conditions known to manifest in voice (voice disorders, neurological and neurodegenerative disorders, mood and psychiatric disorders, and respiratory disorders). Standardized collection protocols were used; raw audio was converted to mono, resampled to 16 kHz with a Butterworth anti-aliasing filter, and used to compute STFT spectrograms and features (openSMILE, Praat/Parselmouth, torchaudio). Transcriptions were generated using Whisper Large but free speech transcripts were removed for release.
Addresses the lack of large, high-quality, diverse, standardized, multi-institutional voice datasets linked to health biomarkers suitable for AI research.
Name
Bridge2AI-Voice Team
Name
Adult cohort v1.0
Is Data Split
no
Is Subpopulation
Adult cohort only in v1.0
Name
Participant-session-recording linkage
Description
Participants may have multiple sessions; sessions include multiple recordings/tasks.
State/province removed; country of data collection retained.
Free speech transcripts removed.
Original audio waveforms omitted from v1.0; only derived data released.
Name
Sensitive data
Description
Contains health-related and demographic information.
Name
Data acquisition
Description
Standardized protocol with voice tasks (e.g., sustained vowel), clinical and targeted questionnaires.
Participants enrolled from specialty clinics into predefined cohorts.
Was Directly Observed
yes (voice recordings)
Was Reported By Subjects
yes (questionnaires)
Was Inferred Derived
yes (features and spectrograms derived from raw audio)
Was Validated Verified
Standardized acquisition protocols were used; preprocessing and feature extraction applied consistently.
Name
Collection mechanisms
Description
Custom tablet application; headset used for data collection when possible.
Data exported and converted from REDCap using an open-source library.
Name
Data collectors
Description
Project investigators at five North American sites; participant screening against inclusion/exclusion criteria.
Name
Ethics and IRB/REB review
Description
Data collection and sharing approved by the University of South Florida Institutional Review Board.
Submitted for review to the University of Toronto Research Ethics Board.
Name
Audio preprocessing and feature extraction
Description
Mono conversion; resampled to 16 kHz with a Butterworth anti-aliasing filter.
STFT spectrograms using 25 ms window, 10 ms hop, 512-point FFT.
Acoustic features via openSMILE.
Phonetic/prosodic measures via Parselmouth and Praat.
Derived audio features via torchaudio.
Used Software
Name
URL
b2aiprep
https://github.com/sensein/b2aiprep
openSMILE
Praat
Parselmouth
torchaudio
Name
Transcription
Description
Transcriptions generated using OpenAI Whisper Large; free speech transcripts removed from release.
Name
De-identification and release filtering
Description
Removal of HIPAA Safe Harbor identifiers.
Removal of state/province; retention of country only.
Exclusion of free speech transcripts and original audio waveforms from v1.0.
Name
Raw audio
Description
Collected but not released in v1.0; derived spectrograms and features provided instead.
Name
Hosting and support
Description
Health Data Nexus; Temerty Centre for AI Research and Education in Medicine (supported by the Temerty Foundation).
🚀
Uses
What (other) tasks could the dataset be used for?
Name
Intended tasks
Response
Development and evaluation of AI methods for health-related voice analytics using derived spectrograms and acoustic/phonetic features; exploration of associations between voice signals and clinical/demographic factors.
Name
License and access terms
Description
License
Bridge2AI Voice Registered Access License.
Access Policy
Only credentialed users who sign the DUA can access the files.