Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before
3OT2OD032720-01S1
Response
Create an ethically sourced, diverse, multi-institutional voice dataset linked to health information to enable AI research on voice as a biomarker of health.
📊
Composition
What do the instances represent?
Counts
Data Type
Instance Type
Representation
Sampling Strategies
12523
Derived spectrograms (513 x N), static acoustic/phonetic features (OpenSMILE, Praat, Parselmouth, to...
Audio recordings and sessions (time-frequency representations and derived features; raw waveforms om...
Voice-derived recordings
{'strategies': ['Participants prospectively selected into 5 predetermined clinical groups (Respiratory, Voice, Neurological, Mood/Psychiatric, Pediatric); v1.0 includes adult cohort only.'], 'source_data': ['Five clinical sites in North America'], 'is_representative': ['No; targeted clinical cohorts'], 'why_not_representative': ['Cohorts intentionally sampled for disorders with known voice manifestations; not a population-representative sample.']}
306
Demographics, validated health questionnaires, acoustic confounders, and disease-specific clinical i...
Individuals (adult cohort in v1.0)
Participants
Identification
Adult cohort only in v1.0
Disorder Cohorts
Voice, Neurological/Neurodegenerative, Mood/Psychiatric, Respiratory, Pediatric (pediatric not included in v1.0)
Distribution
306 adult participants across five North American sites
Participants selected based on membership in predefined clinical cohorts
Description
spectrograms.parquet (Parquet; time-frequency representations per recording)
static_features.tsv (tab-delimited; one row per recording)
static_features.json (data dictionary for features)
phenotype.tsv (tab-delimited; one row per participant)
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically sourced dataset to enable research on the human voice as a biomarker of health. Version 1.0 provides 12,523 recordings for 306 adult participants collected across five sites in North America, selected based on conditions known to manifest in the voice waveform (voice disorders, neurological/neurodegenerative disorders, mood and psychiatric disorders, and respiratory disorders). This initial release contains low-risk derived data (e.g., spectrograms and acoustic/phonetic features) and detailed demographic, clinical, and validated questionnaire data. Original audio waveforms are not included in v1.0. Standardized collection protocols, de-identification under HIPAA Safe Harbor, and data access via registered, credentialed workflows are used to protect participants and enable responsible research.
Address the lack of large, standardized, demographically diverse, multi-institutional voice datasets linked to health biomarkers to improve external validity and clinical utility in voice AI research.
Strategies
Targeted, multi-site clinical recruitment into 5 disorder cohort categories; adult cohort only in v1.0.
Source Data
Specialty clinics across five North American sites
Is Representative
No; targeted clinical cohorts
Why Not Representative
Designed to cover diverse voice-related conditions rather than represent the general population.
Description
Data directly observed via standardized voice recording tasks (e.g., sustained vowel phonation) plus participant-reported questionnaires and clinical data; derived features computed from raw audio.
Was Directly Observed
yes
Was Reported By Subjects
yes
Was Inferred Derived
yes
Was Validated Verified
yes
Description
Standardized protocol using a custom tablet application; headset used for data collection when possible; data exported from REDCap using an open-source library; multiple sessions for some participants as needed.
Description
Project investigators at specialty clinics across five North American sites
Description
Data collection and sharing approved by the University of South Florida Institutional Review Board; submitted for review to the University of Toronto Research Ethics Board.
Description
Raw audio converted to monaural and resampled to 16 kHz with a Butterworth anti-aliasing filter.
Spectrograms computed via short-time FFT (25 ms window, 10 ms hop, 512-point FFT).
Acoustic features extracted with OpenSMILE; phonetic/prosodic features computed with Parselmouth and Praat; additional features via torchaudio.
Used Software
Name
URL
openSMILE
https://audeering.github.io/opensmile/
Praat
https://www.fon.hum.uva.nl/praat/
Parselmouth
https://parselmouth.readthedocs.io/
torchaudio
https://pytorch.org/audio
b2aiprep
https://github.com/sensein/b2aiprep
Description
HIPAA Safe Harbor identifiers removed.
State and province removed; country of data collection retained.
Transcripts of free speech audio removed to reduce re-identification risk.
Description
Machine-generated transcriptions produced using OpenAI Whisper Large model for certain tasks.
Clinical information and validated questionnaire responses collected during visits.
Description
Health-related data linked to voice-derived features.
Voice as a potential biometric/biobehavioral marker.
Description
HIPAA Safe Harbor identifiers removed.
State and province removed; country retained.
Free speech transcripts removed.
Original audio waveforms omitted from v1.0.
Description
Dataset is available to external, credentialed users under registered access terms (DUA and required training).
Description
Health Data Nexus (Temerty Centre for AI Research and Education in Medicine)
Supported by the Temerty Foundation
🚀
Uses
What (other) tasks could the dataset be used for?
Response
Support AI methods development and clinical research using derived voice representations (e.g., spectrograms and acoustic features) linked with demographic, clinical, and questionnaire data across targeted disorder cohorts.
Description
License
Bridge2AI Voice Registered Access License.
Access Policy
Only credentialed users who sign the DUA can access files.
Data Use Agreement
Bridge2AI Voice Registered Access Agreement.
Required Training
TCPS 2: CORE 2022.
Access provided via Health Data Nexus credentialed workflow.
Older versions discoverable via DOI versioning; documentation website will communicate updates and changes.
🔄
Maintenance
How will the dataset be maintained?
1.0
Description
Future releases aim to include original voice waveforms with additional security safeguards; updates and documentation via https://docs.b2ai-voice.org.
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema