Create an ethically sourced flagship dataset to enable AI research on voice as a biomarker of health by linking derived voice data with demographic, clinical, and validated questionnaire information.
Grantor
Grant Name
Grant Number
National Institutes of Health (NIH)
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
Counts
Data Type
Instance Type
Name
Representation
12523
Derived spectrograms (513 x N), acoustic, phonetic, and prosodic features extracted from raw audio; ...
Audio-derived data (per recording)
Voice-derived recordings
Voice-derived recordings (spectrograms/features) linked to metadata
306
Tabular phenotype data (one row per participant) with data dictionary
Participant-level records
Participants
Participants with linked demographic, clinical, and questionnaire data
Name
Adult cohort
Identification
v1.0 includes adults only across five disease cohorts
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information v1.0
Bridge2AI-Voice is a comprehensive collection of data derived from voice recordings linked to corresponding clinical information to enable research on voice as a biomarker of health. Version 1.0 provides 12,523 recordings for 306 participants collected across five sites in North America. Participants were selected based on conditions known to manifest in the voice waveform (voice, neurological, mood/psychiatric, and respiratory disorders). This initial release contains low-risk derived data (e.g., spectrograms and extracted features) and detailed demographic/clinical/questionnaire data; raw audio waveforms are not included in v1.0.
Addresses the lack of large, diverse, multi-institutional voice datasets linked to health information with standardized collection protocols and ethical oversight; mitigates prior limitations such as small datasets and limited demographic diversity reporting.
Name
Sampling approach
Is Sample
True
Is Random
False
Source Data
Patients presenting at specialty clinics/institutions at five North American sites
Is Representative
False
Why Not Representative
Purposeful selection into five disease cohorts; v1.0 includes adults only
Strategies
Deterministic cohort-based enrollment using inclusion/exclusion screening
Name
Data collection mechanisms
Description
Standardized protocol using a custom tablet application; headset used when possible
Clinical/demographic and validated questionnaires collected in-app
Data exported and converted from REDCap; processing via open-source b2aiprep library
Name
Data collection personnel
Description
Project investigators at specialty clinics and institutions screened and enrolled participants
Name
Collection timeframe
Description
Collected across five sites in North America; specific calendar dates not specified in this record
Name
Ethics and IRB/REB review
Description
Approved by the University of South Florida Institutional Review Board
Submitted for review to the University of Toronto Research Ethics Board
Name
Audio preprocessing and feature derivation
Description
Raw audio converted to mono, resampled to 16 kHz with Butterworth anti-aliasing filter
Spectrograms via STFT (25 ms window, 10 ms hop, 512-point FFT)
Acoustic features via OpenSMILE
Phonetic/prosodic features via Parselmouth and Praat (F0, formants, voice quality)
Transcriptions generated using OpenAI Whisper Large (free speech transcripts not included in v1.0)
Used Software
Name
OpenSMILE
Parselmouth
Praat
OpenAI Whisper Large
torchaudio
b2aiprep
Name
Transcription and derived annotations
Description
Automatic transcription using OpenAI Whisper Large for certain tasks; free speech transcripts were removed prior to release
Used Software
Name
OpenAI Whisper Large
Name
Raw audio availability
Description
Raw audio waveforms were not distributed in v1.0; only derived spectrograms and features are provided. Future releases aim to include audio with additional safeguards.
External Resources
Name
https://docs.b2ai-voice.org
Documentation website
https://doi.org/10.5281/zenodo.14148755
Bridge2AI Voice REDCap (v3.20.0)
Name
Confidentiality considerations
Description
Contains clinical and questionnaire data linked to participants; v1.0 includes only low-risk derived data and excludes raw audio
Name
Sensitive data elements
Description
Health-related data (demographics, clinical information, validated questionnaires)
State/province removed; country of data collection retained
Free speech transcripts removed
Raw audio waveforms omitted from this release
Name
Health Data Nexus
Description
Supported by the Temerty Centre for AI Research and Education in Medicine (Temerty Foundation)
Mixed (tabular phenotype/features and array-based parquet spectrograms)
🚀
Uses
What (other) tasks could the dataset be used for?
Name
Intended tasks
Response
Develop and evaluate AI/ML methods for detecting, characterizing, and monitoring health conditions from voice-derived representations and associated clinical data.
Name
Access, license, and terms
Description
License
Bridge2AI Voice Registered Access License
Data Use Agreement
Bridge2AI Voice Registered Access Agreement
Access
Credentialed users only; must sign DUA
Required Training
TCPS 2: CORE 2022
This is a restricted-access resource distributed via Health Data Nexus