Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before
3OT2OD032720-01S1
Response
Create an ethically sourced flagship voice dataset to enable AI research on voice as a biomarker of health and support clinical insights.
π
Composition
What do the instances represent?
Counts
Data Type
Instance Type
Name
Representation
12523
Spectrogram matrices (513 x N), static acoustic/phonetic/prosodic feature vectors
recording-derived instance
Recordings-derived instances
Derived artifacts from voice recordings (e.g., spectrograms and engineered features)
306
Demographics, clinical information, and validated questionnaire responses
participant
Participants
Individual study participants enrolled across five North American sites
Identification
Adult cohort only in v1.0; participants selected based on known conditions with voice manifestations (voice, neurological, mood/psychiatric, respiratory).
Distribution
Counts by subgroup not specified.
Description
spectrograms.parquet β Parquet file with 513 x N spectrogram matrices and identifiers (participant_id, session_id, task_name).
phenotype.tsv β Tab-delimited participant-level data (demographics, acoustic confounders, validated questionnaires).
phenotype.json β Data dictionary for phenotype.tsv.
static_features.tsv β Recording-level engineered features (OpenSMILE, Praat, Parselmouth, torchaudio).
static_features.json β Data dictionary for static_features.tsv.
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically sourced dataset of data derived from voice recordings linked to corresponding clinical and demographic information, intended to enable AI research on voice as a biomarker of health. Version 1.0 provides 12,523 recordings for 306 participants collected across five sites in North America. Participants were selected based on conditions known to manifest in the voice waveform, including voice disorders, neurological and neurodegenerative disorders, mood and psychiatric disorders, and respiratory disorders. The initial release contains low-risk derived data (e.g., spectrograms and engineered features) and does not include original audio waveforms. Detailed demographic, clinical, and validated questionnaire data are also available.
Address the need for a large, high-quality, multi-institutional, and diverse voice database linked to health biomarkers with standardized collection protocols.
Name
Adult cohort v1.0
Description
Initial dataset release includes only the adult cohort with derived data from voice recordings and linked clinical/phenotype data.
Is Subpopulation
yes
Strategies
Purposeful sampling of patients at specialty clinics into five predetermined disease/cohort groups (Respiratory, Voice, Neurological, Mood, Pediatric)
Is Sample
yes
Is Random
no
Source Data
Patients presenting at specialty clinics/institutions across five North American sites
Is Representative
no
Why Not Representative
Clinic-based cohort with inclusion/exclusion criteria; participants selected based on known conditions affecting voice.
Description
Participants may have one or more sessions; sessions include multiple recording tasks. Each spectrogram/feature row links via participant_id, session_id, and task_name.
Description
No recommended train/validation/test splits are provided in v1.0.
State and province removed; country of data collection retained.
Transcripts of free speech audio removed.
Original audio waveforms omitted in v1.0; only derived data are distributed.
Description
Health-related information (clinical and questionnaire data) linked to voice-derived features.
Description
Voice recordings directly observed; questionnaires reported by subjects; multiple derived features computed from raw audio.
Was Directly Observed
yes
Was Reported By Subjects
yes
Was Inferred Derived
yes
Was Validated Verified
Standardized data collection protocol; derived features generated via established tools (OpenSMILE, Praat/Parselmouth, torchaudio). Ethics approvals in place.
Description
Standardized protocol; data collected via custom tablet application using headset when possible; export and conversion from REDCap using an open-source library.
Description
Project investigators at specialty clinics/institutions across five sites in North America; participants consented prior to data collection.
Description
Data collection and sharing approved by the University of South Florida Institutional Review Board; submitted for review to the University of Toronto Research Ethics Board.
Description
Eligible patients provided informed consent for data collection and for sharing acquired research data prior to participation.
Description
Participants were screened and informed as part of the consent process for the data collection initiative and data sharing.
Description
Raw audio converted to mono and resampled to 16 kHz with a Butterworth anti-aliasing filter; derived spectrograms, acoustic, phonetic/prosodic features, and transcriptions generated.
Used Software
Name
URL
openSMILE
Praat
Parselmouth
torchaudio
OpenAI Whisper Large
b2aiprep
https://github.com/sensein/b2aiprep
Description
De-identification and removal of identifiers per HIPAA Safe Harbor; removal of free speech transcripts; exclusion of original audio in v1.0.
Description
Automatic transcription using OpenAI Whisper Large; validated questionnaires administered during clinical data collection.
Description
Original audio waveforms were collected but are not included in v1.0; only derived data are available. Future releases aim to include audio with additional safeguards.
Description
None specified.
Description
Health Data Nexus (Temerty Centre for AI Research and Education in Medicine)
π
Uses
What (other) tasks could the dataset be used for?
Response
Develop and evaluate AI/ML models using derived voice features (spectrograms, acoustic, phonetic, an...
Study associations between voice-derived features and health conditions (voice, neurological, mood/p...
Description
Not specified.
Description
Screening/monitoring of respiratory conditions using cough/breath/voice-derived features; detection and characterization of voice disorders; analysis of neurological and mood/psychiatric condition markers in voice features.
Description
Users should note that v1.0 includes only de-identified derived features without raw audio; future inclusion of audio will require additional safeguards.
Description
Not specified.
Description
Access Policy
Only credentialed users who sign the Data Use Agreement (DUA) can access files.
License
Bridge2AI Voice Registered Access License.
Data Use Agreement
Bridge2AI Voice Registered Access Agreement.
Required Training
TCPS 2: CORE 2022.
Versioned DOIs provided for citation and discovery.