Create an ethically sourced flagship dataset to enable AI research on the use of voice as a biomarker of health by linking diverse voice recordings with clinical and demographic information.
Grantor
Grant Name
Grant Number
National Institutes of Health (NIH)
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
Name
Instance class
Representation
Voice recordings linked to participant-level clinical, demographic, and questionnaire data; derived spectrogram matrices and acoustic/phonetic/prosodic features per recording.
Instance Type
Participants, recording sessions, and recordings (per task); derived data instances per recording.
Data Type
Derived features (e.g., spectrograms, acoustic, phonetic, prosodic) from standardized audio; raw audio waveforms are withheld in v1.0.
Counts
12,523
Name
Adult cohort
Identification
Adult participants meeting inclusion criteria within predefined disease cohorts
Distribution
306 participants; 12,523 recordings across multiple recording tasks
Description
Parquet (.parquet) for spectrograms
Tab-delimited text (.tsv) for phenotype and static features
JSON (.json) data dictionaries
Description
2024-11-27 (initial public release v1.0)
🔍
Collection Process
How was the data acquired?
bridge2ai-voice-v1.0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice v1.0 is an ethically sourced, multi-institutional voice dataset focused on the use of voice as a biomarker of health. The initial release provides 12,523 recordings for 306 adult participants collected across five sites in North America. Participants were selected from five predefined cohorts where voice/speech changes are associated with disease: voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, respiratory disorders, and pediatric voice/speech disorders (note: v1.0 includes adults only). This release contains low-risk derived data (e.g., spectrograms and acoustic/phonetic/prosodic features) and corresponding demographic, clinical, and validated questionnaire data; raw audio waveforms are not included in v1.0. Documentation: "https://docs.b2ai-voice.org"
Address the lack of large, high-quality, multi-institutional, diverse voice datasets linked to other health biomarkers and collected under standardized, ethically governed protocols.
Name
Cohort sampling strategy
Is Sample
True
Is Random
False
Source Data
Patients at specialty clinics across five sites in North America
Is Representative
False
Why Not Representative
Participants were recruited based on membership in predefined disease cohorts (voice, neurological, mood/psychiatric, respiratory, pediatric), and v1.0 includes adult cohort only.
Strategies
Purposive cohort-based sampling within specialty clinics
Description
Health-related clinical information and demographics are included; identifiers removed per HIPAA Safe Harbor.
Description
Data were collected at specialty clinics using a standardized protocol capturing demographics, health questionnaires (including validated instruments), targeted confounder questions, disease-specific information, and voice tasks (e.g., sustained vowel).
Data reported by subjects (questionnaires) and directly observed (voice recordings); some data are indirectly derived (features, transcriptions).
Was Directly Observed
yes (voice recordings)
Was Reported By Subjects
yes (questionnaires)
Was Inferred Derived
yes (features, transcriptions)
Description
Custom tablet-based data collection application; headset used for recording when possible; data exported from REDCap using an open-source library developed by the team.
Used Software
Name
Bridge2AI-Voice tablet data collection app
REDCap
b2aiprep
Description
Project investigators and clinical teams at five North American sites
Description
Data collection and sharing approved by the University of South Florida Institutional Review Board; submission to the University of Toronto Research Ethics Board noted.
Description
Raw audio converted to monaural, resampled to 16 kHz with a Butterworth anti-aliasing filter; spectrograms computed via STFT (25 ms window, 10 ms hop, 512-point FFT).
Acoustic features extracted with OpenSMILE; phonetic/prosodic features computed with Parselmouth and Praat; transcriptions generated using OpenAI Whisper Large.
Code to preprocess and merge source data provided via the b2aiprep library.
Used Software
Name
openSMILE
Parselmouth
Praat
torchaudio
OpenAI Whisper Large
b2aiprep
Description
HIPAA Safe Harbor identifiers removed; state/province removed; country retained.
Transcripts of free speech audio removed from the release.
Only derived data (e.g., spectrograms, features) included in v1.0; audio waveforms omitted.
Description
Automated transcriptions generated using OpenAI Whisper Large; transcripts of free speech audio not released in v1.0.
Used Software
Name
OpenAI Whisper Large
Description
Raw voice audio collected but withheld from v1.0 distribution; planned for future releases with additional safeguards.
Description
Health Data Nexus (Temerty Centre for AI Research and Education in Medicine)
Description
HIPAA Safe Harbor de-identification applied; removal of direct identifiers and fine-grained dates; state/province removed; country retained; free-speech transcripts removed; no raw audio in v1.0.
Versioned DOIs indicate archival/versioning support via persistent identifiers
Restrictions
Registered access with DUA and required training
Description
ID
Media Type
Name
Path
Title
Parquet dataset containing 513 x N spectrogram matrices per recording, with participant_id, session_...
bridge2ai-voice-v1.0-spectrograms
application/x-parquet
spectrograms.parquet
spectrograms.parquet
Derived spectrograms
Tab-delimited file with one row per participant including demographics, acoustic confounders, and re...
bridge2ai-voice-v1.0-phenotype
text/tab-separated-values
phenotype.tsv
phenotype.tsv
Participant phenotype and questionnaires
JSON data dictionary mapping column names to descriptions for phenotype.tsv.
bridge2ai-voice-v1.0-phenotype-dict
application/json
phenotype.json
phenotype.json
Phenotype data dictionary
Tab-delimited file with one row per unique recording; includes features derived using openSMILE, Pra...
bridge2ai-voice-v1.0-static-features
text/tab-separated-values
static_features.tsv
static_features.tsv
Recording-level static features
JSON data dictionary mapping feature column names to descriptions for static_features.tsv.
bridge2ai-voice-v1.0-static-features-dict
application/json
static_features.json
static_features.json
Static features data dictionary
🚀
Uses
What (other) tasks could the dataset be used for?
Name
Task
Response
Support AI-driven analysis of voice (e.g., feature extraction, modeling, and evaluation) for health-related research across disease cohorts where voice/speech changes are clinically relevant.
Description
Restricting v1.0 to low-risk derived data (no raw audio) mitigates privacy risks but may limit certain modeling tasks requiring waveforms; future releases aim to include audio with additional safeguards.
Name
Bridge2AI Voice Registered Access
Description
Access restricted to credentialed users who sign the Bridge2AI Voice Registered Access Agreement (DUA) under the Bridge2AI Voice Registered Access License.