Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before
3OT2OD032720-01S1
Response
Create an ethically sourced, diverse, multi-institutional dataset of voice linked to clinical and questionnaire data to enable AI research on voice as a biomarker of health and support clinically meaningful insights.
📊
Composition
What do the instances represent?
Counts
Data Type
Instance Type
Missing Information
Representation
Sampling Strategies
12523
Derived features (spectrograms, acoustic, phonetic/prosodic); no raw audio waveforms in v1.0
Audio-derived instance
{'missing': ['Original audio waveforms'], 'why_missing': ['Omitted in initial low-risk release; planned for future releases with additional safeguards']}, {'missing': ['Transcripts of free speech audio'], 'why_missing': ['Removed during de-identification']}
Voice-derived recordings (spectrograms/features) per recording
{'is_sample': ['yes'], 'is_random': ['no'], 'source_data': ['Patients presenting at specialty clinics across five North American sites'], 'is_representative': ['no'], 'why_not_representative': ['Participants were selected based on membership in predefined disease cohorts'], 'strategies': ['Targeted enrollment by predefined disease categories']}
306
Demographics, clinical, and validated questionnaire responses
Participant
Participant-level phenotype records
{'is_sample': ['yes'], 'is_random': ['no'], 'source_data': ['Specialty clinics at five North American sites'], 'is_representative': ['no'], 'why_not_representative': ['Cohort-based selection for voice-relevant conditions']}
Participants selected based on membership in predefined disease cohorts
Description
Parquet
TSV
JSON
Description
2024-11-27
🔍
Collection Process
How was the data acquired?
bridge2ai-voice-v1.0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information v1.0
Bridge2AI-Voice is a multi-site, ethically sourced dataset of human voice linked to clinical and questionnaire information to enable research on voice as a biomarker of health. Version 1.0 includes 12,523 recordings from 306 participants across five North American sites, focusing on cohorts with known voice-relevant conditions (voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, and respiratory disorders). The initial release contains only low-risk, derived data (e.g., spectrograms and acoustic features) and de-identified phenotype tables; original audio waveforms and free-speech transcripts are not included. Data were collected via a standardized protocol using a custom tablet application, with preprocessing that included resampling to 16 kHz with a Butterworth anti-aliasing filter, STFT-based spectrogram generation, extraction of acoustic and phonetic/prosodic features (OpenSMILE, Parselmouth/Praat), and automatic transcription using OpenAI Whisper Large. Documentation: "https://docs.b2ai-voice.org."
Addresses the lack of large, high-quality, standardized, and demographically diverse voice datasets linked to clinical and questionnaire data across multiple institutions.
Description
Recordings are linked to participants and sessions via participant_id and session_id; tasks identified by task_name
Free speech transcripts removed; original audio waveforms omitted in v1.0
Description
Contains de-identified health-related information (demographics, clinical variables, questionnaire responses)
Description
De-identified clinical and questionnaire data; identifiable communications removed
Description
Standardized protocol with voice tasks (e.g., sustained vowel), demographics, health and targeted questionnaires
Custom tablet application; headset used when possible
Data exported from REDCap via open-source tooling
Was Directly Observed
yes
Was Reported By Subjects
yes
Was Inferred Derived
yes
Description
Custom tablet app and headset for data capture; REDCap used for data entry/export
Description
Project investigators at five North American sites; patients presenting at specialty clinics were screened and consented
Description
Data collection and sharing approved by the University of South Florida IRB; submitted for review to the University of Toronto Research Ethics Board
Description
Raw audio converted to mono and resampled to 16 kHz with a Butterworth anti-aliasing filter
Spectrograms computed via short-time FFT (25 ms window, 10 ms hop, 512-point FFT)
Acoustic features extracted with OpenSMILE
Phonetic and prosodic features computed with Parselmouth and Praat
Transcriptions generated with OpenAI Whisper Large model
Used Software
Name
URL
OpenSMILE
Parselmouth
Praat
torchaudio
OpenAI Whisper Large
b2aiprep
https://github.com/sensein/b2aiprep
Description
HIPAA Safe Harbor de-identification (removal of identifiers and fine-grained dates)
Removal of free speech transcripts
Omission of original audio waveforms from initial release
Description
Automatic Transcription Using Openai Whisper Large (note
free speech transcripts are not included in v1.0)
Description
Original audio waveforms collected; not distributed in v1.0. Future releases aim to include voice data with additional precautions.
Description
Health Data Nexus
Temerty Centre for AI Research and Education in Medicine
mixed (Parquet dense arrays and tabular TSV/JSON)
Description
Format
ID
Media Type
Name
Path
Title
Parquet dataset containing time-frequency spectrograms derived from voice recordings; includes parti...
spectrograms-parquet
application/x-parquet
spectrograms.parquet
spectrograms.parquet
Spectrograms (derived from voice waveforms)
Tab-delimited file with one row per participant containing demographics, acoustic confounders, and v...
phenotype-tsv
text/tab-separated-values
phenotype.tsv
phenotype.tsv
Phenotype table
JSON data dictionary describing columns in phenotype.tsv; includes a one-sentence description for ea...
JSON
phenotype-json
application/json
phenotype.json
phenotype.json
Phenotype data dictionary
Tab-delimited file with one row per recording containing features derived from OpenSMILE, Praat/Pars...
static-features-tsv
text/tab-separated-values
static_features.tsv
static_features.tsv
Static audio-derived features
JSON data dictionary describing features in static_features.tsv; includes a description for each fea...
JSON
static-features-json
application/json
static_features.json
static_features.json
Static features data dictionary
🚀
Uses
What (other) tasks could the dataset be used for?
Response
Research and development of AI/ML methods for analyzing voice-derived features and their associations with health conditions; exploratory and hypothesis-driven studies on voice as a biomarker.
Description
Bridge2AI Voice Registered Access License
Bridge2AI Voice Registered Access Agreement (DUA)
Access Policy
Only credentialed users who sign the DUA can access the files