Create an ethically sourced, diverse voice dataset linked to health information to enable AI research and evaluate voice as a biomarker of health across multiple clinical conditions.
Grantor
Grant Name
Grant Number
National Institutes of Health
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
ID
instances-voice-derived
Name
Instance
Representation
Voice recordings and corresponding derived data per session/participant
Instance Type
Participants, recording sessions, and task-specific recordings (e.g., sustained phonation); adult cohort in v1.0
Data Type
Derived data only in v1.0: spectrograms (513 x N), acoustic, phonetic/prosodic features, and automatic transcriptions; demographic, clinical, and validated questionnaire responses
Counts
12,523
Label
Clinical cohort membership by disease category (voice, neurological, mood/psychiatric, respiratory); detailed labels vary by disease-specific questionnaires and phenotype fields
Sampling Strategies
ID
sampling-1
Name
SamplingStrategy
Is Sample
Yes; participants recruited from specialty clinics at five North American sites
Is Random
No; condition-based recruitment per predefined cohorts
Source Data
Patients presenting at participating specialty clinics
Is Representative
Not claimed; targeted clinical cohorts
Why Not Representative
Targeted enrollment to capture voice-manifesting conditions
Strategies
Deterministic cohort-based inclusion with screening per inclusion/exclusion criteria
Missing Information
ID
missing-1
Name
MissingInfo
Missing
Some participants may have multiple sessions; completeness may vary by session
Why Missing
Operational needs; subset required multiple visits to complete collection
306 participants; 12,523 recordings; across five North American sites
ID
distfmt-1
Name
DistributionFormat
Description
Credentialed access via Health Data Nexus
Distributed files in Parquet (spectrograms), TSV (phenotype and features), and JSON (data dictionaries)
ID
distdate-1
Name
DistributionDate
Description
2024-11-27 (v1.0 release)
🔍
Collection Process
How was the data acquired?
bridge2ai-voice-v1-0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information (v1.0)
Bridge2AI-Voice is a comprehensive collection of data derived from voice recordings with corresponding clinical information to enable AI research into voice as a biomarker of health. Version 1.0 provides 12,523 recordings from 306 adult participants collected across five North American sites, selected based on conditions that manifest in voice (voice disorders, neurological disorders, mood disorders, respiratory disorders, pediatric cohort planned for future releases). The initial release contains low-risk derived data (e.g., spectrograms and engineered features) and detailed demographic, clinical, and validated questionnaire data; original audio waveforms are omitted in v1.0.
Raw voice recordings collected during standardized clinical sessions; v1.0 distributes only derived data (spectrograms, engineered features, and data dictionaries) with original audio omitted.
ID
gap-1
Name
AddressingGap
Response
Provide a large, high-quality, multi-institutional, demographically diverse voice dataset with linked clinical information and standardized collection protocols, addressing prior limitations of small, non-diverse datasets and heterogeneous protocols.
ID
rel-1
Name
Relationships
Description
Each record links participant_id to session_id and task_name; one row per recording for features; one row per participant for phenotype
REDCap-based data capture (Bridge2AI Voice REDCap v3.20.0; Zenodo DOI 10.5281/zenodo.14148755); raw audio waveforms not distributed in v1.0
ID
maint-1
Name
Maintainer
Description
Hosted and made discoverable via Health Data Nexus; project documentation at https://docs.b2ai-voice.org
partially (mixture of tabular TSV/JSON and array-based Parquet)
Description
ID
Is Tabular
Keywords
Media Type
Name
Path
Title
Parquet dataset containing time-frequency representations (spectrograms) for each recording with par...
spectrograms-parquet
False
spectrogram, parquet
application/x-parquet
spectrograms.parquet
spectrograms.parquet
Derived Spectrograms
Tab-delimited file with one row per participant including demographics, acoustic confounders, and re...
phenotype-tsv
True
phenotype, demographics, questionnaires
text/tab-separated-values
phenotype.tsv
phenotype.tsv
Phenotype Table
JSON data dictionary describing columns in phenotype.tsv; each key maps to column metadata with a on...
phenotype-json
True
data dictionary, phenotype
application/json
phenotype.json
phenotype.json
Phenotype Data Dictionary
Tab-delimited file with one row per recording containing features derived from openSMILE, Praat, Par...
static-features-tsv
True
features, acoustics, phonetics
text/tab-separated-values
static_features.tsv
static_features.tsv
Engineered Acoustic/Phonetic Features
JSON data dictionary describing feature columns in static_features.tsv; each key maps to column meta...
static-features-json
True
data dictionary, features
application/json
static_features.json
static_features.json
Features Data Dictionary
ID
subset-adult-v1-0
Name
Adult cohort (v1.0)
Title
Adult Cohort Subset
Description
Initial public release contains adult participants only; pediatric cohort planned for future releases.
Is Data Split
False
Is Subpopulation
True
🚀
Uses
What (other) tasks could the dataset be used for?
ID
task-1
Name
Intended task
Response
Development and evaluation of AI methods for health-related voice biomarker discovery and analysis using derived acoustic, phonetic, prosodic, and transcriptional features.
ID
terms-1
Name
LicenseAndUseTerms
Description
Bridge2AI Voice Registered Access License
Bridge2AI Voice Registered Access Agreement (DUA)
Access Policy
Only credentialed users who sign the DUA can access files