Create a large, diverse, ethically sourced, multi-institutional dataset to enable AI research on voice as a biomarker of health, linked to clinical and demographic information.
Grantor
Grant Name
Grant Number
National Institutes of Health (NIH)
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
Name
distribution-formats
Description
application/x-parquet
text/tab-separated-values
application/json
Description
Name
2024-11-27
v1.0-release
2025-01-17
v1.1-release
Name
recordings-and-phenotype
Representation
Derived representations from voice recordings (spectrograms, MFCCs, acoustic/phonetic/prosodic feature vectors) and participant-level phenotype (demographics, clinical data, validated questionnaires).
Instance Type
Participants and their recording sessions (participants, sessions, and recordings; one or more sessions per participant).
Data Type
Non-raw, derived audio representations (STFT-based spectrograms; MFCCs; engineered features via OpenSMILE, Parselmouth/Praat; torchaudio-derived features) plus tabular phenotype data and data dictionaries.
Counts
12,523
Label
Health condition cohorts, demographic attributes, and validated questionnaire responses; task metadata (e.g., task_name), participant_id, and session_id.
Sampling Strategies
Name
cohort-based-enrollment
Strategies
Deterministic cohort inclusion based on predefined disease groups at specialty clinics.
Is Sample
True
Is Random
False
Source Data
Specialty clinics across five North American sites.
Is Representative
Not statistically representative of the general population (cohort-based).
Why Not Representative
Focused on predefined disorders to enable targeted biomarker research.
Pediatric voice and speech disorders (planned; adult cohort only in v1.0/v1.1)
Distribution
Adult cohort only in v1.0 and v1.1; 306 participants across five North American sites.
🔍
Collection Process
How was the data acquired?
bridge2ai-voice
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a multi-site, ethically-sourced dataset linking derived data from human voice recordings to detailed demographic, clinical, and validated questionnaire information. The dataset enables research on voice as a biomarker of health across multiple condition cohorts (voice disorders, neurological and neurodegenerative disorders, mood and psychiatric disorders, respiratory disorders, and pediatric voice/speech disorders). The initial releases (v1.0, v1.1) contain low-risk derived data (e.g., spectrograms, MFCCs, engineered acoustic/phonetic/prosodic features) and phenotype tables; raw audio waveforms are not included in public distributions. As of v1.1, the dataset contains 12,523 recordings from 306 adult participants collected across five sites in North America under a standardized protocol, with de-identification following HIPAA Safe Harbor.
Address the lack of large, diverse, standardized, multi-institutional voice datasets with linked clinical data and clear ethical frameworks; overcome small sample sizes, limited demographic diversity reporting, and inconsistent collection protocols in prior literature.
Role
Name
ORCID
Affiliation
Contributor
Bridge2AI-Voice Project Team
-
Bridge2AI Program (NIH Common Fund initiative)
Name
participant-session-recording-linkage
Description
Participant_id and session_id link phenotype rows to recording-derived features; multiple sessions may exist per participant.
Name
recommended-splits
Description
No official train/validation/test splits are provided in these releases.
Public repository for preprocessing code (b2aiprep).
Restrictions
Not applicable to dataset access (software/resources are open; dataset itself is registered/credentialed access).
Name
confidentiality
Description
Dataset releases are designed as low risk; raw audio and free-speech transcripts are withheld; HIPAA Safe Harbor identifiers removed.
Name
content-warnings
Warnings
None stated.
Name
sensitive-health-data
Description
Health-related information (disease cohorts, questionnaires) and voice-derived features are potentially sensitive and handled under registered/credentialed access with DUA.
REDCap used for data capture and export; conversion and integration performed with open-source tooling (b2aiprep).
Name
collection-teams
Description
Project investigators and clinical teams at specialty clinics across five North American sites.
Name
timeframe
Description
Not explicitly stated; adult cohort collected prior to v1.0 (Nov 2024) and v1.1 (Jan 2025) releases.
Name
hosting-and-access-ethics
Description
Registered/credentialed access, signed DUA, and training (where required) used to ensure ethical data sharing and minimize risks.
Name
data-protection
Description
Dataset release strategy minimizes risk via HIPAA Safe Harbor de-identification and withholding of raw audio and free-speech transcripts; controlled/registered access with DUA.
Description
Name
Used Software
Monaural conversion; resampling to 16 kHz; Butterworth anti-aliasing filter.
Removal of HIPAA Safe Harbor identifiers; removal of state/province; removal of free-speech transcripts; omission of raw audio from public releases.
Name
questionnaires-and-task-labels
Description
Validated questionnaires and clinical/phenotypic fields; recording task labels (e.g., task_name) associated with each session/recording.
Name
raw-audio
Description
Original audio waveforms exist but are not publicly distributed; controlled access requests can be directed to DACO@b2ai-voice.org (per PhysioNet notice).
Name
third-party-distribution
Description
Dataset is distributed to third parties under registered/credentialed access with a signed data use agreement; some hosts require documented training (e.g., TCPS 2: CORE 2022 on Health Data Nexus).
Description
ID
Media Type
Name
Path
Title
Short-time FFT spectrograms (513 x N) derived from 16 kHz monaural audio.
spectrograms-parquet
application/x-parquet
spectrograms.parquet
spectrograms.parquet
Spectrograms
60 Mel-frequency cepstral coefficients (60 x N) derived from spectrograms (added in v1.1).
mfcc-parquet
application/x-parquet
mfcc.parquet
mfcc.parquet
MFCCs
One row per recording with features from OpenSMILE, Parselmouth/Praat, and torchaudio-derived measur...
static-features-tsv
text/tab-separated-values
static_features.tsv
static_features.tsv
Engineered acoustic/phonetic/prosodic features
Column-level metadata and descriptions for static_features.tsv.
static-features-dict
application/json
static_features.json
static_features.json
Data dictionary for engineered features
Participant-level demographics, acoustic confounders, clinical data, and validated questionnaire res...
phenotype-tsv
text/tab-separated-values
phenotype.tsv
phenotype.tsv
Phenotype table
Column-level metadata and descriptions for phenotype.tsv.
phenotype-dict
application/json
phenotype.json
phenotype.json
Data dictionary for phenotype
Name
erratum
Description
None specified.
🚀
Uses
What (other) tasks could the dataset be used for?
Name
registered-access-license
Description
Files are distributed under the Bridge2AI Voice Registered Access License and require acceptance of the Bridge2AI Voice Registered Access Agreement (DUA).
Access Is Restricted To Registered/credentialed Users Who Sign The Dua; On Health Data Nexus, Tcps 2
CORE 2022 training is required for access.
Name
target-tasks
Response
AI model development and evaluation for disease screening, risk stratification, and monitoring using voice-derived representations (classification, regression, and representation learning); methodological research on voice/speech biomarkers.
Name
usage-tracking
Description
Not specified; users are requested to cite the appropriate DOI for each used release (e.g., v1.0 on Health Data Nexus; v1.1 on PhysioNet).
Name
citations
Description
DOI (v1.0): https://doi.org/10.57764/qb6h-em84; DOI (v1.1): https://doi.org/10.13026/249v-w155
Name
potential-uses
Description
Multimodal fusion with clinical data; fair and robust model development; domain shift and cohort generalization studies.
Name
risk-mitigation
Description
Cohort-based design and site effects may impact generalizability; users should consider fairness, demographic diversity, and cohort balance in downstream tasks.
Name
inappropriate-uses
Description
Not specified in source; comply with DUA and ethical standards; avoid attempts at re-identification.
Older versions may be retained for citation purposes; some earlier-version files may no longer be downloadable after newer releases on hosting platforms (e.g., PhysioNet 2.x).