Create an ethically sourced, diverse, multi-institutional voice dataset linked to health information to enable AI research on voice as a biomarker of health.
Grantor
Grant Name
Grant Number
National Institutes of Health
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
Name
Instance description
Representation
Voice-derived data linked to demographic, clinical, and validated questionnaire information.
Instance Type
Participants with one or more recording sessions; one row per recording in features; one row per participant in phenotype.
Data Type
Derived data including:
- Spectrograms (513 x N time-frequency representations from 16 kHz monaural audio)
- Acoustic features (e.g., openSMILE)
- Phonetic/prosodic features (e.g., Praat/Parselmouth)
- Phenotype/demographics/clinical/questionnaire responses
Note: Raw audio waveforms are not included in v1.0.
Counts
12,523
Label
Task identifiers (task_name) per recording; no explicit diagnostic labels distributed in v1.0.
Sampling Strategies
Name
Sampling
Is Sample
Yes; participants selected from specialty clinics.
Is Random
No; condition-based targeted enrollment.
Source Data
Patients presenting at participating specialty clinics across five North American sites.
Is Representative
Not intended to be representative of the general population; targeted by condition cohorts.
Why Not Representative
Cohort design targets specific diseases and clinical populations.
Strategies
Targeted clinical cohort recruitment based on predefined disease categories (voice, neurological/neurodegenerative, mood/psychiatric, respiratory; pediatric planned but not included in v1.0).
Distribution
Identification
Name
Distribution by cohort not provided in the release page.
Pediatric cohort planned for future releases; not present in v1.0.
As of v1.0, only adult participants are included.
Adult cohort v1.0
Name
Formats
Description
Parquet (.parquet) for spectrograms
TSV (.tsv) for phenotype and static features
JSON (.json) data dictionaries for phenotype and features
Name
DistributionDate
Description
2024-11-27
🔍
Collection Process
How was the data acquired?
bridge2ai-voice-1.0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice v1.0 is a comprehensive collection of data derived from voice recordings with corresponding clinical information to enable research on voice as a biomarker of health. The initial release provides 12,523 recordings for 306 participants collected across five sites in North America. Participants were selected based on known conditions which manifest within the voice waveform including voice disorders, neurological disorders, mood disorders, and respiratory disorders. As of v1.0, only data from the adult cohort is available. This initial release contains data considered low risk, including derivations such as spectrograms and engineered features, but not the original voice recordings. Detailed demographic, clinical, and validated questionnaire data are also made available. Data collection used a standardized protocol with a custom tablet application, headset recording when possible, and REDCap; preprocessing included resampling to 16 kHz with a Butterworth anti-aliasing filter and derivation of spectrograms and acoustic/phonetic/prosodic features.
Addresses the need for a large, high-quality, diverse, multi-institutional voice dataset linked to other health biomarkers with standardized data collection, ethical sourcing, and detailed clinical/phenotypic context.
Name
Cohort sampling
Is Sample
Yes; clinical cohort-based sample.
Is Random
No.
Source Data
Specialty clinic patients screened against inclusion/exclusion criteria and consented.
Is Representative
Not of the general population; designed for disease-relevant cohorts.
Representative Verification
Standardized protocol across sites; not intended as population-representative sampling.
Strategies
Deterministic cohort selection based on predefined conditions.
None beyond dataset access controls for data files.
Name
Confidentiality
Description
Dataset is released in de-identified, low-risk form; original voice recordings and free-speech transcripts are not included in v1.0.
Name
ContentWarning
Warnings
None noted.
Name
Sensitive elements
Description
Clinical and health-related questionnaire responses.
Voice-related features may be considered biometric information.
Name
Deidentification
Description
HIPAA Safe Harbor identifiers removed.
State and province removed; country of data collection retained.
Free speech transcripts removed.
Raw audio waveforms omitted; only derived spectrograms and features included in v1.0.
Dataset characterized as low risk for this initial release.
Name
Acquisition
Description
Directly observed voice recordings; subject-reported validated questionnaires; derived features from raw audio.
Was Directly Observed
True
Was Reported By Subjects
True
Was Inferred Derived
True
Was Validated Verified
True
Name
CollectionMechanism
Description
Standardized multi-site protocol; custom tablet application; headset used when possible; data captured into REDCap and exported; preprocessing and merges performed with the open-source b2aiprep library.
Used Software
Name
URL
Version
Bridge2AI Voice REDCap
https://doi.org/10.5281/zenodo.14148755
v3.20.0
b2aiprep
https://github.com/sensein/b2aiprep
Name
DataCollector
Description
Project investigators at participating specialty clinics; screening against inclusion/exclusion criteria; patient consent obtained prior to data collection.
Name
EthicalReview
Description
Data collection and sharing approved by the University of South Florida Institutional Review Board (IRB).
Submitted for review to the University of Toronto Research Ethics Board (REB).
Name
Preprocessing
Description
Raw audio converted to monaural, resampled to 16 kHz with a Butterworth anti-aliasing filter.
Spectrograms computed with STFT (25 ms window, 10 ms hop, 512-point FFT).
Acoustic features extracted (e.g., openSMILE); phonetic/prosodic features computed (Praat/Parselmouth); additional features via torchaudio.
Used Software
Name
openSMILE
Praat
Parselmouth
torchaudio
b2aiprep
Name
Cleaning
Description
De-identification per HIPAA Safe Harbor; removal of state/province; removal of free speech transcripts; omission of original audio waveforms for v1.0.
Name
Labeling
Description
Automatic speech transcriptions generated using Whisper Large during processing; free speech transcripts were removed prior to release.
Used Software
Name
OpenAI Whisper Large
Name
RawData
Description
Raw audio waveforms were not released in v1.0; only derived representations (spectrograms and features) are distributed.
Name
IPRestrictions
Description
None stated beyond the registered access license and DUA requirements.
Name
Maintainer
Description
Health Data Nexus (Temerty Centre for AI Research and Education in Medicine), supported by the Temerty Foundation.
Mixed; spectrograms in Parquet, features and phenotype in TSV with JSON data dictionaries.
🚀
Uses
What (other) tasks could the dataset be used for?
Name
Task
Response
AI method development and evaluation for detecting or characterizing health conditions from voice-derived representations (e.g., classification across disease cohorts).
Name
OtherTask
Description
Benchmarking of voice-based health AI models; feature analysis and biomarker discovery; method development for fair, robust voice-based health inference.
Name
FutureUseImpact
Description
Initial release omits raw audio to reduce privacy risk; users should consider potential sensitivity of voice-derived biometric features and clinical phenotypes in downstream applications.
Name
LicenseAndUseTerms
Description
License
Bridge2AI Voice Registered Access License.
Access requires signing the Bridge2AI Voice Registered Access Agreement (DUA).
Access Is Restricted To Credentialed Users Who Complete Required Training (tcps 2
Future releases aim to include original voice recordings with additional precautions to ensure data security; pediatric cohort planned for future inclusion.
Generated on 2025-11-09 10:17:35 using Bridge2AI Data Sheets Schema