Create an ethically sourced flagship dataset to enable AI research and support insights into the use of voice as a biomarker of health across multiple clinical domains.
Grantor
Grant Name
Grant Number
National Institutes of Health (NIH)
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
Name
Instance description
Representation
Voice recordings (and derived artifacts) per participant session, with linked clinical and questionnaire data at the participant level.
Instance Type
Participants, recording sessions, and per-session derived voice data
Data Type
Derived audio representations (spectrograms), acoustic/phonetic/prosodic features, and tabular phenotype data; original audio waveforms are not included in v1.0.
Counts
12,523
Sampling Strategies
Name
Sampling approach
Is Sample
Yes, selected from patients presenting at specialty clinics
Is Random
False
Source Data
Specialty clinics across five North American sites
Is Representative
Not intended to be representative of the general population
Why Not Representative
Cohort intentionally focused on predefined clinical groups with known voice manifestations
Strategies
Consecutive/screened enrollment within predefined disease cohorts at participating sites
Name
Clinical cohorts
Identification
Voice disorders
Neurological/neurodegenerative disorders
Mood/psychiatric disorders
Respiratory disorders
Pediatric (planned for future releases; not in v1.0)
Distribution
v1.0 includes adult participants across the listed cohorts; detailed cohort counts not provided here.
Name
File formats (v1.0)
Description
spectrograms.parquet (Parquet; derived spectrograms with participant_id, session_id, task_name)
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically-sourced dataset linking derived voice data to clinical and phenotypic information to enable research into voice as a biomarker of health. Version 1.0 contains 12,523 recordings from 306 adult participants collected across five North American sites, focusing on conditions known to manifest in voice signals (voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, and respiratory disorders). To minimize re-identification risk, this initial release includes derived artifacts (e.g., spectrograms and acoustic/phonetic/prosodic features) and tabular phenotype data, but not original audio waveforms or free-speech transcripts. Data collection followed a standardized multi-site protocol with informed consent and IRB oversight.
The pressing need for a large, high-quality, multi-institutional, diverse voice dataset linked to other health biomarkers to fuel reproducible voice-AI research with clinical relevance.
ID
bridge2ai-voice-v1.0-adult
Name
Adult cohort (v1.0)
Description
Version 1.0 includes only the adult cohort across five North American sites.
Is Subpopulation
adult
Name
Participant-to-session linkage
Description
Each participant may have one or multiple recording sessions; sessions are linked to participants and tasks.
Name
Data splits
Description
No recommended train/validation/test splits are provided in v1.0.
Contains de-identified clinical and questionnaire responses that could be considered confidential; access controlled.
Name
Potentially sensitive health-related context
Warnings
Health condition information may be sensitive for some users.
Name
Health-related data
Description
Health condition categories and clinical/questionnaire responses linked to participants.
Name
De-identification
Description
HIPAA Safe Harbor identifiers removed.
State and province removed; country of data collection retained.
Free-speech transcripts removed.
Original audio waveforms omitted from v1.0; only derived artifacts released.
Name
Instance acquisition
Description
Data directly collected from patients at specialty clinics using a standardized protocol.
Was Directly Observed
Yes (voice tasks and recordings)
Was Reported By Subjects
Yes (validated and targeted questionnaires)
Was Inferred Derived
Yes (acoustic/phonetic/prosodic features, spectrograms, ASR transcripts; transcripts of free speech removed from release)
Was Validated Verified
Standardized multi-site protocol applied; IRB approval obtained; specific data validation steps beyond protocol not detailed.
Name
Collection protocol and tooling
Description
Standardized multi-site protocol; custom tablet application used; headset used for data collection when possible; REDCap used for data entry/export.
Name
Data collection personnel
Description
Project investigators at participating specialty clinics across five North American sites.
Name
Collection timeframe
Description
Multi-site data collection prior to the v1.0 publication; specific start/end dates not provided.
Name
Ethics oversight
Description
Data collection and sharing approved by the University of South Florida Institutional Review Board; submitted for review to the University of Toronto Research Ethics Board.
Name
Audio preprocessing and feature extraction
Description
Raw Audio Converted To Mono, Resampled To 16 Khz With Butterworth Anti Aliasing Filter; Derived Artifacts Generated
Spectrograms via STFT (25 ms window, 10 ms hop, 512-point FFT)
Acoustic features via OpenSMILE
Phonetic/prosodic features via Parselmouth and Praat
Transcriptions via OpenAI Whisper Large (free-speech transcripts removed in release)
Used Software
Name
Version
OpenSMILE
Parselmouth
Praat
TorchAudio
2.1
OpenAI Whisper Large
Name
De-identification and release scoping
Description
HIPAA Safe Harbor removal of identifiers; removal of state/province; exclusion of free-speech transcripts; omission of original audio waveforms from v1.0.
Name
Automated transcription
Description
ASR transcriptions generated using OpenAI's Whisper Large model; free-speech transcripts were not released.
Name
Raw audio recordings
Description
Original audio waveforms were collected but are omitted from v1.0; planned for future releases subject to additional safeguards.
Name
IP and third-party restrictions
Description
No specific third-party IP restrictions stated for released files; access governed by registered access license and DUA.
Name
Export control and regulatory restrictions
Description
No export control restrictions stated.
bibo:draft
Derived from raw audio recordings collected under a standardized clinical protocol; v1.0 releases derived artifacts (spectrograms and features) and de-identified phenotype tables.
🚀
Uses
What (other) tasks could the dataset be used for?
Name
Primary task
Response
Development and evaluation of AI methods on derived voice representations (e.g., spectrograms and acoustic/phonetic/prosodic features) linked to clinical and phenotypic data for health-related research.
Name
Considerations for future use
Description
v1.0 contains only derived, low-risk data without raw audio, which may limit tasks requiring waveform-level analysis; future releases may alter risk profile when audio is included.
Name
Access and use terms
Description
Access is restricted to credentialed users who sign the Bridge2AI Voice Registered Access Agreement (DUA).
Required Training
TCPS 2: CORE 2022.
Files are distributed under the Bridge2AI Voice Registered Access License.