Enable AI research and critical insights into the use of voice as a biomarker of health via an ethically sourced, diverse, multi-institutional dataset linked to clinical information.
Grantor
Grant Name
Grant Number
National Institutes of Health (NIH)
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
Counts
Data Type
ID
Instance Type
Name
Representation
12523
Spectrograms (513xN), acoustic/phonetic/prosodic features; raw waveforms not included in v1.0.
instance-recordings
participants, sessions, and recordings
Derived recording instances
Derived voice data elements per recording (e.g., spectrogram matrices, static features) linked to se...
306
Tabular phenotype data with one row per participant; associated data dictionary.
instance-participants
participants
Participant instances
Adult participants with demographic, clinical, and validated questionnaire responses.
ID
subp-1
Name
Disease cohorts
Identification
Voice disorders
Neurological and neurodegenerative disorders
Mood and psychiatric disorders
Respiratory disorders
Pediatric voice and speech disorders (adult cohort only in v1.0)
Distribution
Adult cohort only; 306 participants across five sites in North America.
ID
distfmt-1
Name
Distribution formats and files
Description
spectrograms.parquet (Parquet; spectrogram matrices and metadata)
phenotype.tsv (tab-delimited; one row per participant)
phenotype.json (data dictionary for phenotype)
static_features.tsv (tab-delimited; one row per recording with features)
static_features.json (data dictionary for features)
ID
distdate-1
Name
Initial release date
Description
2024-11-27 (v1.0 first release)
ID
access-1
Name
Access modality
Description
Restricted (Credentialed Access) via Health Data Nexus; DUA and training required; no public download URLs provided for files.
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically sourced dataset enabling research on the human voice as a biomarker of health. The v1.0 release provides 12,523 recordings from 306 adult participants collected across five North American sites, with data derived from voice recordings (e.g., spectrograms, acoustic/phonetic/prosodic features) and linked clinical, demographic, and validated questionnaire information. Raw audio waveforms and free speech transcripts are not included in v1.0; only low-risk derivations are provided.
Addresses the lack of large, high-quality, diverse, and standardized multi-institutional voice datasets linked to other health biomarkers for robust AI research.
Role
Name
ORCID
Affiliation
Principal Investigator
Alistair Johnson
person-alistair-johnson
-
Principal Investigator
Jean-Christophe Bélisle-Pipon
person-jean-christophe-belisle-pipon
-
Principal Investigator
David Dorr
person-david-dorr
-
Principal Investigator
Satrajit Ghosh
person-satrajit-ghosh
-
Principal Investigator
Philip Payne
person-philip-payne
-
Principal Investigator
Maria Powell
person-maria-powell
-
Principal Investigator
Anaïs Rameau
person-anais-rameau
-
Principal Investigator
Vardit Ravitsky
person-vardit-ravitsky
-
Principal Investigator
Alexandros Sigaras
person-alexandros-sigaras
-
Principal Investigator
Olivier Elemento
person-olivier-elemento
-
Principal Investigator
Yael Bensoussan
person-yael-bensoussan
-
Bridge2AI-Voice Team
ID
sampling-1
Name
Clinical cohort sampling
Is Sample
True
Is Random
False
Source Data
Patients at specialty clinics across five North American sites
Is Representative
No (targeted cohorts by condition)
Why Not Representative
Participants selected based on five predetermined disease groups (voice, neurological/neurodegenerative, mood/psychiatric, respiratory, pediatric)
Strategies
Targeted clinical cohort sampling based on known voice-related conditions
ID
rel-1
Name
Instance relationships
Description
Recordings are nested within sessions and participants; features and spectrograms link via participant_id, session_id, and task_name.
ID
confidential-1
Name
Clinical and questionnaire data
Description
Contains clinical, demographic, and questionnaire information; released in de-identified, low-risk form.
ID
sensitive-1
Name
Sensitive health-related data
Description
Health, demographic, and questionnaire data linked to voice recordings (de-identified).
ID
deid-1
Name
De-identification
Description
HIPAA Safe Harbor identifiers removed (e.g., names, detailed geography, dates below year, contact and ID numbers, biometric identifiers).
State and province removed; country of data collection retained.
Transcripts of free speech audio removed prior to release.
Audio waveforms omitted from v1.0; only derived spectrograms and features released.
ID
acq-1
Name
Data acquisition
Description
Standardized protocol including demographics, health and targeted questionnaires, disease-specific information, and voice tasks (e.g., sustained vowel).
Data captured via custom tablet application with headset when possible; consent obtained prior to collection.
Was Directly Observed
Yes (voice recordings, tasks)
Was Reported By Subjects
Yes (validated questionnaires)
Was Inferred Derived
Yes (features, spectrograms, ASR transcriptions)
Was Validated Verified
Standardized protocol and validated questionnaires; processing pipeline described and open-sourced.
ID
mech-1
Name
Collection mechanisms
Description
Custom tablet app; headset-based recording when possible; data exported from REDCap and converted using an open-source library (b2aiprep).
ID
collectors-1
Name
Data collectors
Description
Project investigators at specialty clinics screened patients and obtained consent; most participants completed a single session, some multiple sessions.
ID
irb-1
Name
Ethical review
Description
Data collection and sharing approved by University of South Florida Institutional Review Board; submitted for review to University of Toronto Research Ethics Board.
Description
ID
Name
Used Software
Raw audio converted to mono and resampled to 16 kHz with a Butterworth anti-aliasing filter; STFT spectrograms computed with 25 ms window, 10 ms hop, 512-point FFT.
prep-1
Audio standardization and spectrogram extraction
{'id': 'sw-torchaudio', 'name': 'Torchaudio'}
Temporal and acoustic characteristics extracted using OpenSMILE.
prep-2
Acoustic feature extraction
{'id': 'sw-opensmile', 'name': 'OpenSMILE'}
Fundamental frequency, formants, and voice quality computed using Parselmouth and Praat.
Source data exported from REDCap and merged into phenotype and feature files; accompanying JSON data dictionaries provide variable descriptions; processing code available in b2aiprep.
ID
label-1
Name
Automatic transcription
Description
ASR transcriptions generated using OpenAI Whisper Large; free speech transcripts removed from released data.
Used Software
ID
sw-whisper
Name
OpenAI Whisper Large
ID
raw-1
Name
Raw audio waveforms
Description
Raw audio collected but omitted from v1.0 release; future releases aim to include voice data with additional security precautions.
REDCap resource record (Zenodo citation provided in references)
b2aiprep open-source processing library
Future Guarantees
Versioned DOIs available for releases.
Archival
DOIs provided (versioned and latest).
Restrictions
Access via Health Data Nexus under registered access with DUA and required training.
Description
Dialect
Format
ID
Media Type
Name
Path
Title
Parquet dataset with participant_id, session_id, task_name, and 513xN spectrogram matrices derived f...
subset-spectrograms-parquet
application/x-parquet
spectrograms.parquet
spectrograms.parquet
Spectrograms derived from raw audio
Tab-delimited table with demographics, acoustic confounders, and validated questionnaire responses (...
delimiter: header: true
subset-phenotype-tsv
text/tab-separated-values
phenotype.tsv
phenotype.tsv
Participant phenotype data
JSON data dictionary providing descriptions of columns in phenotype.tsv.
JSON
subset-phenotype-json
application/json
phenotype.json
phenotype.json
Phenotype data dictionary
Tab-delimited table with one row per recording containing features derived from openSMILE, Praat, Pa...
delimiter: header: true
subset-static-features-tsv
text/tab-separated-values
static_features.tsv
static_features.tsv
Static acoustic features per recording
JSON data dictionary providing feature descriptions for static_features.tsv.
JSON
subset-static-features-json
application/json
static_features.json
static_features.json
Static features data dictionary
🚀
Uses
What (other) tasks could the dataset be used for?
ID
task-1
Name
Intended tasks
Response
Research and development of AI methods for disease-related voice/speech changes, feature discovery, and health prediction using voice-derived representations linked to clinical data.
ID
license-1
Name
Access, license, and use terms
Description
License
Bridge2AI Voice Registered Access License.
Access Policy
Only credentialed users who sign the Data Use Agreement (DUA) can access the files.
Data Use Agreement
Bridge2AI Voice Registered Access Agreement.
Required Training
TCPS 2: CORE 2022.
ID
future-impact-1
Name
Potential impacts on future use
Description
Absence of raw audio in v1.0 may limit certain signal processing and modeling tasks; inclusion planned in future releases with additional safeguards.