Enable AI research into health-related acoustic markers using ethically sourced, diverse voice data linked to clinical information.
Used Software
Response
Enable future research in artificial intelligence using voice as a biomarker of health.
Grantor
Grant Name
Grant Number
NIH
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before.
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
ID
instance:recordings-and-derived-data
Name
Voice recordings (derived) and participant data
Description
Instances represent derived data from voice recordings (e.g., spectrograms and static acoustic/phonetic/prosodic features) linked to participant-level phenotype and clinical questionnaire data.
Used Software
Representation
Derived voice data linked to clinical and demographic information.
Instance Type
Participants, recording sessions, and derived data per recording.
Data Type
Derived spectrograms (513 x N), static acoustic/phonetic/prosodic features, and participant phenotype/clinical questionnaire responses.
Counts
12,523
Label
Participants selected from predefined disease cohorts; no explicit classification labels provided in v1.0.
Sampling Strategies
ID
sampling:clinical-cohorts
Name
Cohort-based sampling at specialty clinics
Description
Non-random sample of patients selected from five predetermined clinical groups at specialty clinics and institutions.
Used Software
Is Sample
yes
Is Random
no
Source Data
Specialty clinics in North America with predefined disorder cohorts
Is Representative
Not intended to be representative of the general population
Why Not Representative
Purposeful enrichment for conditions known to manifest in the voice waveform
Strategies
Deterministic cohort-based inclusion from predefined groups
Missing Information
ID
missing:raw-audio
Name
Raw audio and free-speech transcripts removed
Description
Raw audio waveforms and transcripts of free speech are not included in v1.0 to reduce risk and protect privacy.
Used Software
Missing
Raw audio waveforms
Transcripts of free speech audio
Why Missing
Privacy and de-identification; low-risk initial release with only derived data
ID
subpops:cohorts
Name
Disorder cohorts
Description
Identification
Voice disorders
Neurological and neurodegenerative disorders
Mood and psychiatric disorders
Respiratory disorders
Pediatric voice and speech disorders (not included in v1.0)
Distribution
v1.0 contains adult cohort data only
Used Software
Description
ID
Name
Used Software
Restricted-access database; files available to credentialed users who sign the DUA and complete required training.
distfmt:portal
Credentialed access via Health Data Nexus
Spectrograms stored as Parquet
distfmt:parquet
Parquet
phenotype.tsv and static_features.tsv
distfmt:tsv
TSV
phenotype.json and static_features.json data dictionaries
distfmt:json
JSON
ID
distdate:2024-11-27
Name
Initial release date
Description
2024-11-27
Used Software
🔍
Collection Process
How was the data acquired?
bridge2ai-voice-v1-0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
The human voice contains complex acoustic markers which have been linked to important health conditions including dementia, mood disorders, and cancer. When viewed as a biomarker, voice is a promising characteristic to measure as it is simple to collect, cost-effective, and has broad clinical utility. The Bridge2AI-Voice project seeks to create an ethically sourced flagship dataset to enable future research in artificial intelligence and support critical insights into the use of voice as a biomarker of health. Bridge2AI-Voice v1.0, the initial release, provides 12,523 recordings for 306 participants collected across five sites in North America. Participants were selected based on known conditions which manifest within the voice waveform including voice disorders, neurological disorders, mood disorders, and respiratory disorders. The initial release contains data considered low risk, including derivations such as spectrograms but not the original voice recordings. Detailed demographic, clinical, and validated questionnaire data are also made available.
Addresses the lack of large, diverse, multi-institutional voice datasets with standardized collection protocols and linked clinical/demographic data for AI research.
Used Software
Response
Provide a large, high-quality, multi-institutional and diverse voice dataset with standardized protocols linked to health information.
ID
relationships:participant-session
Name
Participant-session linkage
Description
Spectrogram entries include participant_id, session_id, and task_name linking recordings to participants and sessions.
Used Software
ID
splits:na
Name
Data splits (not specified)
Description
No recommended train/validation/test splits are provided in v1.0.
Used Software
ID
anomalies:none-reported
Name
No anomalies reported
Description
No specific errors, noise sources, or redundancies were reported in the release notes.
Used Software
Description
ID
Name
Used Software
external_resources: ['https://docs.b2ai-voice.org'] future_guarantees: ['Not stated'] archival: ['DOI versioning is provided']
Dataset includes clinical and demographic information linked to recordings; v1.0 includes de-identified, low-risk derived data only.
Used Software
ID
sensitive:health-data
Name
Health-related data
Description
Demographics, clinical information, and questionnaire responses related to health conditions
Used Software
ID
acquisition:direct-recording-and-questionnaires
Name
Direct recording with standardized protocol and questionnaires
Description
Raw audio recorded via a custom tablet application with headset when possible; demographic and disease-specific data collected via validated questionnaires and clinical instruments.
Multiple tasks including sustained vowel phonation; sessions per participant as needed.
Used Software
Was Directly Observed
yes (raw audio recordings; derived data released)
Was Reported By Subjects
yes (validated questionnaires)
Was Inferred Derived
yes (derived spectrograms and features from raw audio; ASR transcripts initially generated then removed prior to release)
Custom tablet application and headset; REDCap for data management
Description
Standardized protocol; custom data collection app on tablet with headset when possible; REDCap used for clinical/phenotype data; export and conversion performed with an open-source library (b2aiprep).
Used Software
ID
Name
URL
software:redcap
REDCap
software:b2aiprep
b2aiprep
https://github.com/sensein/b2aiprep
ID
collectors:project-investigators
Name
Project investigators at specialty clinics and institutions
Description
Patients screened for inclusion/exclusion; investigators obtained consent and conducted data collection sessions.
Used Software
Description
ID
Name
Used Software
Data collection and sharing approved by the University of South Florida Institutional Review Board.
ethics:usf-irb
University of South Florida IRB
Submission for review to the University of Toronto Research Ethics Board.
ethics:utoronto-reb
University of Toronto Research Ethics Board
ID
prep:audio-standardization
Name
Audio standardization and feature derivation
Description
Raw audio converted to mono, resampled to 16 kHz with a Butterworth anti-aliasing filter; short-time FFT spectrograms computed (25 ms window, 10 ms hop, 512-point FFT).
Acoustic features extracted with OpenSMILE; phonetic and prosodic features computed using Parselmouth and Praat; ASR transcriptions generated using Whisper Large (transcripts of free speech removed before release).
Used Software
ID
Name
URL
software:opensmile
OpenSMILE
software:parselmouth
Parselmouth
software:praat
Praat
software:torchaudio
Torchaudio
software:whisper
OpenAI Whisper Large
software:b2aiprep
b2aiprep
https://github.com/sensein/b2aiprep
ID
cleaning:deid-and-removals
Name
De-identification and removal of sensitive fields
Description
HIPAA Safe Harbor identifiers removed; state/province removed (country retained); transcripts of free speech removed; raw audio waveforms omitted from v1.0.
Used Software
ID
labeling:asr-internal-then-removed
Name
ASR transcription (internal) then removal
Description
Transcriptions were generated using Whisper Large during processing; transcripts of free speech audio were removed in the released dataset.
Used Software
ID
software:whisper
Name
OpenAI Whisper Large
ID
raw:audio
Name
Raw audio waveforms
Description
Raw audio recorded during sessions; not included in v1.0 release. Future releases may include voice data with additional precautions for data security.
Used Software
ID
deid:hipaa-safe-harbor
Name
HIPAA Safe Harbor de-identification
Description
Removal of HIPAA Safe Harbor identifiers
Removal of state/province (country retained)
Removal of transcripts of free speech
Omission of raw audio in v1.0 (derived data only)
mixed
Role
Name
ORCID
Affiliation
Contributor
spectrograms.parquet
subset:spectrograms-parquet
-
Contributor
phenotype.tsv
subset:phenotype-tsv
-
Contributor
phenotype.json
subset:phenotype-json
-
Contributor
static_features.tsv
subset:static-features-tsv
-
Contributor
static_features.json
subset:static-features-json
-
🚀
Uses
What (other) tasks could the dataset be used for?
ID
task:acoustic-feature-analysis
Name
Acoustic feature analysis and modeling
Description
Support AI/ML methods for extracting prognostic and diagnostic information from voice-derived spectrograms and features.
Used Software
Response
AI/ML research on voice-derived spectrograms and features linked to health data.
ID
othertasks:health-ai
Name
Additional health AI tasks
Description
Potential use in screening, monitoring, and characterization of conditions that manifest in voice and speech.
Used Software
ID
future:low-risk-derivatives
Name
Low-risk derivative-only release considerations
Description
Initial release limits data to derived spectrograms and features (no raw audio or free-speech transcripts) to reduce privacy risks and support ethical use.
Used Software
ID
terms:registered-access
Name
Registered access license and DUA
Description
License
Bridge2AI Voice Registered Access License
Data Use Agreement
Bridge2AI Voice Registered Access Agreement
Access Policy
Only credentialed users who sign the DUA can access the files