Research enablement for voice as a biomarker of health
Description
Enable future research in artificial intelligence using ethically sourced, clinically linked voice-derived data to investigate acoustic markers of health conditions.
Response
Create a flagship dataset to support AI research on the human voice as a biomarker across multiple health domains.
Grantor
Grant Name
Grant Number
National Institutes of Health (NIH)
Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before
3OT2OD032720-01S1
📊
Composition
What do the instances represent?
Counts
Data Type
ID
Instance Type
Missing Information
Name
Representation
Sampling Strategies
12523
Derived spectrograms (513 x N), engineered acoustic/phonetic/prosodic features; transcription metada...
instance-recordings
Participants, sessions, and recording-derived instances
{'id': 'sampling-1', 'strategies': ['Condition-focused cohort inclusion across five North American sites; non-random selection'], 'is_sample': ['yes'], 'is_random': ['no'], 'source_data': ['Specialty clinics and institutions across five sites in North America'], 'is_representative': ['no'], 'why_not_representative': ['Condition-focused recruitment rather than population sampling']}
Individual adult participants with linked demographics, clinical data, and validated questionnaire r...
ID
subpop-1
Name
Adult cohort (v1.0)
Identification
Adult participants only in v1.0; pediatric cohort planned for future releases
Distribution
Participants selected across five condition cohorts (voice, neurological, mood/psychiatric, respiratory; pediatric planned)
Description
ID
Name
spectrograms.parquet — dense, derived spectrogram tensors (513 x N) per recording with participant_id, session_id, task_name metadata
distfmt-1
Parquet spectrograms
phenotype.tsv — participant-level demographics, acoustic confounders, validated questionnaires (tab-delimited), phenotype.json — data dictionary for phenotype fields
distfmt-2
Phenotype data
static_features.tsv — one row per recording with features, static_features.json — data dictionary for features
distfmt-3
Engineered acoustic/phonetic/prosodic features
ID
distdate-1
Name
Initial public release (credentialed access)
Description
v1.0 released 2024-11-27
ID
access-1
Name
Credentialed access
Description
Access requires credentialing, DUA signature, and completion of TCPS 2: CORE 2022 training
🔍
Collection Process
How was the data acquired?
bridge2ai-voice-v1.0
Bridge2AI-Voice v1.0
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
The Bridge2AI-Voice project provides an ethically sourced, diverse dataset of data derived from human voice recordings linked to clinical and demographic information to enable AI research on voice as a biomarker of health. Version 1.0 includes 12,523 recordings for 306 adult participants across five North American sites. This initial release contains low-risk derived data (e.g., spectrograms and engineered features) and detailed demographic/clinical/questionnaire information; original audio waveforms and free speech transcripts are not included. Data were collected under a standardized multi-institutional protocol with de-identification following HIPAA Safe Harbor.
Addresses the lack of large, high-quality, diverse, multi-institutional voice datasets linked to health biomarkers with standardized protocols and ethical oversight.
Response
Create a diverse, ethically sourced, clinically linked voice dataset with standardized collection and documentation.
Role
Name
ORCID
Affiliation
Principal Investigator
Alistair Johnson
person-alistair-johnson
-
Principal Investigator
Jean-Christophe Bélisle-Pipon
person-jean-christophe-belisle-pipon
-
Principal Investigator
David Dorr
person-david-dorr
-
Principal Investigator
Satrajit Ghosh
person-satrajit-ghosh
-
Principal Investigator
Philip Payne
person-philip-payne
-
Principal Investigator
Maria Powell
person-maria-powell
-
Principal Investigator
Anaïs Rameau
person-anais-rameau
-
Principal Investigator
Vardit Ravitsky
person-vardit-ravitsky
-
Principal Investigator
Alexandros Sigaras
person-alexandros-sigaras
-
Principal Investigator
Olivier Elemento
person-olivier-elemento
-
Principal Investigator
Yael Bensoussan
person-yael-bensoussan
-
ID
sampling-overall
Name
Cohort-based clinical recruitment
Is Sample
yes
Is Random
no
Source Data
Specialty clinics and institutions at five North American sites
Is Representative
no
Representative Verification
Why Not Representative
Purposeful inclusion of participants with conditions associated with voice changes
HIPAA Safe Harbor de-identification and restricted content
Description
HIPAA Safe Harbor identifiers removed (e.g., names, detailed dates, contact numbers, emails, IPs, SSNs, MRNs, plan IDs, device IDs, license/account numbers, vehicle IDs, URLs, full-face photos/biometrics, and other unique identifiers)
State/province removed; country of data collection retained
Transcripts of free speech audio removed
Original audio waveforms omitted from v1.0; only spectrograms and other derived features are released
ID
sensitive-1
Name
Health-related data
Description
Dataset includes clinical conditions and questionnaire responses linked to participants; distributed in de-identified form
ID
confidential-1
Name
Restricted content and clinical linkages
Description
Clinical and demographic linkages present; access is restricted via registered access with DUA and required training
ID
acquisition-1
Name
Standardized clinical protocol via custom app
Description
Data collected with standardized protocol including demographics, validated questionnaires, condition-specific items, and voice tasks (e.g., sustained vowel phonation)
Recording sessions conducted via custom tablet application; headset used when possible; some participants had multiple sessions
Was Directly Observed
yes (voice tasks, recordings)
Was Reported By Subjects
yes (questionnaires)
Was Inferred Derived
yes (spectrograms, engineered features, ASR transcriptions)
Was Validated Verified
Standardized multi-site protocol with IRB/REB oversight
ID
collectmech-1
Name
Custom data collection application and headset
Description
Custom tablet application; headset microphone when possible; protocol detailed in referenced documentation and publications
ID
datacollect-1
Name
Project investigators at specialty clinics and institutions
Description
Participants screened for inclusion/exclusion; consent obtained prior to standardized data collection
ID
irb-1
Name
Institutional Review and Ethics
Description
Data collection and sharing approved by the University of South Florida Institutional Review Board; submitted to the University of Toronto Research Ethics Board
ID
preprocess-1
Name
Audio standardization and feature extraction
Description
Raw audio converted to mono, resampled to 16 kHz with Butterworth anti-aliasing filter
Spectrograms via STFT (25 ms window, 10 ms hop, 512-point FFT)
Acoustic features via OpenSMILE
Phonetic/prosodic features via Parselmouth and Praat (e.g., f0, formants, voice quality)
Transcriptions generated using OpenAI Whisper Large
Used Software
ID
Name
URL
sw-opensmile
openSMILE
https://audeering.github.io/opensmile/
sw-praat
Praat
https://www.fon.hum.uva.nl/praat/
sw-parselmouth
Parselmouth (Python interface to Praat)
https://parselmouth.readthedocs.io/
sw-torchaudio
torchaudio
https://pytorch.org/audio
sw-whisper
OpenAI Whisper (Large)
https://github.com/openai/whisper
sw-b2aiprep
b2aiprep (data preprocessing library)
https://github.com/sensein/b2aiprep
sw-librosa
librosa
https://librosa.org
ID
raw-1
Name
Original audio waveforms (not released in v1.0)
Description
Raw audio was collected and preprocessed but is not included in v1.0; future releases aim to include voice data with additional security precautions
Archival
External Resources
ID
Name
Restrictions
DOI assigned for dataset releases
https://docs.b2ai-voice.org
ext-docs
Documentation website
https://doi.org/10.5281/zenodo.14148755
ext-redcap-zenodo
Bridge2AI Voice REDCap (v3.20.0) metadata/tools
Independent resource; referenced for tooling/context
ID
regulatory-1
Name
Access policy and required training
Description
Only Credentialed Users Who Sign The Dua And Complete Tcps 2
CORE 2022 may access files
ID
maint-1
Name
Health Data Nexus hosting
Description
Hosted via Health Data Nexus (Temerty Centre for AI Research and Education in Medicine)
mixed
🚀
Uses
What (other) tasks could the dataset be used for?
ID
task-1
Name
Voice biomarker AI research
Description
AI/ML analysis of derived voice representations linked to clinical data to study associations with health conditions.
Response
Development and evaluation of AI methods using derived voice representations and linked phenotypes.
ID
license-terms-1
Name
Registered access and DUA
Description
Access restricted to credentialed users
Data Use Agreement required (Bridge2AI Voice Registered Access Agreement)
Exclusion of original audio in v1.0 may limit tasks requiring waveform-level processing or re-annotation; derived features and spectrograms support many AI analyses
Future releases aim to include original voice data (audio waveforms) with additional security precautions; pediatric cohort planned for future versions
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema