healthnexus tab acknowledgements d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before3OT2OD032720-01S1
  • Response
    Create an ethically sourced, diverse, multi-institutional voice dataset linked to health information to enable AI research on voice as a biomarker of health.
📊

Composition

What do the instances represent?

CountsData TypeInstance TypeRepresentationSampling Strategies
12523Derived spectrograms (513 x N), static acoustic/phonetic features (OpenSMILE, Praat, Parselmouth, to...Audio recordings and sessions (time-frequency representations and derived features; raw waveforms om...Voice-derived recordings{'strategies': ['Participants prospectively selected into 5 predetermined clinical groups (Respiratory, Voice, Neurological, Mood/Psychiatric, Pediatric); v1.0 includes adult cohort only.'], 'source_data': ['Five clinical sites in North America'], 'is_representative': ['No; targeted clinical cohorts'], 'why_not_representative': ['Cohorts intentionally sampled for disorders with known voice manifestations; not a population-representative sample.']}
306Demographics, validated health questionnaires, acoustic confounders, and disease-specific clinical i...Individuals (adult cohort in v1.0)Participants
  • Identification
    • Adult cohort only in v1.0
    • Disorder Cohorts
      Voice, Neurological/Neurodegenerative, Mood/Psychiatric, Respiratory, Pediatric (pediatric not included in v1.0)
    Distribution
    • 306 adult participants across five North American sites
    • Participants selected based on membership in predefined clinical cohorts
  • Description
    • spectrograms.parquet (Parquet; time-frequency representations per recording)
    • static_features.tsv (tab-delimited; one row per recording)
    • static_features.json (data dictionary for features)
    • phenotype.tsv (tab-delimited; one row per participant)
    • phenotype.json (data dictionary for phenotype)
  • Description
    • First Public Release (v1.0)
      2024-11-27
🔍

Collection Process

How was the data acquired?

Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically sourced dataset to enable research on the human voice as a biomarker of health. Version 1.0 provides 12,523 recordings for 306 adult participants collected across five sites in North America, selected based on conditions known to manifest in the voice waveform (voice disorders, neurological/neurodegenerative disorders, mood and psychiatric disorders, and respiratory disorders). This initial release contains low-risk derived data (e.g., spectrograms and acoustic/phonetic features) and detailed demographic, clinical, and validated questionnaire data. Original audio waveforms are not included in v1.0. Standardized collection protocols, de-identification under HIPAA Safe Harbor, and data access via registered, credentialed workflows are used to protect participants and enable responsible research.
2024-11-27
  • voice
  • bridge2ai
  • audio
RoleNameORCIDAffiliation
Principal InvestigatorAlistair Johnson--
Principal InvestigatorJean-Christophe Bélisle-Pipon--
Principal InvestigatorDavid Dorr--
Principal InvestigatorSatrajit Ghosh--
Principal InvestigatorPhilip Payne--
Principal InvestigatorMaria Powell--
Principal InvestigatorAnaïs Rameau--
Principal InvestigatorVardit Ravitsky--
Principal InvestigatorAlexandros Sigaras--
Principal InvestigatorOlivier Elemento--
Principal InvestigatorYael Bensoussan--
  • Response
    Address the lack of large, standardized, demographically diverse, multi-institutional voice datasets linked to health biomarkers to improve external validity and clinical utility in voice AI research.
  1. Strategies
    • Targeted, multi-site clinical recruitment into 5 disorder cohort categories; adult cohort only in v1.0.
    Source Data
    • Specialty clinics across five North American sites
    Is Representative
    • No; targeted clinical cohorts
    Why Not Representative
    • Designed to cover diverse voice-related conditions rather than represent the general population.
  1. Description
    • Data directly observed via standardized voice recording tasks (e.g., sustained vowel phonation) plus participant-reported questionnaires and clinical data; derived features computed from raw audio.
    Was Directly Observed
    yes
    Was Reported By Subjects
    yes
    Was Inferred Derived
    yes
    Was Validated Verified
    yes
  • Description
    • Standardized protocol using a custom tablet application; headset used for data collection when possible; data exported from REDCap using an open-source library; multiple sessions for some participants as needed.
  • Description
    • Project investigators at specialty clinics across five North American sites
  • Description
    • Data collection and sharing approved by the University of South Florida Institutional Review Board; submitted for review to the University of Toronto Research Ethics Board.
  • Description
    • Raw audio converted to monaural and resampled to 16 kHz with a Butterworth anti-aliasing filter.
    • Spectrograms computed via short-time FFT (25 ms window, 10 ms hop, 512-point FFT).
    • Acoustic features extracted with OpenSMILE; phonetic/prosodic features computed with Parselmouth and Praat; additional features via torchaudio.
    Used Software
    NameURL
    openSMILEhttps://audeering.github.io/opensmile/
    Praathttps://www.fon.hum.uva.nl/praat/
    Parselmouthhttps://parselmouth.readthedocs.io/
    torchaudiohttps://pytorch.org/audio
    b2aiprephttps://github.com/sensein/b2aiprep
  • Description
    • HIPAA Safe Harbor identifiers removed.
    • State and province removed; country of data collection retained.
    • Transcripts of free speech audio removed to reduce re-identification risk.
  • Description
    • Machine-generated transcriptions produced using OpenAI Whisper Large model for certain tasks.
    Used Software
  • Description
    • Original audio waveforms were collected but are omitted from v1.0 release; only spectrograms and other derived features are provided.
  • Description
    • Clinical information and validated questionnaire responses collected during visits.
  • Description
    • Health-related data linked to voice-derived features.
    • Voice as a potential biometric/biobehavioral marker.
Description
  • HIPAA Safe Harbor identifiers removed.
  • State and province removed; country retained.
  • Free speech transcripts removed.
  • Original audio waveforms omitted from v1.0.
Description
Dataset is available to external, credentialed users under registered access terms (DUA and required training).
  • Description
    • Health Data Nexus (Temerty Centre for AI Research and Education in Medicine)
    • Supported by the Temerty Foundation
🚀

Uses

What (other) tasks could the dataset be used for?

  • Response
    Support AI methods development and clinical research using derived voice representations (e.g., spectrograms and acoustic features) linked with demographic, clinical, and questionnaire data across targeted disorder cohorts.
Description
  • License
    Bridge2AI Voice Registered Access License.
  • Access Policy
    Only credentialed users who sign the DUA can access files.
  • Data Use Agreement
    Bridge2AI Voice Registered Access Agreement.
  • Required Training
    TCPS 2: CORE 2022.
  • Access provided via Health Data Nexus credentialed workflow.
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
Description
🔄

Maintenance

How will the dataset be maintained?

1.0
Description
  • Future releases aim to include original voice waveforms with additional security safeguards; updates and documentation via https://docs.b2ai-voice.org.
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema