healthnexus tab release-notes d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before3OT2OD032720-01S1
  • Response
    Create an ethically sourced, diverse, multi-institutional voice dataset linked to health information to enable AI research on voice as a biomarker of health.
📊

Composition

What do the instances represent?

  1. Representation
    Voice-derived data (spectrograms and acoustic/phonetic/prosodic features) and associated phenotype/clinical data.
    Instance Type
    Participants, sessions, and recordings; derived data per recording; phenotype per participant.
    Data Type
    Derived features from audio (spectrograms 513xN; static acoustic features), plus tabular phenotype data and accompanying data dictionaries.
    Counts
    12,523
    Label
    Not applicable; dataset includes derived features and task labels; free speech transcripts removed.
    Sampling Strategies
    1. Is Sample
      • True
      Is Random
      • False
      Source Data
      • Patients at specialty clinics across five North American sites.
      Is Representative
      • No (targeted disease cohorts rather than a general population sample).
      Representative Verification
      • Not applicable for targeted cohort recruitment.
      Why Not Representative
      • Participants selected based on membership in predefined disease cohorts (respiratory, voice, neurological, mood); adult cohort only in v1.0.
      Strategies
      • Targeted recruitment at specialty clinics using inclusion/exclusion criteria and standardized protocols.
    Missing Information
    • Missing
      • Original audio waveforms (omitted in v1.0).
      • Transcripts of free speech audio (removed).
      Why Missing
      • Privacy protection and de-identification for low-risk release.
  • Identification
    • Adult cohort only in v1.0.
    • Disease Cohorts
      voice disorders; neurological/neurodegenerative disorders; mood/psychiatric disorders; respiratory disorders.
    Distribution
    • 306 participants across five North American sites; typically one session per participant, with some participants having multiple sessions.
  • Description
    • Parquet
    • TSV
    • JSON
  • Description
    • 2024-11-27
🔍

Collection Process

How was the data acquired?

Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information v1.0
Bridge2AI-Voice is an ethically sourced, multi-institutional dataset linking voice-derived data with clinical and demographic information to enable research on voice as a biomarker of health. The v1.0 release includes 12,523 recordings from 306 adult participants collected across five sites in North America. Participants were selected from disease cohorts where voice and speech changes are clinically relevant (voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, and respiratory disorders). This initial release provides low-risk derived data (e.g., spectrograms and acoustic/phonetic/prosodic features) and detailed phenotype data; original audio waveforms and free speech transcripts are not included to protect privacy. Data collection followed a standardized protocol with IRB/REB oversight, and preprocessing used established audio analysis tools. Access is credentialed and governed by a registered-access license, data use agreement, and required training.
2024-11-27
  • voice
  • bridge2ai
  • audio
RoleNameORCIDAffiliation
ContributorBridge2AI-Voice Team-Health Data Nexus
  • Response
    Addresses the lack of large, high-quality, diverse, multi-institutional voice datasets with standardized protocols and linked health information needed for clinically relevant AI research.
Description
  • HIPAA Safe Harbor identifiers removed (e.g., names, contact details, precise dates, device IDs, medical identifiers, and other unique identifiers).
  • State and province removed; country of data collection retained.
  • Transcripts of free speech audio removed.
  • Original audio waveforms omitted in v1.0; only spectrograms and derived features are provided.
  • Description
    • Health-related demographic, clinical, and validated questionnaire data.
  • Description
    • Clinical information and questionnaire responses distributed under registered access with DUA.
  1. Description
    • Voice recordings collected in clinic using a standardized protocol via a custom tablet application; headset microphone used when possible.
    • Demographics, clinical questionnaires, and confounders collected via the same application; exported from REDCap.
    Was Directly Observed
    Yes (voice audio recordings; derived spectrograms).
    Was Reported By Subjects
    Yes (questionnaires and self-reported data).
    Was Inferred Derived
    Yes (acoustic/phonetic/prosodic features; automatic transcriptions).
    Was Validated Verified
    Standardized multi-site protocol with IRB/REB oversight; derived features computed using established tools.
  • Description
    • Custom tablet application for data capture with headset when possible.
    • Data export and conversion from REDCap using the open-source b2aiprep library.
  • Description
    • Project investigators at specialty clinics across five North American sites; recruitment based on inclusion/exclusion criteria prior to clinic visits.
  • Description
    • Collected during clinic visits across five North American sites; typically single-session per participant with some multi-session participants.
  • Description
    • Approved by the University of South Florida Institutional Review Board (IRB).
    • Submitted for review to the University of Toronto Research Ethics Board (REB).
  • Description
    • Raw audio converted to mono and resampled to 16 kHz with a Butterworth anti-aliasing filter.
    • Spectrograms computed via STFT using 25 ms window, 10 ms hop, 512-point FFT (yielding 513xN spectrograms).
    • Acoustic features extracted with OpenSMILE; phonetic/prosodic features computed with Parselmouth and Praat.
    • Transcriptions generated using OpenAI's Whisper Large model (free speech transcripts removed in release).
    Used Software
    NameURL
    OpenSMILEhttps://audeering.github.io/opensmile/
    Parselmouthhttps://parselmouth.readthedocs.io/
    Praathttp://www.fon.hum.uva.nl/praat/
    Torchaudiohttps://pytorch.org/audio
    OpenAI Whisper Largehttps://openai.com/research/whisper
    b2aiprephttps://github.com/sensein/b2aiprep
  • Description
    • De-identification per HIPAA Safe Harbor; removal of state/province; removal of free speech transcripts; omission of original audio waveforms in v1.0.
  • Description
    • Automatic speech transcription using OpenAI Whisper Large (with free speech transcripts excluded from release).
    • Validated questionnaires for demographic and clinical variables.
    Used Software
    Name
    OpenAI Whisper Large
    Custom tablet data collection application
  • Description
    • Raw audio recordings were collected but are not released in v1.0; only derived spectrograms and features are provided. Future releases may include voice data with additional safeguards.
  1. External Resources
    Future Guarantees
    • Not specified.
    Archival
    • Zenodo record for REDCap configuration; dataset distributions are versioned with DOIs through Health Data Nexus.
    Restrictions
    • Registered/credentialed access with DUA and required training.
Description
  • Use restricted under registered-access license and DUA; credentialed access required.
  • Completion Of Tcps 2
    CORE 2022 training required prior to access.
  • Description
    • Health Data Nexus (Temerty Centre for AI Research and Education in Medicine)
🚀

Uses

What (other) tasks could the dataset be used for?

Response
AI model development and evaluation for detecting, characterizing, or monitoring health conditions f...
Research on associations between voice markers and clinical phenotypes using standardized, multi-sit...
  • Description
    • Initial dataset release (v1.0); prior downstream uses not listed.
Description
  • License
    Bridge2AI Voice Registered Access License.
  • Data Use Agreement
    Bridge2AI Voice Registered Access Agreement.
  • Access Policy
    Only credentialed users who sign the DUA can access the files.
  • Required Training
    TCPS 2: CORE 2022.
📤

Distribution

How will the dataset be distributed?

🔄

Maintenance

How will the dataset be maintained?

1.0
Description
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema