healthnexus tab acknowledgements d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • ID
    purpose-voice-biomarker
    Name
    Purpose
    Response
    Create an ethically sourced flagship dataset to enable AI research on voice as a biomarker of health and support insights into links between acoustic markers and health conditions.
GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before3OT2OD032720-01S1
📊

Composition

What do the instances represent?

  1. ID
    instances-v1-0
    Name
    Instance definition
    Representation
    Voice recordings (not released in v1.0) and derived data (spectrograms, acoustic/phonetic/prosodic features), with participant-level phenotype data (demographics, clinical, validated questionnaires).
    Instance Type
    Participants and recording sessions/recordings
    Data Type
    Derived spectrograms and features (from standardized audio), plus participant phenotype data; no raw audio in v1.0.
    Counts
    12,523
    Label
    Not specified; dataset includes clinical and questionnaire variables that can be used as targets.
    Sampling Strategies
    1. ID
      sampling-v1-0
      Name
      Sampling strategy
      Is Sample
      • yes
      Is Random
      • no
      Source Data
      • Specialty clinics at five North American sites
      Is Representative
      • Not explicitly validated as representative of the general population
      Why Not Representative
      • Clinic-based, condition-targeted sampling across predefined cohorts
  1. ID
    subpops-cohorts
    Name
    Cohorts
    Identification
    • Predefined disease cohorts: voice disorders, neurological disorders, mood disorders, respiratory disorders, and pediatric (pediatric not included in v1.0).
    Distribution
    • v1.0 includes adult cohort only; 306 participants; 12,523 recordings.
  • Parquet
  • TSV
  • JSON
  • 2024-11-27
🔍

Collection Process

How was the data acquired?

Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically sourced dataset derived from voice recordings and linked to health information to enable research on voice as a biomarker of health. Version 1.0 provides 12,523 recordings for 306 adult participants collected across five sites in North America. Participants were selected from cohorts where conditions manifest in the voice waveform, including voice disorders, neurological disorders, mood disorders, respiratory disorders, and a pediatric cohort (pediatric data not included in v1.0). The initial release contains low-risk, de-identified derived data (e.g., spectrograms and acoustic/phonetic/prosodic features) and detailed demographic, clinical, and validated questionnaire data; original audio waveforms and free-speech transcripts are not included in v1.0.
en
2024-11-27
2024-11-27
  • voice
  • bridge2ai
  • audio
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  • ID
    gap-ethical-diverse-voice
    Name
    Addressing gap
    Response
    Addresses the lack of large, high-quality, multi-institutional, diverse, ethically sourced voice datasets linked to health information.
  • ID
    mech-protocol
    Name
    Collection mechanisms
    Description
    • Standardized multi-site protocol with a custom tablet application; headset used for data collection when possible.
  • ID
    collectors-sites
    Name
    Data collectors
    Description
    • Project investigators and site personnel at five North American sites
  • ID
    irb-approvals
    Name
    Ethical review
    Description
    • Data collection and sharing approved by the University of South Florida Institutional Review Board; submitted for review to the University of Toronto Research Ethics Board.
  1. ID
    acquisition-methods
    Name
    Instance acquisition
    Description
    • Directly observed voice tasks (e.g., sustained phonation) recorded under a standardized protocol; demographic, clinical, and validated questionnaires collected via app; derived features computed from standardized audio.
    Was Directly Observed
    yes
    Was Reported By Subjects
    yes
    Was Inferred Derived
    yes
    Was Validated Verified
    yes
  1. ID
    preprocessing-audio
    Name
    Audio preprocessing and feature derivation
    Description
    • Raw audio standardized to mono and 16 kHz with a Butterworth anti-aliasing filter; STFT spectrograms with 25 ms window, 10 ms hop, 512-point FFT.
    • Acoustic features via OpenSMILE; phonetic/prosodic features via Parselmouth/Praat; additional audio processing via torchaudio; transcriptions generated using OpenAI Whisper Large (transcripts of free speech removed from release).
    Used Software
    IDNameURLVersion
    software-opensmileopenSMILEhttps://audeering.github.io/opensmile/
    software-praatPraathttps://www.fon.hum.uva.nl/praat/
    software-parselmouthParselmouth (Python interface to Praat)https://parselmouth.readthedocs.io/
    software-torchaudiotorchaudiohttps://pytorch.org/audio/2.1
    software-whisperOpenAI Whisper Largehttps://github.com/openai/whisper
  1. ID
    labeling-transcription
    Name
    Transcription (removed in v1.0 release)
    Description
    • Automatic transcriptions were generated using OpenAI Whisper Large as part of derivation; free-speech transcripts were removed from the public release for de-identification.
    Used Software
  • ID
    raw-audio
    Name
    Raw data availability
    Description
    • Original audio waveforms are omitted from v1.0; only derived data are released. Future releases aim to include voice data with additional safeguards.
  • ID
    confidentiality-low-risk
    Name
    Confidentiality
    Description
    • Dataset is de-identified and considered low risk; HIPAA Safe Harbor identifiers removed; state/province removed; only country retained.
  • ID
    sensitive-health
    Name
    Sensitive elements
    Description
    • De-identified demographic, clinical, and validated questionnaire data are included.
RoleNameORCIDAffiliation
ContributorMaintainersmaintainers-host-
ID
deid-hipaa-safe-harbor
Name
De-identification
Description
  • HIPAA Safe Harbor identifiers removed.
  • State and province removed; country retained.
  • Transcripts of free speech audio removed.
  • Audio waveforms omitted from v1.0; only derived features/spectrograms released.
partially (mixed: tabular phenotype/features + array-like spectrograms)
DescriptionFormatIDMedia TypeNamePath
Parquet file storing dense spectrogram data derived from standardized audio; includes participant_id...file-spectrograms-parquetapplication/x-parquetspectrograms.parquetspectrograms.parquet
Tab-delimited participant-level data (demographics, acoustic confounders, validated questionnaires);...file-phenotype-tsvtext/tab-separated-valuesphenotype.tsvphenotype.tsv
Data dictionary describing columns in phenotype.tsv.JSONfile-phenotype-jsonapplication/jsonphenotype.jsonphenotype.json
Derived acoustic/phonetic/prosodic features with one row per recording.file-static-features-tsvtext/tab-separated-valuesstatic_features.tsvstatic_features.tsv
Data dictionary describing columns in static_features.tsv.JSONfile-static-features-jsonapplication/jsonstatic_features.jsonstatic_features.json
ID
ip-restrictions
Name
IP restrictions
Description
  • Not specified.
ID
export-regulatory
Name
Export/regulatory restrictions
Description
  • Not specified.
Source clinical/phenotype data collected via custom app and exported from REDCap (Bridge2AI Voice REDCap v3.20.0; https://doi.org/10.5281/zenodo.14148755). Project documentation: "https://docs.b2ai-voice.org/"
🚀

Uses

What (other) tasks could the dataset be used for?

  • ID
    task-voice-ai
    Name
    Target tasks
    Response
    AI/ML research using derived voice features and clinical/phenotype data, such as disorder detection, stratification, and exploration of voice–health associations.
ID
license-registered-access
Name
License and use terms
Description
  • Access policy: Only credentialed users who sign the DUA can access the files.
  • License: Bridge2AI Voice Registered Access License.
  • Data Use Agreement: Bridge2AI Voice Registered Access Agreement.
  • Required training: TCPS 2: CORE 2022.
  • ID
    other-tasks
    Name
    Other potential tasks
    Description
    • Exploratory analyses of voice–health associations; development/evaluation of AI models for condition screening and monitoring using derived voice features.
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
ID
versioning
Name
Version access
Description
  • Latest version DOI: https://doi.org/10.57764/3sg0-7440. Version 1.0 DOI: https://doi.org/10.57764/qb6h-em84.
🔄

Maintenance

How will the dataset be maintained?

2024-11-27
1.0
ID
update-plan
Name
Update plan
Description
  • v1.0 is the initial release. Future releases aim to include voice audio with additional precautions to ensure data security.
Generated on 2025-11-09 10:17:35 using Bridge2AI Data Sheets Schema