healthnexus tab versions d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • Name
    Purpose
    Response
    Create an ethically sourced, diverse, multi-institutional voice dataset linked to health information to enable AI research on voice as a biomarker of health.
GrantorGrant NameGrant Number
National Institutes of HealthBridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before3OT2OD032720-01S1
📊

Composition

What do the instances represent?

  1. Name
    Instance description
    Representation
    Voice-derived data linked to demographic, clinical, and validated questionnaire information.
    Instance Type
    Participants with one or more recording sessions; one row per recording in features; one row per participant in phenotype.
    Data Type
    Derived data including: - Spectrograms (513 x N time-frequency representations from 16 kHz monaural audio) - Acoustic features (e.g., openSMILE) - Phonetic/prosodic features (e.g., Praat/Parselmouth) - Phenotype/demographics/clinical/questionnaire responses Note: Raw audio waveforms are not included in v1.0.
    Counts
    12,523
    Label
    Task identifiers (task_name) per recording; no explicit diagnostic labels distributed in v1.0.
    Sampling Strategies
    1. Name
      Sampling
      Is Sample
      • Yes; participants selected from specialty clinics.
      Is Random
      • No; condition-based targeted enrollment.
      Source Data
      • Patients presenting at participating specialty clinics across five North American sites.
      Is Representative
      • Not intended to be representative of the general population; targeted by condition cohorts.
      Why Not Representative
      • Cohort design targets specific diseases and clinical populations.
      Strategies
      • Targeted clinical cohort recruitment based on predefined disease categories (voice, neurological/neurodegenerative, mood/psychiatric, respiratory; pediatric planned but not included in v1.0).
DistributionIdentificationName
Distribution by cohort not provided in the release page.Participants selected into disease cohort categories (voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, respiratory disorders).Disease cohorts
Pediatric cohort planned for future releases; not present in v1.0.As of v1.0, only adult participants are included.Adult cohort v1.0
  • Name
    Formats
    Description
    • Parquet (.parquet) for spectrograms
    • TSV (.tsv) for phenotype and static features
    • JSON (.json) data dictionaries for phenotype and features
  • Name
    DistributionDate
    Description
    • 2024-11-27
🔍

Collection Process

How was the data acquired?

bridge2ai-voice-1.0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice v1.0 is a comprehensive collection of data derived from voice recordings with corresponding clinical information to enable research on voice as a biomarker of health. The initial release provides 12,523 recordings for 306 participants collected across five sites in North America. Participants were selected based on known conditions which manifest within the voice waveform including voice disorders, neurological disorders, mood disorders, and respiratory disorders. As of v1.0, only data from the adult cohort is available. This initial release contains data considered low risk, including derivations such as spectrograms and engineered features, but not the original voice recordings. Detailed demographic, clinical, and validated questionnaire data are also made available. Data collection used a standardized protocol with a custom tablet application, headset recording when possible, and REDCap; preprocessing included resampling to 16 kHz with a Butterworth anti-aliasing filter and derivation of spectrograms and acoustic/phonetic/prosodic features.
English
2024-11-27
  • voice
  • bridge2ai
  • audio
  • biomarker
  • spectrogram
  • clinical
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  • Name
    AddressingGap
    Response
    Addresses the need for a large, high-quality, diverse, multi-institutional voice dataset linked to other health biomarkers with standardized data collection, ethical sourcing, and detailed clinical/phenotypic context.
  1. Name
    Cohort sampling
    Is Sample
    • Yes; clinical cohort-based sample.
    Is Random
    • No.
    Source Data
    • Specialty clinic patients screened against inclusion/exclusion criteria and consented.
    Is Representative
    • Not of the general population; designed for disease-relevant cohorts.
    Representative Verification
    • Standardized protocol across sites; not intended as population-representative sampling.
    Strategies
    • Deterministic cohort selection based on predefined conditions.
  • Name
    DataAnomaly
    Description
    • None reported in the release notes.
  1. Name
    Documentation and tooling
    External Resources
    Future Guarantees
    • Not specified.
    Archival
    • Not specified.
    Restrictions
    • None beyond dataset access controls for data files.
  • Name
    Confidentiality
    Description
    • Dataset is released in de-identified, low-risk form; original voice recordings and free-speech transcripts are not included in v1.0.
  • Name
    ContentWarning
    Warnings
    • None noted.
  • Name
    Sensitive elements
    Description
    • Clinical and health-related questionnaire responses.
    • Voice-related features may be considered biometric information.
Name
Deidentification
Description
  • HIPAA Safe Harbor identifiers removed.
  • State and province removed; country of data collection retained.
  • Free speech transcripts removed.
  • Raw audio waveforms omitted; only derived spectrograms and features included in v1.0.
  • Dataset characterized as low risk for this initial release.
  1. Name
    Acquisition
    Description
    • Directly observed voice recordings; subject-reported validated questionnaires; derived features from raw audio.
    Was Directly Observed
    True
    Was Reported By Subjects
    True
    Was Inferred Derived
    True
    Was Validated Verified
    True
  • Name
    CollectionMechanism
    Description
    • Standardized multi-site protocol; custom tablet application; headset used when possible; data captured into REDCap and exported; preprocessing and merges performed with the open-source b2aiprep library.
    Used Software
    NameURLVersion
    Bridge2AI Voice REDCaphttps://doi.org/10.5281/zenodo.14148755v3.20.0
    b2aiprephttps://github.com/sensein/b2aiprep
  • Name
    DataCollector
    Description
    • Project investigators at participating specialty clinics; screening against inclusion/exclusion criteria; patient consent obtained prior to data collection.
  • Name
    EthicalReview
    Description
    • Data collection and sharing approved by the University of South Florida Institutional Review Board (IRB).
    • Submitted for review to the University of Toronto Research Ethics Board (REB).
  • Name
    Preprocessing
    Description
    • Raw audio converted to monaural, resampled to 16 kHz with a Butterworth anti-aliasing filter.
    • Spectrograms computed with STFT (25 ms window, 10 ms hop, 512-point FFT).
    • Acoustic features extracted (e.g., openSMILE); phonetic/prosodic features computed (Praat/Parselmouth); additional features via torchaudio.
    Used Software
    Name
    openSMILE
    Praat
    Parselmouth
    torchaudio
    b2aiprep
  • Name
    Cleaning
    Description
    • De-identification per HIPAA Safe Harbor; removal of state/province; removal of free speech transcripts; omission of original audio waveforms for v1.0.
  • Name
    Labeling
    Description
    • Automatic speech transcriptions generated using Whisper Large during processing; free speech transcripts were removed prior to release.
    Used Software
    • Name
      OpenAI Whisper Large
  • Name
    RawData
    Description
    • Raw audio waveforms were not released in v1.0; only derived representations (spectrograms and features) are distributed.
Name
IPRestrictions
Description
  • None stated beyond the registered access license and DUA requirements.
  • Name
    Maintainer
    Description
    • Health Data Nexus (Temerty Centre for AI Research and Education in Medicine), supported by the Temerty Foundation.
Mixed; spectrograms in Parquet, features and phenotype in TSV with JSON data dictionaries.
🚀

Uses

What (other) tasks could the dataset be used for?

  • Name
    Task
    Response
    AI method development and evaluation for detecting or characterizing health conditions from voice-derived representations (e.g., classification across disease cohorts).
  • Name
    OtherTask
    Description
    • Benchmarking of voice-based health AI models; feature analysis and biomarker discovery; method development for fair, robust voice-based health inference.
  • Name
    FutureUseImpact
    Description
    • Initial release omits raw audio to reduce privacy risk; users should consider potential sensitivity of voice-derived biometric features and clinical phenotypes in downstream applications.
Name
LicenseAndUseTerms
Description
  • License
    Bridge2AI Voice Registered Access License.
  • Access requires signing the Bridge2AI Voice Registered Access Agreement (DUA).
  • Access Is Restricted To Credentialed Users Who Complete Required Training (tcps 2
    CORE 2022).
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
🔄

Maintenance

How will the dataset be maintained?

1.0
Name
UpdatePlan
Description
  • Future releases aim to include original voice recordings with additional precautions to ensure data security; pediatric cohort planned for future inclusion.
Generated on 2025-11-09 10:17:35 using Bridge2AI Data Sheets Schema