healthnexus tab metadata d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • Name
    Purpose
    Response
    Create an ethically sourced, diverse, multi-institutional voice dataset linked to health information to enable future AI research and insights into voice as a biomarker of health.
GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before3OT2OD032720-01S1
📊

Composition

What do the instances represent?

  1. Name
    Instance structure
    Representation
    Voice recordings linked to clinical and questionnaire information
    Instance Type
    Participants, recording sessions, and derived per-recording data
    Data Type
    Derived data only: STFT spectrograms, acoustic/phonetic/prosodic features, and (non-free-speech) transcriptions; tabular phenotype data per participant.
    Counts
    12,523
    Label
    Dataset includes cohort membership and clinical/questionnaire variables; no explicit task labels provided in this release.
    Sampling Strategies
    1. Name
      Sampling
      Is Sample
      • True
      Is Random
      • False
      Source Data
      • Patients at specialty clinics across five sites in North America
      Is Representative
      • False
      Why Not Representative
      • Targeted enrollment of disease cohorts with known voice manifestations
      Strategies
      • Purposive/clinical cohort-based sampling
    Missing Information
  • Name
    Disease cohorts
    Identification
    • Voice disorders
    • Neurological and neurodegenerative disorders
    • Mood and psychiatric disorders
    • Respiratory disorders
    Distribution
    • Adult cohort only in v1.0; detailed cohort counts not provided.
  • Name
    Release files
    Description
    • spectrograms.parquet (derived spectrogram data)
    • static_features.tsv and static_features.json (per-recording features and data dictionary)
    • phenotype.tsv and phenotype.json (per-participant phenotype data and data dictionary)
  • Name
    Initial release
    Description
    • 2024-11-27
🔍

Collection Process

How was the data acquired?

Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice v1.0 is a restricted-access dataset enabling research into voice as a biomarker of health. The initial release provides 12,523 recordings for 306 adult participants collected across five sites in North America, with accompanying demographic, clinical, and validated questionnaire data. To reduce re-identification risk, only derived data are released (e.g., spectrograms and extracted acoustic/phonetic features); original audio waveforms and free-speech transcripts are not included. Participants were selected from cohorts with conditions known to manifest in voice (voice disorders, neurological and neurodegenerative disorders, mood and psychiatric disorders, and respiratory disorders). Standardized collection protocols were used; raw audio was converted to mono, resampled to 16 kHz with a Butterworth anti-aliasing filter, and used to compute STFT spectrograms and features (openSMILE, Praat/Parselmouth, torchaudio). Transcriptions were generated using Whisper Large but free speech transcripts were removed for release.
en
2024-11-27
  • voice
  • bridge2ai
  • audio
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  • Name
    Addressing gap
    Response
    Addresses the lack of large, high-quality, diverse, standardized, multi-institutional voice datasets linked to health biomarkers suitable for AI research.
  • Name
    Bridge2AI-Voice Team
  • Name
    Adult cohort v1.0
    Is Data Split
    no
    Is Subpopulation
    Adult cohort only in v1.0
  • Name
    Participant-session-recording linkage
    Description
    • Participants may have multiple sessions; sessions include multiple recordings/tasks.
  1. Name
    Documentation
    External Resources
    Future Guarantees
    • Not stated
    Archival
    • DOI-registered dataset landing page at Health Data Nexus
    Restrictions
    • Credentialed access with required training and DUA
  • Name
    Confidentiality
    Description
    • Clinical and demographic data are included; identifiers removed per HIPAA Safe Harbor.
Name
De-identification
Description
  • HIPAA Safe Harbor identifiers removed (e.g., names, detailed dates, contact and device identifiers).
  • State/province removed; country of data collection retained.
  • Free speech transcripts removed.
  • Original audio waveforms omitted from v1.0; only derived data released.
  • Name
    Sensitive data
    Description
    • Contains health-related and demographic information.
  1. Name
    Data acquisition
    Description
    • Standardized protocol with voice tasks (e.g., sustained vowel), clinical and targeted questionnaires.
    • Participants enrolled from specialty clinics into predefined cohorts.
    Was Directly Observed
    yes (voice recordings)
    Was Reported By Subjects
    yes (questionnaires)
    Was Inferred Derived
    yes (features and spectrograms derived from raw audio)
    Was Validated Verified
    Standardized acquisition protocols were used; preprocessing and feature extraction applied consistently.
  • Name
    Collection mechanisms
    Description
    • Custom tablet application; headset used for data collection when possible.
    • Data exported and converted from REDCap using an open-source library.
  • Name
    Data collectors
    Description
    • Project investigators at five North American sites; participant screening against inclusion/exclusion criteria.
  • Name
    Ethics and IRB/REB review
    Description
    • Data collection and sharing approved by the University of South Florida Institutional Review Board.
    • Submitted for review to the University of Toronto Research Ethics Board.
  • Name
    Audio preprocessing and feature extraction
    Description
    • Mono conversion; resampled to 16 kHz with a Butterworth anti-aliasing filter.
    • STFT spectrograms using 25 ms window, 10 ms hop, 512-point FFT.
    • Acoustic features via openSMILE.
    • Phonetic/prosodic measures via Parselmouth and Praat.
    • Derived audio features via torchaudio.
    Used Software
    NameURL
    b2aiprephttps://github.com/sensein/b2aiprep
    openSMILE
    Praat
    Parselmouth
    torchaudio
  • Name
    Transcription
    Description
    • Transcriptions generated using OpenAI Whisper Large; free speech transcripts removed from release.
  • Name
    De-identification and release filtering
    Description
    • Removal of HIPAA Safe Harbor identifiers.
    • Removal of state/province; retention of country only.
    • Exclusion of free speech transcripts and original audio waveforms from v1.0.
  • Name
    Raw audio
    Description
    • Collected but not released in v1.0; derived spectrograms and features provided instead.
  • Name
    Hosting and support
    Description
    • Health Data Nexus; Temerty Centre for AI Research and Education in Medicine (supported by the Temerty Foundation).
🚀

Uses

What (other) tasks could the dataset be used for?

  • Name
    Intended tasks
    Response
    Development and evaluation of AI methods for health-related voice analytics using derived spectrograms and acoustic/phonetic features; exploration of associations between voice signals and clinical/demographic factors.
Name
License and access terms
Description
  • License
    Bridge2AI Voice Registered Access License.
  • Access Policy
    Only credentialed users who sign the DUA can access the files.
  • Data Use Agreement
    Bridge2AI Voice Registered Access Agreement.
  • Required Training
    TCPS 2: CORE 2022.
📤

Distribution

How will the dataset be distributed?

Name
Versioning and access
Description
🔄

Maintenance

How will the dataset be maintained?

1.0
Name
Update plan
Description
  • Future releases aim to include voice waveforms with additional data security precautions.
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema