healthnexus tab abstract d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before3OT2OD032720-01S1
  • Response
    Create an ethically sourced flagship voice dataset to enable AI research on voice as a biomarker of health and support clinical insights.
πŸ“Š

Composition

What do the instances represent?

CountsData TypeInstance TypeNameRepresentation
12523Spectrogram matrices (513 x N), static acoustic/phonetic/prosodic feature vectorsrecording-derived instanceRecordings-derived instancesDerived artifacts from voice recordings (e.g., spectrograms and engineered features)
306Demographics, clinical information, and validated questionnaire responsesparticipantParticipantsIndividual study participants enrolled across five North American sites
  • Identification
    • Adult cohort only in v1.0; participants selected based on known conditions with voice manifestations (voice, neurological, mood/psychiatric, respiratory).
    Distribution
    • Counts by subgroup not specified.
  • Description
    • spectrograms.parquet β€” Parquet file with 513 x N spectrogram matrices and identifiers (participant_id, session_id, task_name).
    • phenotype.tsv β€” Tab-delimited participant-level data (demographics, acoustic confounders, validated questionnaires).
    • phenotype.json β€” Data dictionary for phenotype.tsv.
    • static_features.tsv β€” Recording-level engineered features (OpenSMILE, Praat, Parselmouth, torchaudio).
    • static_features.json β€” Data dictionary for static_features.tsv.
  • Description
    • Initial Public Release
      2024-11-27
πŸ”

Collection Process

How was the data acquired?

Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically sourced dataset of data derived from voice recordings linked to corresponding clinical and demographic information, intended to enable AI research on voice as a biomarker of health. Version 1.0 provides 12,523 recordings for 306 participants collected across five sites in North America. Participants were selected based on conditions known to manifest in the voice waveform, including voice disorders, neurological and neurodegenerative disorders, mood and psychiatric disorders, and respiratory disorders. The initial release contains low-risk derived data (e.g., spectrograms and engineered features) and does not include original audio waveforms. Detailed demographic, clinical, and validated questionnaire data are also available.
2024-11-27
2024-11-27
  • voice
  • bridge2ai
  • audio
Name
Alistair Johnson
Jean-Christophe BΓ©lisle-Pipon
David Dorr
Satrajit Ghosh
Philip Payne
Maria Powell
AnaΓ―s Rameau
Vardit Ravitsky
Alexandros Sigaras
Olivier Elemento
Yael Bensoussan
  • Response
    Address the need for a large, high-quality, multi-institutional, and diverse voice database linked to health biomarkers with standardized collection protocols.
  • Name
    Adult cohort v1.0
    Description
    Initial dataset release includes only the adult cohort with derived data from voice recordings and linked clinical/phenotype data.
    Is Subpopulation
    yes
  1. Strategies
    • Purposeful sampling of patients at specialty clinics into five predetermined disease/cohort groups (Respiratory, Voice, Neurological, Mood, Pediatric)
    Is Sample
    • yes
    Is Random
    • no
    Source Data
    • Patients presenting at specialty clinics/institutions across five North American sites
    Is Representative
    • no
    Why Not Representative
    • Clinic-based cohort with inclusion/exclusion criteria; participants selected based on known conditions affecting voice.
  • Description
    • Participants may have one or more sessions; sessions include multiple recording tasks. Each spectrogram/feature row links via participant_id, session_id, and task_name.
  • Description
    • No recommended train/validation/test splits are provided in v1.0.
  • Description
    • Not specified in the release notes.
  • Description
    • Dataset includes clinical and demographic elements associated with participants; access is restricted and governed by DUA and required training.
  • Warnings
    • None noted.
Description
  • HIPAA Safe Harbor identifiers removed (e.g., names, fine-grained dates, contact info, device/biometric identifiers, etc.).
  • State and province removed; country of data collection retained.
  • Transcripts of free speech audio removed.
  • Original audio waveforms omitted in v1.0; only derived data are distributed.
  • Description
    • Health-related information (clinical and questionnaire data) linked to voice-derived features.
  1. Description
    • Voice recordings directly observed; questionnaires reported by subjects; multiple derived features computed from raw audio.
    Was Directly Observed
    yes
    Was Reported By Subjects
    yes
    Was Inferred Derived
    yes
    Was Validated Verified
    Standardized data collection protocol; derived features generated via established tools (OpenSMILE, Praat/Parselmouth, torchaudio). Ethics approvals in place.
  • Description
    • Standardized protocol; data collected via custom tablet application using headset when possible; export and conversion from REDCap using an open-source library.
  • Description
    • Project investigators at specialty clinics/institutions across five sites in North America; participants consented prior to data collection.
  • Description
    • Data collection and sharing approved by the University of South Florida Institutional Review Board; submitted for review to the University of Toronto Research Ethics Board.
  • Description
    • Eligible patients provided informed consent for data collection and for sharing acquired research data prior to participation.
  • Description
    • Participants were screened and informed as part of the consent process for the data collection initiative and data sharing.
  • Description
    • Raw audio converted to mono and resampled to 16 kHz with a Butterworth anti-aliasing filter; derived spectrograms, acoustic, phonetic/prosodic features, and transcriptions generated.
    Used Software
    NameURL
    openSMILE
    Praat
    Parselmouth
    torchaudio
    OpenAI Whisper Large
    b2aiprephttps://github.com/sensein/b2aiprep
  • Description
    • De-identification and removal of identifiers per HIPAA Safe Harbor; removal of free speech transcripts; exclusion of original audio in v1.0.
  • Description
    • Automatic transcription using OpenAI Whisper Large; validated questionnaires administered during clinical data collection.
  • Description
    • Original audio waveforms were collected but are not included in v1.0; only derived data are available. Future releases aim to include audio with additional safeguards.
Description
  • None specified.
  • Description
    • Health Data Nexus (Temerty Centre for AI Research and Education in Medicine)
πŸš€

Uses

What (other) tasks could the dataset be used for?

Response
Develop and evaluate AI/ML models using derived voice features (spectrograms, acoustic, phonetic, an...
Study associations between voice-derived features and health conditions (voice, neurological, mood/p...
  • Description
    • Not specified.
  • Description
    • Screening/monitoring of respiratory conditions using cough/breath/voice-derived features; detection and characterization of voice disorders; analysis of neurological and mood/psychiatric condition markers in voice features.
  • Description
    • Users should note that v1.0 includes only de-identified derived features without raw audio; future inclusion of audio will require additional safeguards.
  • Description
    • Not specified.
Description
  • Access Policy
    Only credentialed users who sign the Data Use Agreement (DUA) can access files.
  • License
    Bridge2AI Voice Registered Access License.
  • Data Use Agreement
    Bridge2AI Voice Registered Access Agreement.
  • Required Training
    TCPS 2: CORE 2022.
  • Versioned DOIs provided for citation and discovery.
πŸ“€

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
Description
  • Versioned Dois Are Provided
    v1.0 DOI (https://doi.org/10.57764/qb6h-em84) and a latest-version DOI (https://doi.org/10.57764/3sg0-7440).
πŸ”„

Maintenance

How will the dataset be maintained?

1.0
2024-11-27
Description
  • Future releases aim to include original voice audio with additional data security precautions.
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema