healthnexus tab references d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • Name
    Purpose
    Response
    Create an ethically sourced, diverse, multi-institutional voice dataset linked to clinical information to enable AI research on voice as a biomarker of health.
GrantorGrant NameGrant Number
National Institutes of HealthBridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before.3OT2OD032720-01S1
📊

Composition

What do the instances represent?

  1. Name
    Instance description
    Representation
    Voice recordings with derived spectrograms/features and participant-level phenotype/clinical data.
    Instance Type
    participants, sessions, and recordings (participant_id, session_id, task_name).
    Data Type
    Derived spectrogram arrays; acoustic, phonetic, and prosodic features; limited transcriptions; tabular phenotype and feature files.
    Counts
    12,523
    Label
    Cohort/disease group membership and clinical/phenotype variables; no raw audio included in v1.0.
    Sampling Strategies
    1. Name
      Targeted clinical cohort sampling
      Is Sample
      • sample from patients presenting at specialty clinics at five North American sites
      Is Random
      • False
      Source Data
      • Adults with conditions affecting voice (voice, neurological/neurodegenerative, mood/psychiatric, respiratory disorders)
      Is Representative
      • not stated
      Why Not Representative
      • Targeted sampling of specific disorders; pediatric cohort not included in v1.0
      Strategies
      • targeted clinical cohort sampling at participating sites
  • Name
    Cohort categories
    Identification
    • Disease Cohorts Identified At Enrollment
      Voice disorders; Neurological and Neurodegenerative; Mood and Psychiatric; Respiratory; Pediatric (planned but not included in v1.0).
    Distribution
    • Adult cohort only in v1.0; 306 participants across five sites in North America.
  • Name
    Distribution formats
    Description
    • Parquet (spectrograms.parquet)
    • TSV (phenotype.tsv, static_features.tsv)
    • JSON (phenotype.json, static_features.json)
  • Name
    Initial release
    Description
    • 2024-11-27 (v1.0)
🔍

Collection Process

How was the data acquired?

Bridge2AI-Voice v1.0
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
The Bridge2AI-Voice project presents a comprehensive, ethically sourced dataset enabling research on voice as a biomarker of health. Bridge2AI-Voice v1.0 provides 12,523 recordings for 306 adult participants collected across five sites in North America, with corresponding demographic, clinical, and validated questionnaire information. The initial release is considered low risk and includes derived data (e.g., spectrograms, acoustic/phonetic/prosodic features) and data dictionaries; original voice audio waveforms are not included. Participants were enrolled in predetermined groups reflecting conditions that manifest in voice (voice disorders, neurological and neurodegenerative disorders, mood and psychiatric disorders, respiratory disorders), with pediatric cohorts planned but not included in v1.0. Data were collected via a standardized protocol and subsequently de-identified using HIPAA Safe Harbor principles.
en
2024-11-27
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  • voice
  • bridge2ai
  • audio
  • biomarker
  • spectrograms
  • clinical
  • health
  • Name
    AddressingGap
    Response
    Address the lack of large, high-quality, multi-institutional, demographically diverse voice datasets linked to health and clinical data.
  • Name
    Sensitive health data
    Description
    • Contains health-related clinical and demographic questionnaire information.
Name
De-identification status
Description
  • HIPAA Safe Harbor identifiers removed (e.g., names, contact details, finer-than-year dates, identifiers).
  • State and province removed; country of data collection retained.
  • Transcripts of free speech audio removed.
  • Original audio waveforms omitted in v1.0; only derived spectrograms and features released.
  1. Name
    Data acquisition
    Description
    • Standardized protocol at specialty clinics; demographic, clinical, and validated questionnaires; voice tasks (e.g., sustained vowel).
    • Custom tablet application used; headset microphone when possible.
    • Some participants completed multiple sessions.
    Was Directly Observed
    yes (voice recordings)
    Was Reported By Subjects
    yes (validated questionnaires)
    Was Inferred Derived
    yes (spectrograms, acoustic/phonetic/prosodic features, transcriptions)
    Was Validated Verified
    Standardized data collection protocol; IRB/REB oversight.
  • Name
    Collection mechanisms
    Description
    • Custom tablet application for standardized data capture; headset microphone when feasible.
    • Export and conversion from REDCap using an open-source b2aiprep library.
    Used Software
    NameURLVersion
    REDCap3.20.0
    b2aiprephttps://github.com/sensein/b2aiprep
  • Name
    Data collection team
    Description
    • Project investigators at five North American clinical sites.
  • Name
    Ethics and review
    Description
    • Data collection and sharing approved by the University of South Florida Institutional Review Board.
    • Submitted for review to the University of Toronto Research Ethics Board.
  • Name
    Audio preprocessing and feature extraction
    Description
    • Raw audio converted to mono and resampled to 16 kHz with a Butterworth anti-aliasing filter.
    • Spectrograms computed via STFT with 25 ms window, 10 ms hop, 512-point FFT.
    • Acoustic features via OpenSMILE; phonetic/prosodic features via Parselmouth and Praat.
    • Transcriptions generated with OpenAI's Whisper Large model (free-speech transcripts later removed from release).
    Used Software
    NameVersion
    OpenSMILE
    Parselmouth
    Praat
    Torchaudio2.1
    OpenAI Whisper Large
  • Name
    De-identification and release filtering
    Description
    • HIPAA Safe Harbor removal; state/province removed; country retained.
    • Free-speech transcripts removed; only derived data released (no raw audio).
  • Name
    Transcription
    Description
    • Automatic transcriptions generated using OpenAI's Whisper Large model; free-speech transcripts removed prior to release.
    Used Software
    • Name
      OpenAI Whisper Large
  • Name
    Raw audio availability
    Description
    • Raw audio waveforms are not included in v1.0; only spectrograms and derived features are provided. Future releases aim to include voice waveforms with additional security precautions.
  1. Name
    External resources
    External Resources
    Archival
    • Version-specific and latest DOIs provided.
    Restrictions
    • Registered, credentialed access; signed DUA and training required.
Name
Retention limits
Description
  • Not specified in source.
mixed (tabular TSV/JSON dictionaries and array-based Parquet spectrograms)
DescriptionIDMedia TypeNamePathTitle
Parquet dataset containing 513xN spectrograms per recording with participant_id, session_id, and tas...spectrograms.parquetapplication/x-parquetspectrograms.parquetspectrograms.parquetSpectrograms derived from voice waveforms
Tab-delimited participant-level demographics, acoustic confounders, and validated questionnaire resp...phenotype.tsvtext/tab-separated-valuesphenotype.tsvphenotype.tsvParticipant-level phenotype data
JSON data dictionary describing columns in phenotype.tsv.phenotype.jsonapplication/jsonphenotype.jsonphenotype.jsonPhenotype data dictionary
Tab-delimited features derived from raw audio (one row per recording).static_features.tsvtext/tab-separated-valuesstatic_features.tsvstatic_features.tsvRecording-level static features
JSON data dictionary describing columns in static_features.tsv.static_features.jsonapplication/jsonstatic_features.jsonstatic_features.jsonStatic features data dictionary
🚀

Uses

What (other) tasks could the dataset be used for?

  • Name
    Task
    Response
    AI/ML research on voice-based biomarkers for disease screening, monitoring, and prognosis.
  • Name
    Potential uses
    Description
    • Disease screening and therapeutic monitoring for respiratory, voice, neurological, and mood/psychiatric disorders.
    • Development and evaluation of AI methods for voice analysis.
  • Name
    Considerations for future use
    Description
    • v1.0 includes only derived features and spectrograms (no raw audio), which may limit tasks requiring waveforms.
    • Free-speech transcripts removed.
    • Adult-only dataset in v1.0; pediatric cohort not yet included.
  • Name
    Access and licensing
    Description
    • License
      Bridge2AI Voice Registered Access License.
    • Access Policy
      Only credentialed users who sign the Bridge2AI Voice Registered Access Agreement.
    • Required Training
      TCPS 2: CORE 2022.
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
Name
Versioning and access to older versions
Description
🔄

Maintenance

How will the dataset be maintained?

1.0
2024-11-27
Name
Update plan
Description
  • Future releases aim to include original voice audio waveforms with additional security measures; see latest DOI for updates.
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema