healthnexus tab 2 d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • Name
    Purpose
    Response
    Create an ethically sourced flagship dataset to enable AI research and support insights into the use of voice as a biomarker of health across multiple clinical domains.
GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before3OT2OD032720-01S1
📊

Composition

What do the instances represent?

  1. Name
    Instance description
    Representation
    Voice recordings (and derived artifacts) per participant session, with linked clinical and questionnaire data at the participant level.
    Instance Type
    Participants, recording sessions, and per-session derived voice data
    Data Type
    Derived audio representations (spectrograms), acoustic/phonetic/prosodic features, and tabular phenotype data; original audio waveforms are not included in v1.0.
    Counts
    12,523
    Sampling Strategies
    1. Name
      Sampling approach
      Is Sample
      • Yes, selected from patients presenting at specialty clinics
      Is Random
      • False
      Source Data
      • Specialty clinics across five North American sites
      Is Representative
      • Not intended to be representative of the general population
      Why Not Representative
      • Cohort intentionally focused on predefined clinical groups with known voice manifestations
      Strategies
      • Consecutive/screened enrollment within predefined disease cohorts at participating sites
  • Name
    Clinical cohorts
    Identification
    • Voice disorders
    • Neurological/neurodegenerative disorders
    • Mood/psychiatric disorders
    • Respiratory disorders
    • Pediatric (planned for future releases; not in v1.0)
    Distribution
    • v1.0 includes adult participants across the listed cohorts; detailed cohort counts not provided here.
  • Name
    File formats (v1.0)
    Description
    • spectrograms.parquet (Parquet; derived spectrograms with participant_id, session_id, task_name)
    • phenotype.tsv (tab-delimited participant-level demographics, questionnaires, confounders)
    • phenotype.json (data dictionary for phenotype)
    • static_features.tsv (tab-delimited per-recording acoustic/phonetic/prosodic features)
    • static_features.json (data dictionary for features)
  • Name
    Initial public release
    Description
    • 2024-11-27
🔍

Collection Process

How was the data acquired?

Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically-sourced dataset linking derived voice data to clinical and phenotypic information to enable research into voice as a biomarker of health. Version 1.0 contains 12,523 recordings from 306 adult participants collected across five North American sites, focusing on conditions known to manifest in voice signals (voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, and respiratory disorders). To minimize re-identification risk, this initial release includes derived artifacts (e.g., spectrograms and acoustic/phonetic/prosodic features) and tabular phenotype data, but not original audio waveforms or free-speech transcripts. Data collection followed a standardized multi-site protocol with informed consent and IRB oversight.
English
2024-11-27
  • voice
  • bridge2ai
  • audio
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  • Name
    Gap addressed
    Response
    The pressing need for a large, high-quality, multi-institutional, diverse voice dataset linked to other health biomarkers to fuel reproducible voice-AI research with clinical relevance.
  1. ID
    bridge2ai-voice-v1.0-adult
    Name
    Adult cohort (v1.0)
    Description
    Version 1.0 includes only the adult cohort across five North American sites.
    Is Subpopulation
    adult
  • Name
    Participant-to-session linkage
    Description
    • Each participant may have one or multiple recording sessions; sessions are linked to participants and tasks.
  • Name
    Data splits
    Description
    • No recommended train/validation/test splits are provided in v1.0.
  1. Name
    Documentation
    External Resources
    Future Guarantees
    • Not stated
    Archival
    • DOI registered for versioned dataset record
    Restrictions
    • Registered access with DUA and required training
  • Name
    Clinical and questionnaire data
    Description
    • Contains de-identified clinical and questionnaire responses that could be considered confidential; access controlled.
  • Name
    Potentially sensitive health-related context
    Warnings
    • Health condition information may be sensitive for some users.
  • Name
    Health-related data
    Description
    • Health condition categories and clinical/questionnaire responses linked to participants.
Name
De-identification
Description
  • HIPAA Safe Harbor identifiers removed.
  • State and province removed; country of data collection retained.
  • Free-speech transcripts removed.
  • Original audio waveforms omitted from v1.0; only derived artifacts released.
  1. Name
    Instance acquisition
    Description
    • Data directly collected from patients at specialty clinics using a standardized protocol.
    Was Directly Observed
    Yes (voice tasks and recordings)
    Was Reported By Subjects
    Yes (validated and targeted questionnaires)
    Was Inferred Derived
    Yes (acoustic/phonetic/prosodic features, spectrograms, ASR transcripts; transcripts of free speech removed from release)
    Was Validated Verified
    Standardized multi-site protocol applied; IRB approval obtained; specific data validation steps beyond protocol not detailed.
  • Name
    Collection protocol and tooling
    Description
    • Standardized multi-site protocol; custom tablet application used; headset used for data collection when possible; REDCap used for data entry/export.
  • Name
    Data collection personnel
    Description
    • Project investigators at participating specialty clinics across five North American sites.
  • Name
    Collection timeframe
    Description
    • Multi-site data collection prior to the v1.0 publication; specific start/end dates not provided.
  • Name
    Ethics oversight
    Description
    • Data collection and sharing approved by the University of South Florida Institutional Review Board; submitted for review to the University of Toronto Research Ethics Board.
  • Name
    Audio preprocessing and feature extraction
    Description
    • Raw Audio Converted To Mono, Resampled To 16 Khz With Butterworth Anti Aliasing Filter; Derived Artifacts Generated
    • Spectrograms via STFT (25 ms window, 10 ms hop, 512-point FFT)
    • Acoustic features via OpenSMILE
    • Phonetic/prosodic features via Parselmouth and Praat
    • Transcriptions via OpenAI Whisper Large (free-speech transcripts removed in release)
    Used Software
    NameVersion
    OpenSMILE
    Parselmouth
    Praat
    TorchAudio2.1
    OpenAI Whisper Large
  • Name
    De-identification and release scoping
    Description
    • HIPAA Safe Harbor removal of identifiers; removal of state/province; exclusion of free-speech transcripts; omission of original audio waveforms from v1.0.
  • Name
    Automated transcription
    Description
    • ASR transcriptions generated using OpenAI's Whisper Large model; free-speech transcripts were not released.
  • Name
    Raw audio recordings
    Description
    • Original audio waveforms were collected but are omitted from v1.0; planned for future releases subject to additional safeguards.
Name
IP and third-party restrictions
Description
  • No specific third-party IP restrictions stated for released files; access governed by registered access license and DUA.
Name
Export control and regulatory restrictions
Description
  • No export control restrictions stated.
bibo:draft
Derived from raw audio recordings collected under a standardized clinical protocol; v1.0 releases derived artifacts (spectrograms and features) and de-identified phenotype tables.
🚀

Uses

What (other) tasks could the dataset be used for?

  • Name
    Primary task
    Response
    Development and evaluation of AI methods on derived voice representations (e.g., spectrograms and acoustic/phonetic/prosodic features) linked to clinical and phenotypic data for health-related research.
  • Name
    Considerations for future use
    Description
    • v1.0 contains only derived, low-risk data without raw audio, which may limit tasks requiring waveform-level analysis; future releases may alter risk profile when audio is included.
Name
Access and use terms
Description
  • Access is restricted to credentialed users who sign the Bridge2AI Voice Registered Access Agreement (DUA).
  • Required Training
    TCPS 2: CORE 2022.
  • Files are distributed under the Bridge2AI Voice Registered Access License.
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
🔄

Maintenance

How will the dataset be maintained?

2024-11-27T18:11:00
1.0
Name
Update plan
Description
  • Future releases aim to include original voice data with additional security precautions; pediatric cohort planned for future inclusion.
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema