healthnexus tab background d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • Response
    Create an ethically sourced flagship dataset to enable AI research on voice as a biomarker of health and support insights across multiple clinical domains.
GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before3OT2OD032720-01S1
📊

Composition

What do the instances represent?

CountsData TypeInstance TypeRepresentation
12523Derived audio data: spectrograms (513 x N), acoustic features (openSMILE), phonetic and prosodic fea...Audio-derived features and spectrograms per recordingVoice recordings (derived)
306Demographics, clinical information, and validated questionnaire responsesHuman subjects enrolled across five North American sitesParticipants
  1. Description
    • Clinical voice and health data acquired at point of care via standardized tasks and questionnaires; derived features computed from raw audio.
    Was Directly Observed
    yes
    Was Reported By Subjects
    yes
    Was Inferred Derived
    yes
    Was Validated Verified
    Validated questionnaires were used; derived signals followed standardized preprocessing.
  • Identification
    • Adult cohort (v1.0)
    • Disorder Cohorts
      voice disorders, neurological disorders, mood/psychiatric disorders, respiratory disorders
    Distribution
    • Data provided across five North American collection sites; detailed distributions in phenotype files and data dictionary
  • Description
    • spectrograms.parquet (derived spectrograms; participant_id, session_id, task_name, 513xN spectrogram)
    • phenotype.tsv (participant-level demographics, clinical data, validated questionnaires)
    • phenotype.json (data dictionary for phenotype)
    • static_features.tsv (recording-level derived acoustic/phonetic features)
    • static_features.json (data dictionary for features)
  • Description
    • Initial public release on 2024-11-27 (v1.0)
🔍

Collection Process

How was the data acquired?

Bridge2AI-Voice v1.0
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically sourced dataset linking derived voice recordings to clinical and demographic information to advance research on voice as a biomarker of health. Version 1.0 (released Nov 27, 2024) includes 12,523 recordings from 306 adult participants across five North American sites. This initial release contains low-risk derived data (e.g., spectrograms and acoustic/phonetic features) and detailed demographics, clinical data, and validated questionnaire responses. Original audio waveforms are omitted in this release; transcripts of free speech are removed. The dataset supports AI research across cohorts including voice disorders, neurological disorders, mood/psychiatric disorders, and respiratory disorders.
2024-11-27
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  • voice
  • bridge2ai
  • audio
  • VOICE
  • Response
    Lack of large, high-quality, multi-institutional, demographically diverse voice datasets linked to health information and collected under standardized, ethically grounded protocols.
  1. ID
    bridge2ai-voice-adult-v1.0
    Name
    Adult cohort v1.0
    Title
    Adult cohort (derived data only)
    Description
    Adult participants only for the initial release; derived data provided, original audio omitted.
    Is Data Split
    no
    Is Subpopulation
    yes
  1. Is Sample
    • yes
    Is Random
    • no
    Source Data
    • Patients at specialty clinics enrolled into predefined disorder cohorts (adult cohort in v1.0)
    Is Representative
    • no
    Why Not Representative
    • Disorder-focused recruitment selected for specific conditions; not a general population sample
    Strategies
    • Deterministic cohort assignment based on inclusion/exclusion criteria within five sites
  • Description
    • Project investigators at five specialty clinical sites in North America collected data under a standardized protocol.
  • Description
    • Standardized protocol using a custom tablet application; headset used for data collection when possible; REDCap used for source data capture and export.
  • Description
    • Data collection and sharing approved by the University of South Florida IRB; submitted to the University of Toronto Research Ethics Board.
  • Description
    • Raw audio converted to monaural, resampled to 16 kHz with a Butterworth anti-aliasing filter; spectrograms computed via short-time FFT (25 ms window, 10 ms hop, 512-point FFT); acoustic features extracted with openSMILE; phonetic/prosodic features via Parselmouth/Praat; automatic transcriptions via Whisper Large; integration and parquet generation via b2aiprep.
    Used Software
    Name
    openSMILE
    Parselmouth
    Praat
    torchaudio
    Whisper Large
    b2aiprep
  • Description
    • Standardized signal processing (resampling, channel conversion, anti-alias filtering); harmonized integration of sources (REDCap exports) into phenotype and feature tables.
  • Description
    • Automatic transcriptions generated using OpenAI Whisper Large; transcripts of free-speech audio removed in this release.
  • Description
    • Original audio waveforms were collected but are not included in v1.0 distribution; only derived data (e.g., spectrograms, features) are released.
  • External Resources
    Archival
    • Versioned DOIs provided (version-specific and latest)
    Restrictions
    • Access governed by registered access license, DUA, credentialing, and required training
Description
  • HIPAA Safe Harbor identifiers removed (e.g., names, granular dates, contact numbers, IPs, MRNs, etc.)
  • State and province removed; country of data collection retained
  • Transcripts of free-speech audio removed
  • Original audio waveforms omitted from v1.0; only derived data released
  • Description
    • Clinical and demographic information and validated questionnaire responses related to health conditions
  • Description
    • Health-related data subject to ethical oversight; de-identified per HIPAA Safe Harbor and restricted access controls
Description
Distributed to credentialed external users via Health Data Nexus under registered access terms
  • Description
    • Health Data Nexus (Temerty Centre for AI Research and Education in Medicine)
🚀

Uses

What (other) tasks could the dataset be used for?

  • Response
    Develop and evaluate AI/ML methods for health-related inference from voice, including detection, classification, and risk stratification across specified disorder cohorts.
Description
  • License
    Bridge2AI Voice Registered Access License
  • Data Use Agreement
    Bridge2AI Voice Registered Access Agreement
  • Access Policy
    Only credentialed users who sign the DUA can access the files
  • Required Training
    TCPS 2: CORE 2022
  • Versioned Dois
    version-specific (https://doi.org/10.57764/qb6h-em84) and latest (https://doi.org/10.57764/3sg0-7440)
📤

Distribution

How will the dataset be distributed?

🔄

Maintenance

How will the dataset be maintained?

1.0
Description
  • Future releases aim to include original voice data with additional safeguards; continued expansion of cohorts and data elements anticipated
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema