healthnexus d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • Name
    Purpose
    Response
    Create an ethically sourced flagship dataset to enable AI research on voice as a biomarker of health by linking derived voice data with demographic, clinical, and validated questionnaire information.
GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database3OT2OD032720-01S1
📊

Composition

What do the instances represent?

CountsData TypeInstance TypeNameRepresentation
12523Derived spectrograms (513 x N), acoustic, phonetic, and prosodic features extracted from raw audio; ...Audio-derived data (per recording)Voice-derived recordingsVoice-derived recordings (spectrograms/features) linked to metadata
306Tabular phenotype data (one row per participant) with data dictionaryParticipant-level recordsParticipantsParticipants with linked demographic, clinical, and questionnaire data
  • Name
    Adult cohort
    Identification
    • v1.0 includes adults only across five disease cohorts
    Distribution
    • 306 participants; disease-targeted cohorts (voice, neurological, mood/psychiatric, respiratory)
  • Name
    Files in v1.0
    Description
    • spectrograms.parquet (derived spectrograms; 513 x N per recording; includes participant_id, session_id, task_name)
    • phenotype.tsv (participant-level demographics, acoustic confounders, validated questionnaires)
    • phenotype.json (data dictionary for phenotype.tsv)
    • static_features.tsv (one row per recording with acoustic/phonetic/prosodic features)
    • static_features.json (data dictionary for static_features.tsv)
  • Name
    Initial release
    Description
    • 2024-11-27
🔍

Collection Process

How was the data acquired?

Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information v1.0
Bridge2AI-Voice is a comprehensive collection of data derived from voice recordings linked to corresponding clinical information to enable research on voice as a biomarker of health. Version 1.0 provides 12,523 recordings for 306 participants collected across five sites in North America. Participants were selected based on conditions known to manifest in the voice waveform (voice, neurological, mood/psychiatric, and respiratory disorders). This initial release contains low-risk derived data (e.g., spectrograms and extracted features) and detailed demographic/clinical/questionnaire data; raw audio waveforms are not included in v1.0.
en
2024-11-27
  • voice
  • bridge2ai
  • audio
  • biomarker
  • health
  • credentialed access
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  • Name
    Gap addressed
    Response
    Addresses the lack of large, diverse, multi-institutional voice datasets linked to health information with standardized collection protocols and ethical oversight; mitigates prior limitations such as small datasets and limited demographic diversity reporting.
  1. Name
    Sampling approach
    Is Sample
    • True
    Is Random
    • False
    Source Data
    • Patients presenting at specialty clinics/institutions at five North American sites
    Is Representative
    • False
    Why Not Representative
    • Purposeful selection into five disease cohorts; v1.0 includes adults only
    Strategies
    • Deterministic cohort-based enrollment using inclusion/exclusion screening
  • Name
    Data collection mechanisms
    Description
    • Standardized protocol using a custom tablet application; headset used when possible
    • Clinical/demographic and validated questionnaires collected in-app
    • Data exported and converted from REDCap; processing via open-source b2aiprep library
  • Name
    Data collection personnel
    Description
    • Project investigators at specialty clinics and institutions screened and enrolled participants
  • Name
    Collection timeframe
    Description
    • Collected across five sites in North America; specific calendar dates not specified in this record
  • Name
    Ethics and IRB/REB review
    Description
    • Approved by the University of South Florida Institutional Review Board
    • Submitted for review to the University of Toronto Research Ethics Board
  • Name
    Audio preprocessing and feature derivation
    Description
    • Raw audio converted to mono, resampled to 16 kHz with Butterworth anti-aliasing filter
    • Spectrograms via STFT (25 ms window, 10 ms hop, 512-point FFT)
    • Acoustic features via OpenSMILE
    • Phonetic/prosodic features via Parselmouth and Praat (F0, formants, voice quality)
    • Transcriptions generated using OpenAI Whisper Large (free speech transcripts not included in v1.0)
    Used Software
    Name
    OpenSMILE
    Parselmouth
    Praat
    OpenAI Whisper Large
    torchaudio
    b2aiprep
  • Name
    Transcription and derived annotations
    Description
    • Automatic transcription using OpenAI Whisper Large for certain tasks; free speech transcripts were removed prior to release
    Used Software
    • Name
      OpenAI Whisper Large
  • Name
    Raw audio availability
    Description
    • Raw audio waveforms were not distributed in v1.0; only derived spectrograms and features are provided. Future releases aim to include audio with additional safeguards.
External ResourcesName
https://docs.b2ai-voice.orgDocumentation website
https://doi.org/10.5281/zenodo.14148755Bridge2AI Voice REDCap (v3.20.0)
  • Name
    Confidentiality considerations
    Description
    • Contains clinical and questionnaire data linked to participants; v1.0 includes only low-risk derived data and excludes raw audio
  • Name
    Sensitive data elements
    Description
    • Health-related data (demographics, clinical information, validated questionnaires)
  • Name
    De-identification
    Description
    • HIPAA Safe Harbor identifiers removed (e.g., names, detailed dates, contact identifiers, IDs)
    • State/province removed; country of data collection retained
    • Free speech transcripts removed
    • Raw audio waveforms omitted from this release
  • Name
    Health Data Nexus
    Description
    • Supported by the Temerty Centre for AI Research and Education in Medicine (Temerty Foundation)
Mixed (tabular phenotype/features and array-based parquet spectrograms)
🚀

Uses

What (other) tasks could the dataset be used for?

  • Name
    Intended tasks
    Response
    Develop and evaluate AI/ML methods for detecting, characterizing, and monitoring health conditions from voice-derived representations and associated clinical data.
Name
Access, license, and terms
Description
  • License
    Bridge2AI Voice Registered Access License
  • Data Use Agreement
    Bridge2AI Voice Registered Access Agreement
  • Access
    Credentialed users only; must sign DUA
  • Required Training
    TCPS 2: CORE 2022
  • This is a restricted-access resource distributed via Health Data Nexus
📤

Distribution

How will the dataset be distributed?

Name
Versioning and access to latest
Description
🔄

Maintenance

How will the dataset be maintained?

1.0
Name
Update plans
Description
  • Future releases aim to include voice waveforms (raw audio) with additional precautions to ensure data security
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema