healthnexus tab abstract d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • Name
    Primary purpose
    Used Software
    Attributes
    Response
    Create an ethically sourced flagship dataset to enable AI research on voice as a biomarker of health and support clinically meaningful insights.
GrantorGrant NameGrant Number
---
📊

Composition

What do the instances represent?

  • Name
    Recording-level instances
    Used Software
    Attributes
    Representation
    Voice recordings (derived representations) linked to participant clinical and questionnaire data
    Instance Type
    Participants and their associated recording sessions/tasks
    Data Type
    Derived spectrograms (513×N), acoustic features, phonetic and prosodic features, and metadata (demographics, questionnaires). Original audio waveforms excluded in v1.0.
    Counts
    12,523
    Label
    Cohort membership and clinical variables available; no explicit single target label across all instances.
    Sampling Strategies
    Missing Information
  • Name
    Adult cohort and disease categories
    Used Software
    Attributes
    Identification
    • Adult participants only in v1.0
    • Disease Cohorts
      Voice disorders; Neurological and neurodegenerative disorders; Mood and psychiatric disorders; Respiratory disorders
    Distribution
    • 306 adult participants across five sites (detailed distribution not provided)
  • Name
    File formats
    Used Software
    Attributes
    Description
    • Parquet (.parquet)
    • Tab-separated values (.tsv)
    • JSON (.json)
  • Name
    Release date
    Used Software
    Attributes
    Description
    • 2024-11-27 (v1.0)
🔍

Collection Process

How was the data acquired?

bridge2ai-voice-v1-0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically sourced dataset for voice as a biomarker of health. Version 1.0 provides 12,523 recordings for 306 adult participants collected across five North American sites, with corresponding demographic, clinical, and validated questionnaire information. Participants were selected from cohorts with conditions known to manifest in the voice (voice disorders, neurological disorders, mood/psychiatric disorders, respiratory disorders; pediatric cohort planned for future releases). In v1.0, low-risk derived data (e.g., spectrograms, acoustic, phonetic, and prosodic features) are released; original audio waveforms are not included. Data are de-identified under HIPAA Safe Harbor (with additional removals such as state/province and free-speech transcripts). Documentation: https://docs.b2ai-voice.org
English
2024-11-27
  • voice
  • bridge2ai
  • audio
  • biomarker
  • health
  • spectrograms
  • parquet
  • tsv
  • json
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  • Name
    Gap addressed
    Used Software
    Attributes
    Response
    Lack of large, diverse, multi-institutional, ethically sourced voice datasets linked to clinical information for AI research.
  • Name
    Cohort-based clinical sampling
    Used Software
    Attributes
    Is Sample
    • True
    Is Random
    • False
    Source Data
    • Patients presenting at specialty clinics across five North American sites
    Is Representative
    • Not designed to be population-representative
    Representative Verification
    Why Not Representative
    • Purposeful enrollment based on predefined disease cohorts
    Strategies
    • Purposeful cohort-based sampling aligned to predefined conditions
  • Name
    Documentation and related resources
    Used Software
    Attributes
    External Resources
    Future Guarantees
    • Versioned DOIs provided (latest version DOI available)
    Archival
    • Versioned DOI records (v1.0 and latest-version DOI)
    Restrictions
    • Registered/credentialed access under DUA and training requirements
  • Name
    Clinical information
    Used Software
    Attributes
    Description
    • Contains clinical/demographic/questionnaire data considered low risk, distributed under registered access and DUA
Name
De-identification
Used Software
Attributes
Description
  • HIPAA Safe Harbor identifiers removed
  • State and province removed; country retained
  • Free-speech transcripts removed
  • Audio waveforms omitted in v1.0; only derived/low-risk data released
  • Name
    Health-related and biometric-adjacent data
    Used Software
    Attributes
    Description
    • Contains health information and derived data from voice (audio waveforms excluded in v1.0)
  • Name
    Data acquisition
    Used Software
    Attributes
    Description
    • Directly observed audio tasks (e.g., sustained phonation); clinical and questionnaire data collected via custom tablet application
    Was Directly Observed
    True
    Was Reported By Subjects
    True
    Was Inferred Derived
    True
    Was Validated Verified
    Standardized protocol; derived features computed from standardized audio; IRB approval
  • Name
    Collection mechanisms
    Used Software
    Attributes
    Description
    • Custom tablet-based application; headset microphone when possible; standardized multi-site protocol
  • Name
    Data collection teams
    Used Software
    Attributes
    Description
    • Project investigators at specialty clinics across five North American sites
  • Name
    Session structure
    Used Software
    Attributes
    Description
    • Most participants completed data collection in a single session; a subset required multiple sessions
  • Name
    Ethics approvals
    Used Software
    Attributes
    Description
    • Data collection and sharing approved by the University of South Florida IRB; submitted to the University of Toronto Research Ethics Board
  • Name
    Audio standardization and feature derivation
    Used Software
    IDNameURLVersion
    openSMILEopenSMILEhttps://audeering.github.io/opensmile/
    ParselmouthParselmouthhttps://github.com/YannickJadoul/Parselmouth
    PraatPraathttp://www.praat.org/
    Whisper-LargeOpenAI Whisper Largehttps://github.com/openai/whisper
    torchaudioTorchAudiohttps://pytorch.org/audio2.1
    librosalibrosahttps://librosa.org
    b2aiprepb2aiprephttps://github.com/sensein/b2aiprep
    Attributes
    Description
    • Monaural conversion; resampling to 16 kHz with Butterworth anti-aliasing filter
    • Spectrograms via STFT (25 ms window, 10 ms hop, 512-point FFT)
    • Acoustic features via openSMILE
    • Phonetic/prosodic features via Parselmouth and Praat
    • Transcriptions via OpenAI Whisper Large
  • Name
    Transcription
    Used Software
    Attributes
    Description
    • Automatic transcription for certain tasks; free-speech transcripts removed for de-identification
  • Name
    Raw audio availability
    Used Software
    Attributes
    Description
    • Original audio waveforms are not included in v1.0; planned for future releases with additional safeguards
Name
Third-party distribution
Description
Yes. Distributed to credentialed users outside the hosting entity under registered access controls and a DUA.
  • Name
    Hosting and support
    Used Software
    Attributes
    Description
    • Health Data Nexus
    • Temerty Centre for AI Research and Education in Medicine (Temerty Foundation-supported)
Partially (tabular TSV/JSON metadata and features; spectrograms stored in Parquet)
DescriptionIDMedia TypeNamePathTitle
Parquet dataset with spectrograms (513×N) and identifiers (participant_id, session_id, task_name).spectrograms.parquetapplication/x-parquetspectrograms.parquetspectrograms.parquetDerived spectrograms
Tab-delimited table of demographics, acoustic confounders, and validated questionnaire responses (on...phenotype.tsvtext/tab-separated-valuesphenotype.tsvphenotype.tsvParticipant-level phenotype data
JSON data dictionary describing columns in phenotype.tsv.phenotype.jsonapplication/jsonphenotype.jsonphenotype.jsonPhenotype data dictionary
Tab-delimited table of features derived from raw audio (one row per recording).static_features.tsvtext/tab-separated-valuesstatic_features.tsvstatic_features.tsvRecording-level static features
JSON data dictionary describing columns in static_features.tsv.static_features.jsonapplication/jsonstatic_features.jsonstatic_features.jsonStatic features data dictionary
🚀

Uses

What (other) tasks could the dataset be used for?

  • Name
    Intended tasks
    Used Software
    Attributes
    Response
    Voice-based biomarker discovery; disease state classification and screening; analysis of acoustic, phonetic, and prosodic features; model development and validation using derived speech representations.
  • Name
    First release status
    Used Software
    Attributes
    Description
    • v1.0 is the first public release; prior external uses not listed
  • Name
    Potential downstream tasks
    Used Software
    Attributes
    Description
    • Condition screening and monitoring using voice-derived features
    • Multimodal fusion with demographics/clinical variables
    • Robustness, fairness, and domain generalization studies for voice biomarkers
  • Name
    Use considerations
    Used Software
    Attributes
    Description
    • Cohort-based sampling may influence generalizability; users should account for cohort composition and site effects
Name
Access and use terms
Used Software
Attributes
Description
  • Access Policy
    only credentialed users who sign the DUA can access the files
  • License (files)
    Bridge2AI Voice Registered Access License
  • Data Use Agreement
    Bridge2AI Voice Registered Access Agreement
  • Required Training
    TCPS 2 - CORE 2022
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
Name
Version access
Used Software
Attributes
Description
  • Versioned records maintained via DOIs; latest-version DOI available for discovery
🔄

Maintenance

How will the dataset be maintained?

1.0
Name
Update plan
Used Software
Attributes
Description
Generated on 2025-11-09 10:17:35 using Bridge2AI Data Sheets Schema