healthnexus tab conflicts-of-interest d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • Name
    Purpose
    Response
    Create an ethically sourced flagship dataset to enable AI research on the use of voice as a biomarker of health by linking diverse voice recordings with clinical and demographic information.
GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease3OT2OD032720-01S1
📊

Composition

What do the instances represent?

  1. Name
    Instance class
    Representation
    Voice recordings linked to participant-level clinical, demographic, and questionnaire data; derived spectrogram matrices and acoustic/phonetic/prosodic features per recording.
    Instance Type
    Participants, recording sessions, and recordings (per task); derived data instances per recording.
    Data Type
    Derived features (e.g., spectrograms, acoustic, phonetic, prosodic) from standardized audio; raw audio waveforms are withheld in v1.0.
    Counts
    12,523
  • Name
    Adult cohort
    Identification
    • Adult participants meeting inclusion criteria within predefined disease cohorts
    Distribution
    • 306 participants; 12,523 recordings across multiple recording tasks
  • Description
    • Parquet (.parquet) for spectrograms
    • Tab-delimited text (.tsv) for phenotype and static features
    • JSON (.json) data dictionaries
  • Description
    • 2024-11-27 (initial public release v1.0)
🔍

Collection Process

How was the data acquired?

bridge2ai-voice-v1.0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice v1.0 is an ethically sourced, multi-institutional voice dataset focused on the use of voice as a biomarker of health. The initial release provides 12,523 recordings for 306 adult participants collected across five sites in North America. Participants were selected from five predefined cohorts where voice/speech changes are associated with disease: voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, respiratory disorders, and pediatric voice/speech disorders (note: v1.0 includes adults only). This release contains low-risk derived data (e.g., spectrograms and acoustic/phonetic/prosodic features) and corresponding demographic, clinical, and validated questionnaire data; raw audio waveforms are not included in v1.0. Documentation: "https://docs.b2ai-voice.org"
2024-11-27
  • voice
  • bridge2ai
  • audio
  • health
  • biomarker
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  • Name
    AddressingGap
    Response
    Address the lack of large, high-quality, multi-institutional, diverse voice datasets linked to other health biomarkers and collected under standardized, ethically governed protocols.
  1. Name
    Cohort sampling strategy
    Is Sample
    • True
    Is Random
    • False
    Source Data
    • Patients at specialty clinics across five sites in North America
    Is Representative
    • False
    Why Not Representative
    • Participants were recruited based on membership in predefined disease cohorts (voice, neurological, mood/psychiatric, respiratory, pediatric), and v1.0 includes adult cohort only.
    Strategies
    • Purposive cohort-based sampling within specialty clinics
  • Description
    • Health-related clinical information and demographics are included; identifiers removed per HIPAA Safe Harbor.
  1. Description
    • Data were collected at specialty clinics using a standardized protocol capturing demographics, health questionnaires (including validated instruments), targeted confounder questions, disease-specific information, and voice tasks (e.g., sustained vowel).
    • Data reported by subjects (questionnaires) and directly observed (voice recordings); some data are indirectly derived (features, transcriptions).
    Was Directly Observed
    yes (voice recordings)
    Was Reported By Subjects
    yes (questionnaires)
    Was Inferred Derived
    yes (features, transcriptions)
  • Description
    • Custom tablet-based data collection application; headset used for recording when possible; data exported from REDCap using an open-source library developed by the team.
    Used Software
    Name
    Bridge2AI-Voice tablet data collection app
    REDCap
    b2aiprep
  • Description
    • Project investigators and clinical teams at five North American sites
  • Description
    • Data collection and sharing approved by the University of South Florida Institutional Review Board; submission to the University of Toronto Research Ethics Board noted.
  • Description
    • Raw audio converted to monaural, resampled to 16 kHz with a Butterworth anti-aliasing filter; spectrograms computed via STFT (25 ms window, 10 ms hop, 512-point FFT).
    • Acoustic features extracted with OpenSMILE; phonetic/prosodic features computed with Parselmouth and Praat; transcriptions generated using OpenAI Whisper Large.
    • Code to preprocess and merge source data provided via the b2aiprep library.
    Used Software
    Name
    openSMILE
    Parselmouth
    Praat
    torchaudio
    OpenAI Whisper Large
    b2aiprep
  • Description
    • HIPAA Safe Harbor identifiers removed; state/province removed; country retained.
    • Transcripts of free speech audio removed from the release.
    • Only derived data (e.g., spectrograms, features) included in v1.0; audio waveforms omitted.
  • Description
    • Automated transcriptions generated using OpenAI Whisper Large; transcripts of free speech audio not released in v1.0.
    Used Software
    • Name
      OpenAI Whisper Large
  • Description
    • Raw voice audio collected but withheld from v1.0 distribution; planned for future releases with additional safeguards.
  • Description
    • Health Data Nexus (Temerty Centre for AI Research and Education in Medicine)
Description
  • HIPAA Safe Harbor de-identification applied; removal of direct identifiers and fine-grained dates; state/province removed; country retained; free-speech transcripts removed; no raw audio in v1.0.
partial (tabular phenotype/features; array-based spectrograms)
  • External Resources
  • Archival
    • Versioned DOIs indicate archival/versioning support via persistent identifiers
  • Restrictions
    • Registered access with DUA and required training
DescriptionIDMedia TypeNamePathTitle
Parquet dataset containing 513 x N spectrogram matrices per recording, with participant_id, session_...bridge2ai-voice-v1.0-spectrogramsapplication/x-parquetspectrograms.parquetspectrograms.parquetDerived spectrograms
Tab-delimited file with one row per participant including demographics, acoustic confounders, and re...bridge2ai-voice-v1.0-phenotypetext/tab-separated-valuesphenotype.tsvphenotype.tsvParticipant phenotype and questionnaires
JSON data dictionary mapping column names to descriptions for phenotype.tsv.bridge2ai-voice-v1.0-phenotype-dictapplication/jsonphenotype.jsonphenotype.jsonPhenotype data dictionary
Tab-delimited file with one row per unique recording; includes features derived using openSMILE, Pra...bridge2ai-voice-v1.0-static-featurestext/tab-separated-valuesstatic_features.tsvstatic_features.tsvRecording-level static features
JSON data dictionary mapping feature column names to descriptions for static_features.tsv.bridge2ai-voice-v1.0-static-features-dictapplication/jsonstatic_features.jsonstatic_features.jsonStatic features data dictionary
🚀

Uses

What (other) tasks could the dataset be used for?

  • Name
    Task
    Response
    Support AI-driven analysis of voice (e.g., feature extraction, modeling, and evaluation) for health-related research across disease cohorts where voice/speech changes are clinically relevant.
  • Description
    • Restricting v1.0 to low-risk derived data (no raw audio) mitigates privacy risks but may limit certain modeling tasks requiring waveforms; future releases aim to include audio with additional safeguards.
Name
Bridge2AI Voice Registered Access
Description
  • Access restricted to credentialed users who sign the Bridge2AI Voice Registered Access Agreement (DUA) under the Bridge2AI Voice Registered Access License.
  • Required Training
    TCPS 2: CORE 2022.
  • Access Platform
    Health Data Nexus (credentialed access).
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
Description
🔄

Maintenance

How will the dataset be maintained?

1.0
Description
  • Future releases planned to include raw voice data with additional security and privacy precautions; ongoing additions and corrections anticipated.
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema