=== YAML Fixing Applied ===
id: "doi:10.57764/qb6h-em84"
name: Bridge2AI-Voice
title: "Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information"
description: >-
  Bridge2AI-Voice is a comprehensive, ethically sourced dataset of derived data
  from human voice recordings linked with corresponding clinical information to
  enable research on voice as a biomarker of health. Version 1.0 includes
  12,523 recordings from 306 adult participants collected across five sites in
  North America. Participants were selected based on conditions known to
  manifest in the voice waveform, including voice disorders, neurological
  disorders, mood/psychiatric disorders, and respiratory disorders. This
  initial low-risk release includes derived spectrograms and acoustic,
  phonetic/prosodic features; original audio waveforms are not included.
language: en
page: "https://docs.b2ai-voice.org"
issued: 2024-11-27
doi: "doi:10.57764/qb6h-em84"
version: "1.0"
keywords:
  - voice
  - bridge2ai
  - audio
  - biomarker
  - health
license: Bridge2AI Voice Registered Access License
created_by:
  - Bridge2AI-Voice team
last_updated_on: 2024-11-27
was_derived_from: >-
  Clinical voice data collected via standardized protocols at five North
  American sites; source data managed via REDCap; preprocessing and merges via
  the open-source b2aiprep library.
purposes:
  - name: Primary Purpose
    response: >-
      Create a large, ethically sourced, diverse, multi-institutional dataset
      of voice-derived data linked to health information to enable AI research
      on voice as a biomarker of health.
tasks:
  - name: Intended Tasks
    response: >-
      AI/ML research using voice-derived features and spectrograms for health
      applications, including analysis associated with voice, neurological,
      mood/psychiatric, and respiratory disorders.
addressing_gaps:
  - name: Addressed Gap
    response: >-
      Lack of large, high-quality, standardized, multi-institutional and
      demographically diverse voice datasets linked to clinical information for
      reproducible AI research.
creators:
  - name: Author
    principal_investigator:
      name: Alistair Johnson
  - name: Author
    principal_investigator:
      name: Jean-Christophe Bélisle-Pipon
  - name: Author
    principal_investigator:
      name: David Dorr
  - name: Author
    principal_investigator:
      name: Satrajit Ghosh
  - name: Author
    principal_investigator:
      name: Philip Payne
  - name: Author
    principal_investigator:
      name: Maria Powell
  - name: Author
    principal_investigator:
      name: Anaïs Rameau
  - name: Author
    principal_investigator:
      name: Vardit Ravitsky
  - name: Author
    principal_investigator:
      name: Alexandros Sigaras
  - name: Author
    principal_investigator:
      name: Olivier Elemento
  - name: Author
    principal_investigator:
      name: Yael Bensoussan
funders:
  - name: NIH Bridge2AI Funding
    grantor:
      name: National Institutes of Health (NIH)
    grant:
      name: >-
        Bridge2AI: Voice as a Biomarker of Health - Building an ethically
        sourced, bioacoustic database to understand disease like never before
      grant_number: 3OT2OD032720-01S1
instances:
  - name: Dataset Instances
    representation: >-
      Derived data from human voice recordings (spectrograms and extracted
      acoustic/phonetic/prosodic features) linked to participant- and
      session-level clinical and questionnaire data.
    instance_type: >-
      Participants (adult cohort), recording sessions, and per-recording
      derived artifacts (spectrograms, feature vectors).
    data_type: >-
      Derived spectrograms (513×N), acoustic features (e.g., OpenSMILE),
      phonetic/prosodic features (Parselmouth/Praat), torchaudio-derived
      features; structured phenotype data (TSV/JSON data dictionaries).
    counts: 12523
    label: >-
      Participant clinical phenotypes and cohort categories (e.g., voice,
      neurological, mood/psychiatric, respiratory, pediatric; adult cohort
      included in v1.0).
    sampling_strategies:
      - name: Clinical Cohort Sampling
        is_sample:
          - yes
        is_random:
          - no
        source_data:
          - Patients at specialty clinics across five North American sites
        is_representative:
          - no
        why_not_representative:
          - Clinical cohort-based selection focused on conditions affecting voice
        strategies:
          - Eligibility-screened enrollment into predefined disorder cohorts
    missing_information:
      - name: Known Removals
        missing:
          - Original audio waveforms
          - Transcripts of free speech audio
          - HIPAA Safe Harbor identifiers (e.g., names; detailed dates; contact identifiers)
          - State/province fields
        why_missing:
          - Removed for de-identification and privacy protection in v1.0
relationships:
  - name: Participant-Session Linkage
    description:
      - Each spectrogram/feature record links to participant_id and session_id with task_name
splits:
  - name: Release Scope
    description:
      - v1.0 includes adult cohort only; pediatric cohort not included in this release
external_resources:
  - name: Project Resources
    external_resources:
      - Documentation site: https://docs.b2ai-voice.org
      - b2aiprep preprocessing library (open source)
      - Referenced toolkits: OpenSMILE, Praat, Parselmouth, torchaudio, OpenAI Whisper
    archival:
      - Versioned DOIs are provided for dataset releases
confidential_elements:
  - name: Clinical Context
    description:
      - Dataset includes clinical and questionnaire data linked to participants (de-identified)
content_warnings:
  - name: Content Notes
    warnings:
      - None noted
subpopulations:
  - name: Disorder Cohorts
    identification:
      - Voice disorders
      - Neurological and neurodegenerative disorders
      - Mood and psychiatric disorders
      - Respiratory disorders
      - Pediatric voice and speech disorders (cohort identified; not included in v1.0)
    distribution:
      - v1.0 includes 306 adult participants; pediatric data not included in this release
sensitive_elements:
  - name: Health-Related Data
    description:
      - Demographics, clinical information, and validated questionnaires (de-identified)
acquisition_methods:
  - name: Data Acquisition
    description:
      - Direct audio recordings using standardized tasks (e.g., sustained phonation) and clinical questionnaires
      - Derived features computed from standardized preprocessed audio
    was_directly_observed: yes
    was_reported_by_subjects: yes
    was_inferred_derived: yes
    was_validated_verified: >-
      Standardized protocol used; de-identification applied; toolchains and procedures referenced
collection_mechanisms:
  - name: Collection Protocols and Tools
    description:
      - Standardized multi-site protocol
      - Custom tablet application for data capture
      - Headset microphone used when possible
      - Source data exported and converted from REDCap using team-developed library
data_collectors:
  - name: Collection Personnel
    description:
      - Project investigators and clinical teams at specialty clinics across five North American sites
collection_timeframes: []
ethical_reviews:
  - name: Ethical Oversight
    description:
      - Data collection and sharing approved by University of South Florida IRB
      - Submission for review to University of Toronto Research Ethics Board
data_protection_impacts:
  - name: Data Protection and Risk
    description:
      - HIPAA Safe Harbor de-identification
      - Removal of original audio and free-speech transcripts reduces re-identification risk
preprocessing_strategies:
  - name: Audio Preprocessing
    description:
      - Convert to monaural
      - Resample to 16 kHz
      - Apply Butterworth anti-aliasing filter
      - Compute spectrograms via STFT (25 ms window, 10 ms hop, 512-point FFT; 513×N magnitude)
    used_software:
      - name: torchaudio
        version: "2.1"
      - name: b2aiprep
      - name: OpenSMILE
      - name: Parselmouth
      - name: Praat
cleaning_strategies:
  - name: De-identification Removals
    description:
      - Removal of identifiers per HIPAA Safe Harbor
      - Removal of state/province; retention of country only
      - Omission of original audio waveforms and free-speech transcripts in v1.0
labeling_strategies:
  - name: Transcription and Feature Derivation
    description:
      - Transcriptions generated using OpenAI Whisper Large for applicable tasks
      - Acoustic and prosodic feature extraction via OpenSMILE, Parselmouth/Praat
      - Note: Transcripts of free speech audio removed from release
    used_software:
      - name: OpenAI Whisper Large
      - name: OpenSMILE
      - name: Parselmouth
      - name: Praat
raw_sources:
  - name: Original Audio
    description:
      - Raw audio recordings were collected but are not included in v1.0 release; only derived data provided
existing_uses: []
use_repository: []
other_tasks:
  - name: Potential Applications
    description:
      - Biomarker development, disease screening support, therapeutic monitoring research using voice-derived features
future_use_impacts:
  - name: Use Considerations
    description:
      - Exclusion of raw audio and removal of free speech transcripts may limit certain acoustic analyses and model types
      - Clinical cohort sampling could introduce distributional biases; users should assess generalizability
discouraged_uses:
  - name: Cautions
    description:
      - Re-identification attempts or misuse to infer sensitive attributes beyond de-identified scope
distribution_formats:
  - name: Release File Formats
    description:
      - Parquet (spectrograms.parquet)
      - TSV (phenotype.tsv, static_features.tsv)
      - JSON (phenotype.json, static_features.json)
distribution_dates:
  - name: v1.0 Release
    description:
      - 2024-11-27
license_and_use_terms:
  name: Access and Terms
  description:
    - Only credentialed users may access the files
    - Required training: TCPS 2: CORE 2022
    - Must sign the Bridge2AI Voice Registered Access Agreement (DUA)
    - License: Bridge2AI Voice Registered Access License
ip_restrictions: {}
regulatory_restrictions: {}
maintainers:
  - name: Dataset Maintenance
    description:
      - Bridge2AI-Voice team; hosted via Health Data Nexus
errata: []
updates:
  name: Update Plan
  description:
    - Future releases may include voice audio waveforms with additional safeguards
    - Latest version DOI: https://doi.org/10.57764/3sg0-7440
retention_limit: {}
version_access:
  name: Versioning and Access
  description:
    - Versioned DOIs provided; v1.0 doi:10.57764/qb6h-em84, latest doi:10.57764/3sg0-7440
extension_mechanism: {}
is_deidentified:
  name: De-identification Summary
  description:
    - HIPAA Safe Harbor identifiers removed
    - State/province removed; country retained
    - Free-speech transcripts removed
    - Original audio waveforms omitted in v1.0
is_tabular: mixed (Parquet, TSV, JSON)
distribution_formats:
  - name: File Types
    description:
      - Parquet
      - TSV
      - JSON