=== YAML Fixing Applied ===
id: "doi:10.57764/qb6h-em84"
name: Bridge2AI-Voice v1.0
title: "Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information"
description: >-
  Bridge2AI-Voice is a comprehensive, ethically-sourced dataset of derived voice
  data linked to health information to enable research on voice as a biomarker of
  health. Version 1.0 contains 12,523 recordings from 306 participants collected
  across five North American sites. Participants were selected from disease cohorts
  with known voice-related manifestations, including voice disorders, neurological
  disorders, mood/psychiatric disorders, and respiratory disorders. This initial
  release includes low-risk, derived data (e.g., spectrograms and engineered features)
  and detailed demographic/clinical/questionnaire data; original audio waveforms and
  free-speech transcripts are not included.
language: en
page: "https://docs.b2ai-voice.org"
issued: "2024-11-27"
version: "1.0"
created_by:
  - Bridge2AI-Voice team
  - Health Data Nexus
last_updated_on: "2024-11-27"
doi: "doi:10.57764/qb6h-em84"
keywords:
  - voice
  - bridge2ai
  - audio
license: Bridge2AI Voice Registered Access License
is_tabular: "true"
purposes:
  - id: purpose-1
    name: Purpose
    description: Enable AI research on the human voice as a biomarker of health.
    used_software: []
    response: >-
      Create an ethically sourced flagship dataset to support research on voice as
      a biomarker and enable development and evaluation of AI methods using diverse,
      clinically linked voice-derived data.
tasks:
  - id: task-1
    name: Intended tasks
    description: Downstream AI tasks using derived voice data.
    used_software: []
    response: >-
      Voice-based biomarker discovery and model development (e.g., classification/
      regression for health-related conditions), acoustic feature analysis, and
      benchmarking using spectrograms and engineered features.
addressing_gaps:
  - id: gap-1
    name: Addressing gaps
    description: Gaps addressed by the dataset.
    used_software: []
    response: >-
      Addresses the lack of large, high-quality, multi-institutional and diverse
      voice datasets linked to clinical information, collected under standardized
      protocols with ethical oversight.
funders:
  - id: funder-nih-b2ai-voice
    name: Funding
    description: NIH funding support for Bridge2AI-Voice.
    used_software: []
    grantor:
      id: nih
      name: National Institutes of Health (NIH)
    grant:
      id: 3OT2OD032720-01S1
      name: "Bridge2AI: Voice as a Biomarker of Health"
      grant_number: 3OT2OD032720-01S1
instances:
  - id: instance-recordings
    name: Recording-derived instances
    description: Instances representing per-recording derived data.
    used_software: []
    representation: Derived voice recordings (e.g., spectrograms and engineered features)
    instance_type: Recording-level instances
    data_type: Spectrogram matrices and engineered acoustic/phonetic/prosodic features
    counts: 12523
    label: None specified
    sampling_strategies: []
    missing_information: []
  - id: instance-participants
    name: Participant instances
    description: Instances representing per-participant clinical/phenotype data.
    used_software: []
    representation: Participant-level phenotype/demographics/questionnaires
    instance_type: Participant-level instances
    data_type: Tabular phenotype and questionnaire responses
    counts: 306
    label: None specified
    sampling_strategies: []
    missing_information: []
sampling_strategies:
  - id: sampling-1
    name: Cohort sampling
    description: Clinic-based enrollment by predefined disease cohorts.
    used_software: []
    is_sample:
      - "yes"
    is_random:
      - "no"
    source_data:
      - Specialty clinics and institutions across five North American sites
    is_representative:
      - "no"
    representative_verification:
      - Not applicable; cohort-based enrollment
    why_not_representative:
      - Purposeful sampling of disease cohorts rather than population sampling
    strategies:
      - Cohort-based enrollment by inclusion/exclusion criteria and standardized protocol
relationships:
  - id: rel-1
    name: Relationships
    description:
      - Recording sessions are linked to participants (participant_id, session_id) and tasks (task_name).
    used_software: []
splits:
  - id: splits-1
    name: Data splits
    description:
      - No official train/dev/test splits specified in v1.0.
    used_software: []
anomalies:
  - id: anomalies-1
    name: Data anomalies
    description:
      - None reported in v1.0.
    used_software: []
external_resources:
  - id: ext-1
    name: External resources
    description: External links and resources related to the dataset.
    used_software: []
    external_resources:
      - Project documentation: https://docs.b2ai-voice.org
      - Processing code (b2aiprep): https://github.com/sensein/b2aiprep
    future_guarantees:
      - Not specified
    archival:
      - Versioned DOIs are provided (v1.0 and latest version DOIs).
    restrictions:
      - Access restricted to credentialed users with DUA and required training
confidential_elements:
  - id: confidential-1
    name: Confidentiality
    description:
      - Dataset excludes raw audio and removes identifiers following HIPAA Safe Harbor; only low-risk derived data are released in v1.0.
    used_software: []
content_warnings:
  - id: cw-1
    name: Content warnings
    warnings:
      - None noted; sensitive audio content is not included in this release.
    used_software: []
subpopulations:
  - id: subpop-1
    name: Subpopulations
    description: Adult cohort in v1.0; pediatric cohort identified but not included in this release.
    used_software: []
    identification:
      - Adult participants; disease cohorts (voice, neurological, mood/psychiatric, respiratory)
    distribution:
      - Detailed distributions not provided in the release notes
sensitive_elements:
  - id: sensitive-1
    name: Sensitive elements
    description:
      - Health-related clinical and questionnaire data (demographics, validated instruments) linked to recordings.
    used_software: []
acquisition_methods:
  - id: acquisition-1
    name: Instance acquisition
    description:
      - Voice data directly recorded in clinic using a standardized protocol; linked clinical and questionnaire data collected via custom app.
    used_software: []
    was_directly_observed: "yes"
    was_reported_by_subjects: "yes (questionnaire responses)"
    was_inferred_derived: "yes (acoustic/phonetic/prosodic features; ASR transcriptions)"
    was_validated_verified: "yes (standardized protocol; institutional oversight)"
collection_mechanisms:
  - id: mech-1
    name: Collection mechanisms
    description:
      - Custom tablet application for data collection; headset used when possible; standardized clinical protocol across sites.
    used_software: []
data_collectors:
  - id: collectors-1
    name: Data collectors
    description:
      - Project investigators and clinical staff at specialty clinics/institutions across five North American sites.
    used_software: []
collection_timeframes:
  - id: timeframe-1
    name: Collection timeframe
    description:
      - Specific dates not provided; multiple sessions possible for some participants within the study period prior to v1.0 release.
    used_software: []
ethical_reviews:
  - id: ethics-1
    name: Ethical review
    description:
      - Approved by the University of South Florida IRB; submitted for review to the University of Toronto Research Ethics Board.
    used_software: []
data_protection_impacts: []
preprocessing_strategies:
  - id: prep-1
    name: Audio preprocessing and feature extraction
    description:
      - Raw audio converted to mono and resampled to 16 kHz with Butterworth anti-aliasing.
      - Spectrograms computed via STFT (25 ms window, 10 ms hop, 512-point FFT).
      - Acoustic features extracted with OpenSMILE.
      - Phonetic/prosodic features computed with Praat/Parselmouth (e.g., F0, formants, voice quality).
      - ASR transcriptions generated with OpenAI Whisper Large (free-speech transcripts not released).
    used_software:
      - id: sw-opensmile
        name: openSMILE
        url: "https://audeering.github.io/opensmile/"
      - id: sw-praat
        name: Praat
        url: "https://www.fon.hum.uva.nl/praat/"
      - id: sw-parselmouth
        name: Parselmouth
        url: "https://parselmouth.readthedocs.io/"
      - id: sw-torchaudio
        name: TorchAudio
        url: "https://pytorch.org/audio"
      - id: sw-whisper
        name: OpenAI Whisper (Large)
        url: "https://github.com/openai/whisper"
cleaning_strategies:
  - id: clean-1
    name: De-identification and redactions
    description:
      - HIPAA Safe Harbor identifiers removed (e.g., names, detailed dates, contact numbers, SSNs, MRNs, etc.).
      - State/province removed; country of data collection retained.
      - Free-speech transcripts removed.
      - Raw audio waveforms omitted from v1.0 release.
    used_software: []
labeling_strategies:
  - id: label-1
    name: Labeling/annotation
    description:
      - Automatic speech recognition transcriptions generated with Whisper Large; free-speech transcripts not included in release.
    used_software:
      - id: sw-whisper
        name: OpenAI Whisper (Large)
        url: "https://github.com/openai/whisper"
raw_sources:
  - id: raw-1
    name: Raw data availability
    description:
      - Original audio waveforms are not distributed in v1.0; planned for future releases subject to additional safeguards.
    used_software: []
existing_uses:
  - id: uses-1
    name: Existing uses
    description:
      - This is the first public release (v1.0); prior external uses not listed.
    used_software: []
use_repository: []
other_tasks:
  - id: othertasks-1
    name: Potential additional tasks
    description:
      - Acoustic biomarker discovery; robustness and bias assessment; feature benchmarking; cohort comparison studies; task development without raw audio.
    used_software: []
future_use_impacts:
  - id: futureimpact-1
    name: Future use considerations
    description:
      - Cohort-based sampling may limit population representativeness; derived-only release mitigates privacy risks but may limit tasks requiring raw audio.
      - Users should account for cohort selection and site effects when modeling; adhere to ethical and privacy constraints.
    used_software: []
discouraged_uses: []
distribution_formats:
  - id: distfmt-1
    name: Distribution formats
    description:
      - Credentialed access via Health Data Nexus; files provided as Parquet (spectrograms.parquet), TSV (phenotype.tsv, static_features.tsv), and JSON (phenotype.json, static_features.json).
    used_software: []
distribution_dates:
  - id: distdate-1
    name: Distribution date
    description:
      - 2024-11-27 (v1.0 release)
    used_software: []
license_and_use_terms:
  id: lut-1
  name: Access, license, and use terms
  description:
    - Access Policy: Only credentialed users who sign the DUA can access the files.
    - License: Bridge2AI Voice Registered Access License.
    - Data Use Agreement: Bridge2AI Voice Registered Access Agreement.
    - Required training: TCPS 2: CORE 2022.
  used_software: []
ip_restrictions: []
regulatory_restrictions: []
maintainers:
  - id: maint-1
    name: Maintainers
    description:
      - Health Data Nexus; Bridge2AI-Voice team.
    used_software: []
errata: []
updates:
  id: updates-1
  name: Update plan
  description:
    - Future releases aim to include voice audio with additional data security precautions.
    - Latest version DOI provided for continued versioning and access.
  used_software: []
retention_limit:
  id: retention-1
  name: Retention limits
  description:
    - Not specified.
  used_software: []
version_access:
  id: version-1
  name: Version access
  description:
    - Latest version DOI: https://doi.org/10.57764/3sg0-7440; versioned releases will be maintained via DOI links.
  used_software: []
extension_mechanism:
  id: extension-1
  name: Extension/contribution mechanisms
  description:
    - Not specified; code for preprocessing released as open source (b2aiprep).
  used_software:
    - id: sw-b2aiprep
      name: b2aiprep
      url: "https://github.com/sensein/b2aiprep"
is_deidentified:
  id: deid-1
  name: De-identification
  description:
    - HIPAA Safe Harbor de-identification; removal of state/province; retention of country only; exclusion of raw audio and free-speech transcripts in v1.0.
  used_software: []
subsets:
  - id: file-spectrograms-parquet
    name: spectrograms.parquet
    title: Spectrograms (derived from audio)
    description: Parquet dataset of spectrogram matrices per recording with participant_id, session_id, and task_name.
    media_type: application/parquet
    path: spectrograms.parquet
    bytes:
    format:
    encoding:
    purposes: []
    tasks: []
    addressing_gaps: []
    creators: []
    funders: []
    subsets: []
    instances: []
    anomalies: []
    external_resources: []
    confidential_elements: []
    content_warnings: []
    subpopulations: []
    sensitive_elements: []
    acquisition_methods: []
    collection_mechanisms: []
    sampling_strategies: []
    data_collectors: []
    collection_timeframes: []
    ethical_reviews: []
    data_protection_impacts: []
    preprocessing_strategies: []
    cleaning_strategies: []
    labeling_strategies: []
    raw_sources: []
    existing_uses: []
    use_repository: []
    other_tasks: []
    future_use_impacts: []
    discouraged_uses: []
    distribution_formats: []
    distribution_dates: []
    license_and_use_terms:
    ip_restrictions:
    regulatory_restrictions:
    maintainers: []
    errata: []
    updates:
    retention_limit:
    version_access:
    extension_mechanism:
    is_deidentified:
  - id: file-phenotype-tsv
    name: phenotype.tsv
    title: Participant phenotype and questionnaire data
    description: Tab-delimited participant-level demographics, acoustic confounders, and validated questionnaire responses (one row per participant).
    media_type: text/tab-separated-values
    path: phenotype.tsv
  - id: file-phenotype-json
    name: phenotype.json
    title: Data dictionary for phenotype.tsv
    description: JSON data dictionary describing columns in phenotype.tsv.
    media_type: application/json
    path: phenotype.json
    format: JSON
  - id: file-static-features-tsv
    name: static_features.tsv
    title: Engineered acoustic/phonetic/prosodic features
    description: Tab-delimited recording-level features (one row per recording) derived with openSMILE, Praat/Parselmouth, and torchaudio.
    media_type: text/tab-separated-values
    path: static_features.tsv
  - id: file-static-features-json
    name: static_features.json
    title: Data dictionary for static_features.tsv
    description: JSON data dictionary describing engineered feature columns.
    media_type: application/json
    path: static_features.json
    format: JSON