healthnexus tab documentation d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • ID
    purpose-1
    Name
    Purpose
    Response
    Create an ethically sourced, diverse voice dataset linked to health information to enable AI research and evaluate voice as a biomarker of health across multiple clinical conditions.
GrantorGrant NameGrant Number
National Institutes of HealthBridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database3OT2OD032720-01S1
📊

Composition

What do the instances represent?

  1. ID
    instances-voice-derived
    Name
    Instance
    Representation
    Voice recordings and corresponding derived data per session/participant
    Instance Type
    Participants, recording sessions, and task-specific recordings (e.g., sustained phonation); adult cohort in v1.0
    Data Type
    Derived data only in v1.0: spectrograms (513 x N), acoustic, phonetic/prosodic features, and automatic transcriptions; demographic, clinical, and validated questionnaire responses
    Counts
    12,523
    Label
    Clinical cohort membership by disease category (voice, neurological, mood/psychiatric, respiratory); detailed labels vary by disease-specific questionnaires and phenotype fields
    Sampling Strategies
    1. ID
      sampling-1
      Name
      SamplingStrategy
      Is Sample
      • Yes; participants recruited from specialty clinics at five North American sites
      Is Random
      • No; condition-based recruitment per predefined cohorts
      Source Data
      • Patients presenting at participating specialty clinics
      Is Representative
      • Not claimed; targeted clinical cohorts
      Why Not Representative
      • Targeted enrollment to capture voice-manifesting conditions
      Strategies
      • Deterministic cohort-based inclusion with screening per inclusion/exclusion criteria
    Missing Information
    1. ID
      missing-1
      Name
      MissingInfo
      Missing
      • Some participants may have multiple sessions; completeness may vary by session
      Why Missing
      • Operational needs; subset required multiple visits to complete collection
  1. ID
    subpop-1
    Name
    Subpopulation
    Identification
    • Adult participants (v1.0 release)
    • Disease Cohorts
      voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, respiratory disorders
    Distribution
    • 306 participants; 12,523 recordings; across five North American sites
  • ID
    distfmt-1
    Name
    DistributionFormat
    Description
    • Credentialed access via Health Data Nexus
    • Distributed files in Parquet (spectrograms), TSV (phenotype and features), and JSON (data dictionaries)
  • ID
    distdate-1
    Name
    DistributionDate
    Description
    • 2024-11-27 (v1.0 release)
🔍

Collection Process

How was the data acquired?

bridge2ai-voice-v1-0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information (v1.0)
Bridge2AI-Voice is a comprehensive collection of data derived from voice recordings with corresponding clinical information to enable AI research into voice as a biomarker of health. Version 1.0 provides 12,523 recordings from 306 adult participants collected across five North American sites, selected based on conditions that manifest in voice (voice disorders, neurological disorders, mood disorders, respiratory disorders, pediatric cohort planned for future releases). The initial release contains low-risk derived data (e.g., spectrograms and engineered features) and detailed demographic, clinical, and validated questionnaire data; original audio waveforms are omitted in v1.0.
English
2024-11-27
  • voice
  • bridge2ai
  • audio
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
bibo:published
Raw voice recordings collected during standardized clinical sessions; v1.0 distributes only derived data (spectrograms, engineered features, and data dictionaries) with original audio omitted.
  • ID
    gap-1
    Name
    AddressingGap
    Response
    Provide a large, high-quality, multi-institutional, demographically diverse voice dataset with linked clinical information and standardized collection protocols, addressing prior limitations of small, non-diverse datasets and heterogeneous protocols.
  • ID
    rel-1
    Name
    Relationships
    Description
    • Each record links participant_id to session_id and task_name; one row per recording for features; one row per participant for phenotype
ID
deid-1
Name
Deidentification
Description
  • HIPAA Safe Harbor identifiers removed (e.g., names, fine-grained dates, contact and device identifiers, biometrics)
  • State and province removed; country retained
  • Free-speech transcripts removed
  • Original audio waveforms omitted from v1.0 distribution
  • ID
    sens-1
    Name
    SensitiveElement
    Description
    • Contains health-related demographic, clinical, and validated questionnaire data (de-identified)
  1. ID
    acq-1
    Name
    InstanceAcquisition
    Description
    • Standardized clinic-based protocol with demographic and clinical questionnaires and task-based voice recordings
    • Tasks include sustained phonation and other voice/speech tasks
    Was Directly Observed
    Yes (voice recordings, tasks performed in clinic)
    Was Reported By Subjects
    Yes (validated questionnaires and targeted confounder questions)
    Was Inferred Derived
    Yes (spectrograms, engineered acoustic/phonetic/prosodic features, and ASR transcriptions)
    Was Validated Verified
    Validated questionnaires used; standardized multi-site protocol
  • ID
    mech-1
    Name
    CollectionMechanism
    Description
    • Custom tablet application with headset for data collection when possible
    • REDCap used for data capture; export and conversion using open-source library (b2aiprep)
  • ID
    collectors-1
    Name
    DataCollector
    Description
    • Project investigators at participating specialty clinics across five North American sites
  • ID
    irb-1
    Name
    EthicalReview
    Description
    • Data collection and sharing approved by University of South Florida Institutional Review Board
    • Submitted for review to the University of Toronto Research Ethics Board
  1. ID
    prep-1
    Name
    PreprocessingStrategy
    Description
    • Raw audio converted to mono
    • Resampled to 16 kHz with Butterworth anti-aliasing filter
    • Spectrograms via STFT (25 ms window, 10 ms hop, 512-point FFT)
    • Acoustic features via openSMILE
    • Phonetic/prosodic features via Parselmouth and Praat
    • Additional features via torchaudio
    Used Software
    IDNameURL
    software-opensmileopenSMILEhttps://audeering.github.io/opensmile/
    software-parselmouthParselmouthhttps://github.com/YannickJadoul/Parselmouth
    software-praatPraathttps://www.fon.hum.uva.nl/praat/
    software-torchaudiotorchaudiohttps://pytorch.org/audio/stable/
  1. ID
    label-1
    Name
    LabelingStrategy
    Description
    • Automatic speech transcriptions generated using OpenAI Whisper Large
    Used Software
  • ID
    raw-1
    Name
    RawData
    Description
    • REDCap-based data capture (Bridge2AI Voice REDCap v3.20.0; Zenodo DOI 10.5281/zenodo.14148755); raw audio waveforms not distributed in v1.0
  • ID
    maint-1
    Name
    Maintainer
    Description
    • Hosted and made discoverable via Health Data Nexus; project documentation at https://docs.b2ai-voice.org
partially (mixture of tabular TSV/JSON and array-based Parquet)
DescriptionIDIs TabularKeywordsMedia TypeNamePathTitle
Parquet dataset containing time-frequency representations (spectrograms) for each recording with par...spectrograms-parquetFalsespectrogram, parquetapplication/x-parquetspectrograms.parquetspectrograms.parquetDerived Spectrograms
Tab-delimited file with one row per participant including demographics, acoustic confounders, and re...phenotype-tsvTruephenotype, demographics, questionnairestext/tab-separated-valuesphenotype.tsvphenotype.tsvPhenotype Table
JSON data dictionary describing columns in phenotype.tsv; each key maps to column metadata with a on...phenotype-jsonTruedata dictionary, phenotypeapplication/jsonphenotype.jsonphenotype.jsonPhenotype Data Dictionary
Tab-delimited file with one row per recording containing features derived from openSMILE, Praat, Par...static-features-tsvTruefeatures, acoustics, phoneticstext/tab-separated-valuesstatic_features.tsvstatic_features.tsvEngineered Acoustic/Phonetic Features
JSON data dictionary describing feature columns in static_features.tsv; each key maps to column meta...static-features-jsonTruedata dictionary, featuresapplication/jsonstatic_features.jsonstatic_features.jsonFeatures Data Dictionary
  1. ID
    subset-adult-v1-0
    Name
    Adult cohort (v1.0)
    Title
    Adult Cohort Subset
    Description
    Initial public release contains adult participants only; pediatric cohort planned for future releases.
    Is Data Split
    False
    Is Subpopulation
    True
🚀

Uses

What (other) tasks could the dataset be used for?

  • ID
    task-1
    Name
    Intended task
    Response
    Development and evaluation of AI methods for health-related voice biomarker discovery and analysis using derived acoustic, phonetic, prosodic, and transcriptional features.
ID
terms-1
Name
LicenseAndUseTerms
Description
  • Bridge2AI Voice Registered Access License
  • Bridge2AI Voice Registered Access Agreement (DUA)
  • Access Policy
    Only credentialed users who sign the DUA can access files
  • Required Training
    TCPS 2: CORE 2022
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
ID
version-access-1
Name
VersionAccess
Description
🔄

Maintenance

How will the dataset be maintained?

1.0
ID
updates-1
Name
UpdatePlan
Description
  • Future releases aim to include original voice audio with additional safeguards
  • Versioned DOIs provided; latest version DOI available
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema