healthnexus tab usage-notes d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  1. ID
    purpose:voice-biomarker-research
    Name
    Voice as a biomarker of health
    Description
    Enable AI research into health-related acoustic markers using ethically sourced, diverse voice data linked to clinical information.
    Used Software
    Response
    Enable future research in artificial intelligence using voice as a biomarker of health.
GrantorGrant NameGrant Number
NIHBridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before.3OT2OD032720-01S1
📊

Composition

What do the instances represent?

  1. ID
    instance:recordings-and-derived-data
    Name
    Voice recordings (derived) and participant data
    Description
    Instances represent derived data from voice recordings (e.g., spectrograms and static acoustic/phonetic/prosodic features) linked to participant-level phenotype and clinical questionnaire data.
    Used Software
    Representation
    Derived voice data linked to clinical and demographic information.
    Instance Type
    Participants, recording sessions, and derived data per recording.
    Data Type
    Derived spectrograms (513 x N), static acoustic/phonetic/prosodic features, and participant phenotype/clinical questionnaire responses.
    Counts
    12,523
    Label
    Participants selected from predefined disease cohorts; no explicit classification labels provided in v1.0.
    Sampling Strategies
    1. ID
      sampling:clinical-cohorts
      Name
      Cohort-based sampling at specialty clinics
      Description
      Non-random sample of patients selected from five predetermined clinical groups at specialty clinics and institutions.
      Used Software
      Is Sample
      • yes
      Is Random
      • no
      Source Data
      • Specialty clinics in North America with predefined disorder cohorts
      Is Representative
      • Not intended to be representative of the general population
      Why Not Representative
      • Purposeful enrichment for conditions known to manifest in the voice waveform
      Strategies
      • Deterministic cohort-based inclusion from predefined groups
    Missing Information
    1. ID
      missing:raw-audio
      Name
      Raw audio and free-speech transcripts removed
      Description
      Raw audio waveforms and transcripts of free speech are not included in v1.0 to reduce risk and protect privacy.
      Used Software
      Missing
      • Raw audio waveforms
      • Transcripts of free speech audio
      Why Missing
      • Privacy and de-identification; low-risk initial release with only derived data
  1. ID
    subpops:cohorts
    Name
    Disorder cohorts
    Description
    Identification
    • Voice disorders
    • Neurological and neurodegenerative disorders
    • Mood and psychiatric disorders
    • Respiratory disorders
    • Pediatric voice and speech disorders (not included in v1.0)
    Distribution
    • v1.0 contains adult cohort data only
    Used Software
DescriptionIDNameUsed Software
Restricted-access database; files available to credentialed users who sign the DUA and complete required training.distfmt:portalCredentialed access via Health Data Nexus
Spectrograms stored as Parquetdistfmt:parquetParquet
phenotype.tsv and static_features.tsvdistfmt:tsvTSV
phenotype.json and static_features.json data dictionariesdistfmt:jsonJSON
  1. ID
    distdate:2024-11-27
    Name
    Initial release date
    Description
    • 2024-11-27
    Used Software
🔍

Collection Process

How was the data acquired?

bridge2ai-voice-v1-0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
The human voice contains complex acoustic markers which have been linked to important health conditions including dementia, mood disorders, and cancer. When viewed as a biomarker, voice is a promising characteristic to measure as it is simple to collect, cost-effective, and has broad clinical utility. The Bridge2AI-Voice project seeks to create an ethically sourced flagship dataset to enable future research in artificial intelligence and support critical insights into the use of voice as a biomarker of health. Bridge2AI-Voice v1.0, the initial release, provides 12,523 recordings for 306 participants collected across five sites in North America. Participants were selected based on known conditions which manifest within the voice waveform including voice disorders, neurological disorders, mood disorders, and respiratory disorders. The initial release contains data considered low risk, including derivations such as spectrograms but not the original voice recordings. Detailed demographic, clinical, and validated questionnaire data are also made available.
English
2024-11-27
2024-11-27
  • voice
  • bridge2ai
  • audio
RoleNameORCIDAffiliation
ContributorBridge2AI-Voice Teamcreator:bridge2ai-voice-team-
  1. ID
    gap:multisite-diverse-standardized-voice
    Name
    Diverse, multi-institutional, standardized voice dataset
    Description
    Addresses the lack of large, diverse, multi-institutional voice datasets with standardized collection protocols and linked clinical/demographic data for AI research.
    Used Software
    Response
    Provide a large, high-quality, multi-institutional and diverse voice dataset with standardized protocols linked to health information.
  1. ID
    relationships:participant-session
    Name
    Participant-session linkage
    Description
    • Spectrogram entries include participant_id, session_id, and task_name linking recordings to participants and sessions.
    Used Software
  1. ID
    splits:na
    Name
    Data splits (not specified)
    Description
    • No recommended train/validation/test splits are provided in v1.0.
    Used Software
  1. ID
    anomalies:none-reported
    Name
    No anomalies reported
    Description
    • No specific errors, noise sources, or redundancies were reported in the release notes.
    Used Software
DescriptionIDNameUsed Software
external_resources: ['https://docs.b2ai-voice.org']
future_guarantees: ['Not stated']
archival: ['DOI versioning is provided']
ext:docsDocumentation website
external_resources: ['https://doi.org/10.5281/zenodo.14148755']
future_guarantees: ['Not stated']
archival: ['Zenodo DOI']
ext:redcap-zenodoBridge2AI Voice REDCap (v3.20.0)
  1. ID
    confidential:clinical-linkage
    Name
    Clinical and demographic information
    Description
    • Dataset includes clinical and demographic information linked to recordings; v1.0 includes de-identified, low-risk derived data only.
    Used Software
  1. ID
    sensitive:health-data
    Name
    Health-related data
    Description
    • Demographics, clinical information, and questionnaire responses related to health conditions
    Used Software
  1. ID
    acquisition:direct-recording-and-questionnaires
    Name
    Direct recording with standardized protocol and questionnaires
    Description
    • Raw audio recorded via a custom tablet application with headset when possible; demographic and disease-specific data collected via validated questionnaires and clinical instruments.
    • Multiple tasks including sustained vowel phonation; sessions per participant as needed.
    Used Software
    Was Directly Observed
    yes (raw audio recordings; derived data released)
    Was Reported By Subjects
    yes (validated questionnaires)
    Was Inferred Derived
    yes (derived spectrograms and features from raw audio; ASR transcripts initially generated then removed prior to release)
    Was Validated Verified
    yes (validated questionnaires; standardized collection protocol)
  1. ID
    collection:tablet-headset-redcap
    Name
    Custom tablet application and headset; REDCap for data management
    Description
    • Standardized protocol; custom data collection app on tablet with headset when possible; REDCap used for clinical/phenotype data; export and conversion performed with an open-source library (b2aiprep).
    Used Software
    IDNameURL
    software:redcapREDCap
    software:b2aiprepb2aiprephttps://github.com/sensein/b2aiprep
  1. ID
    collectors:project-investigators
    Name
    Project investigators at specialty clinics and institutions
    Description
    • Patients screened for inclusion/exclusion; investigators obtained consent and conducted data collection sessions.
    Used Software
DescriptionIDNameUsed Software
Data collection and sharing approved by the University of South Florida Institutional Review Board.ethics:usf-irbUniversity of South Florida IRB
Submission for review to the University of Toronto Research Ethics Board.ethics:utoronto-rebUniversity of Toronto Research Ethics Board
  1. ID
    prep:audio-standardization
    Name
    Audio standardization and feature derivation
    Description
    • Raw audio converted to mono, resampled to 16 kHz with a Butterworth anti-aliasing filter; short-time FFT spectrograms computed (25 ms window, 10 ms hop, 512-point FFT).
    • Acoustic features extracted with OpenSMILE; phonetic and prosodic features computed using Parselmouth and Praat; ASR transcriptions generated using Whisper Large (transcripts of free speech removed before release).
    Used Software
    IDNameURL
    software:opensmileOpenSMILE
    software:parselmouthParselmouth
    software:praatPraat
    software:torchaudioTorchaudio
    software:whisperOpenAI Whisper Large
    software:b2aiprepb2aiprephttps://github.com/sensein/b2aiprep
  1. ID
    cleaning:deid-and-removals
    Name
    De-identification and removal of sensitive fields
    Description
    • HIPAA Safe Harbor identifiers removed; state/province removed (country retained); transcripts of free speech removed; raw audio waveforms omitted from v1.0.
    Used Software
  1. ID
    labeling:asr-internal-then-removed
    Name
    ASR transcription (internal) then removal
    Description
    • Transcriptions were generated using Whisper Large during processing; transcripts of free speech audio were removed in the released dataset.
    Used Software
    • ID
      software:whisper
      Name
      OpenAI Whisper Large
  1. ID
    raw:audio
    Name
    Raw audio waveforms
    Description
    • Raw audio recorded during sessions; not included in v1.0 release. Future releases may include voice data with additional precautions for data security.
    Used Software
ID
deid:hipaa-safe-harbor
Name
HIPAA Safe Harbor de-identification
Description
  • Removal of HIPAA Safe Harbor identifiers
  • Removal of state/province (country retained)
  • Removal of transcripts of free speech
  • Omission of raw audio in v1.0 (derived data only)
mixed
RoleNameORCIDAffiliation
Contributorspectrograms.parquetsubset:spectrograms-parquet-
Contributorphenotype.tsvsubset:phenotype-tsv-
Contributorphenotype.jsonsubset:phenotype-json-
Contributorstatic_features.tsvsubset:static-features-tsv-
Contributorstatic_features.jsonsubset:static-features-json-
🚀

Uses

What (other) tasks could the dataset be used for?

  1. ID
    task:acoustic-feature-analysis
    Name
    Acoustic feature analysis and modeling
    Description
    Support AI/ML methods for extracting prognostic and diagnostic information from voice-derived spectrograms and features.
    Used Software
    Response
    AI/ML research on voice-derived spectrograms and features linked to health data.
  1. ID
    othertasks:health-ai
    Name
    Additional health AI tasks
    Description
    • Potential use in screening, monitoring, and characterization of conditions that manifest in voice and speech.
    Used Software
  1. ID
    future:low-risk-derivatives
    Name
    Low-risk derivative-only release considerations
    Description
    • Initial release limits data to derived spectrograms and features (no raw audio or free-speech transcripts) to reduce privacy risks and support ethical use.
    Used Software
ID
terms:registered-access
Name
Registered access license and DUA
Description
  • License
    Bridge2AI Voice Registered Access License
  • Data Use Agreement
    Bridge2AI Voice Registered Access Agreement
  • Access Policy
    Only credentialed users who sign the DUA can access the files
  • Required Training
    TCPS 2: CORE 2022
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
ID
versioning:doi-latest
Name
DOI versioning
Description
🔄

Maintenance

How will the dataset be maintained?

2024-11-27
1.0
ID
updates:future-voice-inclusion
Name
Planned future updates
Description
  • v1.0 is the first release; future releases aim to include voice audio with additional data security precautions.
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema