VOICE d4d alldocs

Datasheet for Dataset - Human Readable Format

🔍

Collection Process

How was the data acquired?

Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
The human voice contains complex acoustic markers which have been linked to important health conditions including dementia, mood disorders, and cancer. When viewed as a biomarker, voice is a promising characteristic to measure as it is simple to collect, cost-effective, and has broad clinical utility. Recent advances in artificial intelligence have provided techniques to extract previously unknown prognostically useful information from dense data elements such as images. The Bridge2AI-Voice project seeks to create an ethically sourced flagship dataset to enable future research in artificial intelligence and support critical insights into the use of voice as a biomarker of health. Bridge2AI-Voice provides a comprehensive collection of derived data from voice recordings with corresponding clinical information, demographic data, and validated questionnaires collected across multiple North American sites. Initial releases focus on low-risk, de-identified derived data (e.g., spectrograms, acoustic features) and phenotype data to reduce re-identification risk while enabling broad research utility.
  • voice
  • bridge2ai
  • PhysioNet
  • Health Data Nexus
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anais Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  • Bridge2AI-Voice Consortium
  • PhysioNet
  1. ID
    physionet-b2ai-voice-1.1
    Name
    Bridge2AI-Voice v1.1 (PhysioNet)
    Title
    Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information (version 1.1)
    Description
    Derived audio representations (e.g., spectrograms, MFCCs, acoustic and phonetic/prosodic features) and associated phenotype and questionnaire data from adult participants recruited at specialty clinics across five North American sites. Raw audio is not included in this release to reduce re-identification risk. Common identifiers include participant_id, session_id, and task_name.
    DOI
    doi:10.13026/249v-w155
    Issued
    2025-01-17
    Version
    1.1
    Keywords
    • voice
    • bridge2ai
    • PhysioNet
    • RRID:SCR_007345
    License
    Bridge2AI Voice Registered Access License
    Created By
    • PhysioNet
    • Bridge2AI-Voice Consortium
    Purposes
    • Name
      primary-purpose
      Response
      Enable ethically sourced, large-scale research on voice as a biomarker of health by linking derived voice representations to demographic, clinical, and questionnaire data.
    Tasks
    NameResponse
    example-task-1Development and benchmarking of models associating voice-derived features with health conditions.
    example-task-2Exploration of acoustic, phonetic, and prosodic correlates of disease using de-identified derived da...
    Addressing Gaps
    • Name
      gap-addressed
      Response
      Lack of an ethically sourced, clinically linked, multi-site voice dataset with robust de-identification suitable for AI research.
    Funders
    GrantorGrant NameGrant Number
    National Institutes of HealthBridge2AI: Voice as a Biomarker of Health3OT2OD032720-01S1
    National Institute of Biomedical Imaging and Bioengineering (NIBIB), NIHPhysioNet infrastructure supportR01EB030362
    Instances
    1. Name
      dataset-instances
      Representation
      Derived representations of voice recordings linked to participant-level phenotype and questionnaire data.
      Instance Type
      Participants and their recordings; per-recording static feature rows and per-participant phenotype rows.
      Data Type
      De-identified derived features (spectrograms, MFCCs, acoustic and phonetic/prosodic features) and structured phenotype/questionnaire data.
      Counts
      12,523
      Label
      Clinical and demographic attributes (e.g., condition groups, validated questionnaires) are available as metadata; no explicit machine-learning labels are defined in this release.
      Sampling Strategies
      1. Name
        purposive-sampling
        Is Sample
        • True
        Is Random
        • False
        Source Data
        • Adult patients recruited at specialty clinics across five North American sites
        Is Representative
        • False
        Why Not Representative
        • Participants were selected based on conditions known to manifest in voice, which may affect generalizability.
        Strategies
        • Purposive sampling by predefined condition groups
      Missing Information
      • Name
        privacy-removals
        Missing
        • Free speech transcripts
        • Raw audio waveforms
        Why Missing
        • Removed to reduce re-identification risk and comply with HIPAA Safe Harbor.
    Subpopulations
    • Name
      adult-cohort
      Identification
      • Adults only in v1.1
      Distribution
      • 306 participants; 12,523 recordings across predefined clinical condition groups
    Sensitive Elements
    • Name
      health-data
      Description
      • Contains de-identified health-related data (clinical and questionnaire information).
    Is Deidentified
    Name
    hipaa-safe-harbor
    Description
    • HIPAA Safe Harbor de-identification applied; removal of identifiers (e.g., names, fine-grained dates, contact details, geographic locators below country), removal of free speech transcripts, and omission of raw audio in v1.1.
    Acquisition Methods
    1. Name
      acquisition-overview
      Description
      • Voice recordings directly observed; phenotype/questionnaire data reported by participants; derived features computed from recordings.
      Was Directly Observed
      True
      Was Reported By Subjects
      True
      Was Inferred Derived
      True
      Was Validated Verified
      Standardized protocols and multi-site QA were used; derived features computed via established toolkits.
    Collection Mechanisms
    • Name
      data-capture
      Description
      • Custom tablet application and headset microphones used when possible; data exported from REDCap and converted using an open-source library.
    Data Collectors
    • Name
      site-teams
      Description
      • Researchers and clinicians at five North American specialty-clinic sites.
    Collection Timeframes
    Ethical Reviews
    • Name
      irb-approval
      Description
      • Data collection and sharing approved by the University of South Florida Institutional Review Board.
    Preprocessing Strategies
    • Name
      waveform-prep-and-feature-extraction
      Description
      • Waveforms converted to mono and resampled to 16 kHz with anti-aliasing; spectrograms via STFT (25 ms window, 10 ms hop, 512-point FFT, power); 60-coefficient MFCCs computed; static acoustic features extracted (e.g., openSMILE) and phonetic/prosodic features via Parselmouth/Praat; transcripts generated by OpenAI Whisper Large (free speech transcripts removed prior to release).
      Used Software
      DescriptionNameURL
      Open-source library used to preprocess waveforms and merge phenotype data.b2aiprephttps://github.com/sensein/b2aiprep
      Acoustic feature extraction toolkit.openSMILE
      Speech analysis software for phonetics.Praat
      Python interface to Praat.Parselmouth
      Audio processing components for PyTorch.torchaudio
      ASR model used to generate transcripts (free speech transcripts removed before release).OpenAI Whisper Large
    Cleaning Strategies
    • Name
      de-identification-and-redactions
      Description
      • Removal of HIPAA Safe Harbor identifiers; removal of state/province with retention of country; removal of free speech transcripts; omission of raw audio waveforms in v1.1.
    Labeling Strategies
    • Name
      transcription
      Description
      • Automatic speech recognition using OpenAI Whisper Large; transcripts of free speech removed prior to release.
    Raw Sources
    • Name
      raw-audio-availability
      Description
      • Raw audio waveforms were collected but are not distributed in v1.1; only derived representations are provided.
    Existing Uses
    DescriptionName
    {'Johnson, A., Bélisle-Pipon, J., Dorr, D., Ghosh, S., Payne, P., Powell, M., Rameau, A., Ravitsky, V., Sigaras, A., Elemento, O., & Bensoussan, Y. (2025). Bridge2AI-Voice': 'An ethically-sourced, diverse voice dataset linked to health information (version 1.1). PhysioNet. RRID:SCR_007345. https://doi.org/10.13026/249v-w155'}dataset-citation
    Goldberger, A., et al. (2000). PhysioBank, PhysioToolkit, and PhysioNet. Circulation. RRID:SCR_007345.platform-citation
    Use Repository
    Other Tasks
    Future Use Impacts
    • Name
      derived-only-constraints
      Description
      • The absence of raw audio may limit certain analyses (e.g., new feature extraction requiring original waveforms) but reduces re-identification risk.
    Discouraged Uses
    Distribution Formats
    • Name
      formats
      Description
      • Parquet
      • TSV
      • JSON
    Distribution Dates
    • Name
      initial-release
      Description
      • 2025-01-17
    License And Use Terms
    Name
    access-and-licensing
    Description
    • Platform
      PhysioNet
    • Access
      Registered/Restricted Access; only registered users who sign the Bridge2AI Voice Registered Access Agreement may access files.
    • License
      Bridge2AI Voice Registered Access License; applicable Data Use Agreement required.
    Ip Restrictions
    Regulatory Restrictions
    Maintainers
    • Name
      maintainers
      Description
      • PhysioNet platform team
      • Bridge2AI-Voice consortium
    Errata
    Updates
    Name
    version-history
    Description
    • V1.0 (2024)
      Initial release.
    • V1.1 (2025 01 17)
      Added MFCCs; files for v1.1 are no longer available on the platform.
    • Newer Versions Available
      2.0.0 (2025-04-16), 2.0.1 (2025-08-18).
    Retention Limit
    Version Access
    Name
    older-version-availability
    Description
    • Older version 1.1 files are no longer available; latest version on the platform is 2.0.1.
    Extension Mechanism
    Is Tabular
    yes
  2. ID
    healthdatanexus-voice-1.0
    Name
    Health Data Nexus VOICE 1.0
    Title
    VOICE 1.0
    Description
    Resource page for the VOICE 1.0 project on Health Data Nexus.
    Page
    https://healthdatanexus.ai/content/b2ai-voice/1.0/
    Version
    1.0
    DOI
    doi:10.57764/qb6h-em84
    Keywords
    • VOICE
    • Health Data Nexus
    • b2ai-voice
    • healthdatanexus.ai
    Created By
    • Health Data Nexus
    Distribution Formats
    Distribution Dates
    License And Use Terms
    Name
    unspecified-license
    Description
    • Resource/landing page for the VOICE 1.0 project; licensing for downloadable data not specified on this page.
Generated on 2025-11-09 10:17:35 using Bridge2AI Data Sheets Schema