healthnexus tab files d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

GrantorGrant NameGrant Number
National Institutes of HealthBridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before3OT2OD032720-01S1
  • Response
    Create an ethically sourced, diverse, multi-institutional dataset of voice linked to clinical and questionnaire data to enable AI research on voice as a biomarker of health and support clinically meaningful insights.
📊

Composition

What do the instances represent?

CountsData TypeInstance TypeMissing InformationRepresentationSampling Strategies
12523Derived features (spectrograms, acoustic, phonetic/prosodic); no raw audio waveforms in v1.0Audio-derived instance{'missing': ['Original audio waveforms'], 'why_missing': ['Omitted in initial low-risk release; planned for future releases with additional safeguards']}, {'missing': ['Transcripts of free speech audio'], 'why_missing': ['Removed during de-identification']}Voice-derived recordings (spectrograms/features) per recording{'is_sample': ['yes'], 'is_random': ['no'], 'source_data': ['Patients presenting at specialty clinics across five North American sites'], 'is_representative': ['no'], 'why_not_representative': ['Participants were selected based on membership in predefined disease cohorts'], 'strategies': ['Targeted enrollment by predefined disease categories']}
306Demographics, clinical, and validated questionnaire responsesParticipantParticipant-level phenotype records{'is_sample': ['yes'], 'is_random': ['no'], 'source_data': ['Specialty clinics at five North American sites'], 'is_representative': ['no'], 'why_not_representative': ['Cohort-based selection for voice-relevant conditions']}
  • Identification
    • Adult cohort (v1.0); disease categories include voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, respiratory disorders
    Distribution
    • Participants selected based on membership in predefined disease cohorts
  • Description
    • Parquet
    • TSV
    • JSON
  • Description
    • 2024-11-27
🔍

Collection Process

How was the data acquired?

bridge2ai-voice-v1.0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information v1.0
Bridge2AI-Voice is a multi-site, ethically sourced dataset of human voice linked to clinical and questionnaire information to enable research on voice as a biomarker of health. Version 1.0 includes 12,523 recordings from 306 participants across five North American sites, focusing on cohorts with known voice-relevant conditions (voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, and respiratory disorders). The initial release contains only low-risk, derived data (e.g., spectrograms and acoustic features) and de-identified phenotype tables; original audio waveforms and free-speech transcripts are not included. Data were collected via a standardized protocol using a custom tablet application, with preprocessing that included resampling to 16 kHz with a Butterworth anti-aliasing filter, STFT-based spectrogram generation, extraction of acoustic and phonetic/prosodic features (OpenSMILE, Parselmouth/Praat), and automatic transcription using OpenAI Whisper Large. Documentation: "https://docs.b2ai-voice.org."
2024-11-27
  • voice
  • bridge2ai
  • audio
  • VOICE
Name
Alistair Johnson
Jean-Christophe Bélisle-Pipon
David Dorr
Satrajit Ghosh
Philip Payne
Maria Powell
Anaïs Rameau
Vardit Ravitsky
Alexandros Sigaras
Olivier Elemento
Yael Bensoussan
  • Response
    Addresses the lack of large, high-quality, standardized, and demographically diverse voice datasets linked to clinical and questionnaire data across multiple institutions.
  • Description
    • Recordings are linked to participants and sessions via participant_id and session_id; tasks identified by task_name
Description
  • HIPAA Safe Harbor identifiers removed (e.g., names, contact details, precise dates, etc.)
  • State and province removed; country retained
  • Free speech transcripts removed; original audio waveforms omitted in v1.0
  • Description
    • Contains de-identified health-related information (demographics, clinical variables, questionnaire responses)
  • Description
    • De-identified clinical and questionnaire data; identifiable communications removed
  1. Description
    • Standardized protocol with voice tasks (e.g., sustained vowel), demographics, health and targeted questionnaires
    • Custom tablet application; headset used when possible
    • Data exported from REDCap via open-source tooling
    Was Directly Observed
    yes
    Was Reported By Subjects
    yes
    Was Inferred Derived
    yes
  • Description
    • Custom tablet app and headset for data capture; REDCap used for data entry/export
  • Description
    • Project investigators at five North American sites; patients presenting at specialty clinics were screened and consented
  • Description
    • Data collection and sharing approved by the University of South Florida IRB; submitted for review to the University of Toronto Research Ethics Board
  • Description
    • Raw audio converted to mono and resampled to 16 kHz with a Butterworth anti-aliasing filter
    • Spectrograms computed via short-time FFT (25 ms window, 10 ms hop, 512-point FFT)
    • Acoustic features extracted with OpenSMILE
    • Phonetic and prosodic features computed with Parselmouth and Praat
    • Transcriptions generated with OpenAI Whisper Large model
    Used Software
    NameURL
    OpenSMILE
    Parselmouth
    Praat
    torchaudio
    OpenAI Whisper Large
    b2aiprephttps://github.com/sensein/b2aiprep
  • Description
    • HIPAA Safe Harbor de-identification (removal of identifiers and fine-grained dates)
    • Removal of free speech transcripts
    • Omission of original audio waveforms from initial release
  • Description
    • Automatic Transcription Using Openai Whisper Large (note
      free speech transcripts are not included in v1.0)
  • Description
    • Original audio waveforms collected; not distributed in v1.0. Future releases aim to include voice data with additional precautions.
  • Description
    • Health Data Nexus
    • Temerty Centre for AI Research and Education in Medicine
mixed (Parquet dense arrays and tabular TSV/JSON)
DescriptionFormatIDMedia TypeNamePathTitle
Parquet dataset containing time-frequency spectrograms derived from voice recordings; includes parti...spectrograms-parquetapplication/x-parquetspectrograms.parquetspectrograms.parquetSpectrograms (derived from voice waveforms)
Tab-delimited file with one row per participant containing demographics, acoustic confounders, and v...phenotype-tsvtext/tab-separated-valuesphenotype.tsvphenotype.tsvPhenotype table
JSON data dictionary describing columns in phenotype.tsv; includes a one-sentence description for ea...JSONphenotype-jsonapplication/jsonphenotype.jsonphenotype.jsonPhenotype data dictionary
Tab-delimited file with one row per recording containing features derived from OpenSMILE, Praat/Pars...static-features-tsvtext/tab-separated-valuesstatic_features.tsvstatic_features.tsvStatic audio-derived features
JSON data dictionary describing features in static_features.tsv; includes a description for each fea...JSONstatic-features-jsonapplication/jsonstatic_features.jsonstatic_features.jsonStatic features data dictionary
🚀

Uses

What (other) tasks could the dataset be used for?

  • Response
    Research and development of AI/ML methods for analyzing voice-derived features and their associations with health conditions; exploratory and hypothesis-driven studies on voice as a biomarker.
Description
  • Bridge2AI Voice Registered Access License
  • Bridge2AI Voice Registered Access Agreement (DUA)
  • Access Policy
    Only credentialed users who sign the DUA can access the files
  • Required Training
    TCPS 2: CORE 2022
📤

Distribution

How will the dataset be distributed?

Description
🔄

Maintenance

How will the dataset be maintained?

1.0
Description
  • Future releases aim to include voice data with additional precautions to ensure data security
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema