healthnexus tab citationModal d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  1. ID
    purpose-1
    Name
    Dataset purpose
    Description
    Enable AI research into voice as a biomarker of health across multiple clinical conditions using ethically sourced, diverse, multi-institutional data linked to health information.
    Response
    Create an ethically sourced, diverse voice dataset linked to clinical information to accelerate AI research on voice as a biomarker of health.
GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before3OT2OD032720-01S1
📊

Composition

What do the instances represent?

CountsData TypeIDInstance TypeLabelMissing InformationNameRepresentationSampling Strategies
12523Derived features from raw audio including spectrograms and acoustic/phonetic/prosodic features.inst-recordingsRecordings (per session and task)Task names per recording (e.g., sustained phonation vowel tasks); no raw audio included in v1.0.{'id': 'miss-1', 'name': 'Removed elements', 'missing': ['Raw audio waveforms (omitted in v1.0)', 'Transcripts of free speech audio'], 'why_missing': ['Privacy and low-risk release; de-identification measures']}Voice-derived recordingsVoice recordings-derived data (spectrograms and engineered features) aligned to session and task met...{'id': 'samp-1', 'name': 'Cohort sampling', 'is_sample': [True], 'is_random': [False], 'source_data': ['Patients presenting at specialty clinics/institutions across five North American sites'], 'is_representative': ['not specified'], 'strategies': ['Targeted enrollment of five predetermined clinical groups (respiratory, voice, neurological, mood/psychiatric, pediatric)']}
306Demographics, acoustic confounders, validated questionnaires, disease-specific information collected...inst-participantsParticipants (adult cohort in v1.0)Disease cohort categories and clinical attributes captured in phenotype file.ParticipantsIndividual participants with linked demographic, clinical, and questionnaire data.
DistributionIDIdentificationName
Not specified in this recordsubpop-1v1.0 includes only adult participantsAdult cohort
Per-group counts not specified in this recordsubpop-2Voice disorders, Neurological and neurodegenerative disorders, ... (+3 more)Clinical condition groups
DescriptionIDName
Credentialed access via Health Data Nexus portal; files available after DUA and required training completiondist-portalDistribution channel
Parquet (spectrograms), TSV (phenotype and static features), JSON (data dictionaries)dist-formatsFile formats
  • ID
    distdate-1
    Name
    Initial public release
    Description
    • 2024-11-27 (v1.0)
🔍

Collection Process

How was the data acquired?

b2ai-voice-v1.0
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information v1.0
The Bridge2AI-Voice project seeks to create an ethically sourced flagship dataset to enable future research in artificial intelligence and support critical insights into the use of voice as a biomarker of health. Bridge2AI-Voice v1.0 provides 12,523 recordings for 306 participants collected across five sites in North America. Participants were selected based on known conditions which manifest within the voice waveform including voice disorders, neurological disorders, mood disorders, and respiratory disorders. The initial release contains data considered low risk, including derivations such as spectrograms (not the original voice recordings), as well as detailed demographic, clinical, and validated questionnaire data.
2024-11-27
2024-11-27
English
  • voice
  • bridge2ai
  • audio
  • spectrograms
  • phenotype
  • parquet
  • tsv
  • health
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
bibo:draft
  1. ID
    gap-1
    Name
    Gaps addressed
    Description
    Addresses limitations of prior work with small datasets, limited demographic diversity, and non-standardized protocols.
    Response
    Provides a large, multi-institutional, standardized, ethically sourced dataset with diverse demographics and linked clinical information.
  1. ID
    acq-1
    Name
    Data acquisition
    Description
    • Standardized protocol with demographic information, health questionnaires, targeted acoustic confounders, disease-specific information, and voice recording tasks (e.g., sustained vowel phonation)
    • Data exported from REDCap and converted using an open-source library developed by the team
    Was Directly Observed
    yes (voice tasks recorded via headset/tablet app)
    Was Reported By Subjects
    yes (validated questionnaires and self-reported items)
    Was Inferred Derived
    yes (acoustic/phonetic/prosodic features and spectrograms derived from raw audio; ASR transcriptions generated then removed for free speech)
    Was Validated Verified
    yes (standardized multi-site protocol and IRB/REB oversight)
  • ID
    mech-1
    Name
    Collection mechanisms
    Description
    • Custom tablet application for data capture; headset used when possible
    • Data managed in REDCap; conversion and preprocessing via open-source b2aiprep library
    • Standardized tasks including sustained phonation
  • ID
    collectors-1
    Name
    Data collectors
    Description
    • Project investigators at specialty clinics and institutions; participants screened for inclusion/exclusion prior to visit; consent obtained
  • ID
    timeframe-1
    Name
    Collection timeframe
    Description
    • Data collected across five sites in North America; specific collection dates not provided in this record; first public release on 2024-11-27
DescriptionIDName
Data collection and sharing approved by the University of South Florida Institutional Review Boardethics-usfIRB approval (USF)
Submission to the University of Toronto Research Ethics Board for reviewethics-utorontoREB submission (U of Toronto)
  1. ID
    prep-1
    Name
    Audio preprocessing and feature derivation
    Description
    • Raw audio converted to monaural and resampled to 16 kHz with a Butterworth anti-aliasing filter
    • Spectrograms computed via STFT (25 ms window, 10 ms hop, 512-point FFT)
    • Acoustic features extracted with OpenSMILE
    • Phonetic and prosodic features computed using Parselmouth and Praat (e.g., F0, formants, voice quality)
    • Transcriptions generated with OpenAI Whisper Large (free speech transcripts later removed from release)
    Used Software
    IDNameURLVersion
    sw-opensmileopenSMILEhttps://www.audeering.com/research/opensmile
    sw-praatPraathttps://www.fon.hum.uva.nl/praat/
    sw-parselmouthParselmouthhttps://parselmouth.readthedocs.io/
    sw-torchaudiotorchaudiohttps://pytorch.org/audio2.1
    sw-whisperOpenAI Whisper (Large)https://github.com/openai/whisper
  • ID
    clean-1
    Name
    Data cleaning and packaging
    Description
    • Conversion/merging of source data into phenotype files and spectrogram parquet using b2aiprep
    • Creation of data dictionaries (phenotype.json, static_features.json) describing each column/feature
  • ID
    label-1
    Name
    Labeling/transcription
    Description
    • ASR transcriptions generated using Whisper Large; transcripts of free speech were removed prior to release
  • ID
    raw-1
    Name
    Raw data availability
    Description
    • In v1.0, audio waveforms are omitted; only spectrograms and derived features are provided. Raw audio may be considered for future releases with additional safeguards.
ArchivalExternal ResourcesIDNameRestrictions
DOI landing page provides versioned records (v1.0)https://docs.b2ai-voice.orgext-docsProject documentationCredentialed access with DUA and training required for files
https://doi.org/10.5281/zenodo.14148755ext-redcapBridge2AI Voice REDCap (v3.20.0) - Zenodo
https://github.com/sensein/b2aiprepext-b2aiprepb2aiprep library (open source)
ID
deid-1
Name
De-identification
Description
  • HIPAA Safe Harbor identifiers removed (e.g., names, detailed dates, contact info, SSNs, MRNs, device IDs, URLs, biometric identifiers)
  • State/province removed; country of data collection retained
  • Transcripts of free speech removed
  • Audio waveforms omitted in v1.0; only derived data released
  • ID
    sens-1
    Name
    Sensitive data considerations
    Description
    • Contains demographic, clinical, and questionnaire data; released elements considered low risk after de-identification
  • ID
    maint-1
    Name
    Hosting and maintenance
    Description
    • Health Data Nexus
    • Temerty Centre for AI Research and Education in Medicine (Supported by the Temerty Foundation)
Mixed (Parquet, TSV, JSON)
DescriptionDialectEncodingFormatIDIs Data SplitIs SubpopulationMedia TypeNamePathTitle
Parquet file storing dense time-frequency representations derived from raw audio waveforms with part...UTF-8subset-spectrograms-parquetFalseAdult cohort (v1.0 only)application/x-parquetspectrograms.parquetspectrograms.parquetSpectrograms derived from raw audio
Tab-delimited file containing demographics, acoustic confounders, and responses to validated questio...delimiter:
header: True
UTF-8subset-phenotype-tsvFalseAdult cohort (v1.0 only)text/tab-separated-valuesphenotype.tsvphenotype.tsvPhenotype data (participant-level)
JSON data dictionary detailing each phenotype column and its description.UTF-8JSONsubset-phenotype-jsonFalseAdult cohort (v1.0 only)application/jsonphenotype.jsonphenotype.jsonData dictionary for phenotype data
Tab-delimited file with one row per recording; includes features from openSMILE, Praat, parselmouth,...delimiter:
header: True
UTF-8subset-static-features-tsvFalseAdult cohort (v1.0 only)text/tab-separated-valuesstatic_features.tsvstatic_features.tsvStatic audio features
JSON data dictionary describing each feature present in static_features.tsv.UTF-8JSONsubset-static-features-jsonFalseAdult cohort (v1.0 only)application/jsonstatic_features.jsonstatic_features.jsonData dictionary for static features
🚀

Uses

What (other) tasks could the dataset be used for?

  1. ID
    task-1
    Name
    Clinical voice AI tasks
    Description
    Model development and evaluation for condition detection/monitoring from voice-derived signals.
    Response
    Disease screening, detection, and monitoring tasks related to voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, and respiratory disorders.
ID
terms-1
Name
License and access terms
Description
  • License
    Bridge2AI Voice Registered Access License
  • Data Use Agreement
    Bridge2AI Voice Registered Access Agreement
  • Access Policy
    Only credentialed users who sign the DUA can access the files
  • Required Training
    TCPS 2: CORE 2022
  • ID
    impact-1
    Name
    Considerations for future use
    Description
    • v1.0 excludes raw audio and free speech transcripts, which may limit certain modeling approaches (e.g., end-to-end raw waveform models)
    • Adult-only cohort and targeted clinical groups may affect generalizability to pediatric populations and to conditions outside the five categories
    • Planned inclusion of raw audio in future releases will require continued attention to privacy and security safeguards
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
ID
versions-1
Name
Versioning and access
Description
🔄

Maintenance

How will the dataset be maintained?

1.0
ID
updates-1
Name
Update plan
Description
  • Future releases aim to include original voice data with additional security precautions; pediatric cohort planned for future inclusion
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema