healthnexus tab description d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  1. ID
    purpose-1
    Name
    Research enablement for voice as a biomarker of health
    Description
    Enable future research in artificial intelligence using ethically sourced, clinically linked voice-derived data to investigate acoustic markers of health conditions.
    Response
    Create a flagship dataset to support AI research on the human voice as a biomarker across multiple health domains.
GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before3OT2OD032720-01S1
📊

Composition

What do the instances represent?

CountsData TypeIDInstance TypeMissing InformationNameRepresentationSampling Strategies
12523Derived spectrograms (513 x N), engineered acoustic/phonetic/prosodic features; transcription metada...instance-recordingsParticipants, sessions, and recording-derived instancesRecording-derived instancesAudio recording–derived instances (e.g., spectrogram tensors, engineered features) linked to session...{'id': 'sampling-1', 'strategies': ['Condition-focused cohort inclusion across five North American sites; non-random selection'], 'is_sample': ['yes'], 'is_random': ['no'], 'source_data': ['Specialty clinics and institutions across five sites in North America'], 'is_representative': ['no'], 'why_not_representative': ['Condition-focused recruitment rather than population sampling']}
306Demographics, clinical information, validated questionnaires, task metadatainstance-participantsParticipants (adult cohort, v1.0)ParticipantsIndividual adult participants with linked demographics, clinical data, and validated questionnaire r...
  1. ID
    subpop-1
    Name
    Adult cohort (v1.0)
    Identification
    • Adult participants only in v1.0; pediatric cohort planned for future releases
    Distribution
    • Participants selected across five condition cohorts (voice, neurological, mood/psychiatric, respiratory; pediatric planned)
DescriptionIDName
spectrograms.parquet — dense, derived spectrogram tensors (513 x N) per recording with participant_id, session_id, task_name metadatadistfmt-1Parquet spectrograms
phenotype.tsv — participant-level demographics, acoustic confounders, validated questionnaires (tab-delimited), phenotype.json — data dictionary for phenotype fieldsdistfmt-2Phenotype data
static_features.tsv — one row per recording with features, static_features.json — data dictionary for featuresdistfmt-3Engineered acoustic/phonetic/prosodic features
  • ID
    distdate-1
    Name
    Initial public release (credentialed access)
    Description
    • v1.0 released 2024-11-27
  • ID
    access-1
    Name
    Credentialed access
    Description
    Access requires credentialing, DUA signature, and completion of TCPS 2: CORE 2022 training
🔍

Collection Process

How was the data acquired?

bridge2ai-voice-v1.0
Bridge2AI-Voice v1.0
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
The Bridge2AI-Voice project provides an ethically sourced, diverse dataset of data derived from human voice recordings linked to clinical and demographic information to enable AI research on voice as a biomarker of health. Version 1.0 includes 12,523 recordings for 306 adult participants across five North American sites. This initial release contains low-risk derived data (e.g., spectrograms and engineered features) and detailed demographic/clinical/questionnaire information; original audio waveforms and free speech transcripts are not included. Data were collected under a standardized multi-institutional protocol with de-identification following HIPAA Safe Harbor.
2024-11-27
  • voice
  • bridge2ai
  • audio
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  1. ID
    gap-1
    Name
    Diverse multi-institutional voice dataset gap
    Description
    Addresses the lack of large, high-quality, diverse, multi-institutional voice datasets linked to health biomarkers with standardized protocols and ethical oversight.
    Response
    Create a diverse, ethically sourced, clinically linked voice dataset with standardized collection and documentation.
RoleNameORCIDAffiliation
Principal InvestigatorAlistair Johnsonperson-alistair-johnson-
Principal InvestigatorJean-Christophe Bélisle-Piponperson-jean-christophe-belisle-pipon-
Principal InvestigatorDavid Dorrperson-david-dorr-
Principal InvestigatorSatrajit Ghoshperson-satrajit-ghosh-
Principal InvestigatorPhilip Payneperson-philip-payne-
Principal InvestigatorMaria Powellperson-maria-powell-
Principal InvestigatorAnaïs Rameauperson-anais-rameau-
Principal InvestigatorVardit Ravitskyperson-vardit-ravitsky-
Principal InvestigatorAlexandros Sigarasperson-alexandros-sigaras-
Principal InvestigatorOlivier Elementoperson-olivier-elemento-
Principal InvestigatorYael Bensoussanperson-yael-bensoussan-
  1. ID
    sampling-overall
    Name
    Cohort-based clinical recruitment
    Is Sample
    • yes
    Is Random
    • no
    Source Data
    • Specialty clinics and institutions at five North American sites
    Is Representative
    • no
    Representative Verification
    Why Not Representative
    • Purposeful inclusion of participants with conditions associated with voice changes
    Strategies
    • Deterministic cohort enrollment by predefined disease categories (voice disorders, neurological, mood/psychiatric, respiratory; pediatric planned)
ID
deid-1
Name
HIPAA Safe Harbor de-identification and restricted content
Description
  • HIPAA Safe Harbor identifiers removed (e.g., names, detailed dates, contact numbers, emails, IPs, SSNs, MRNs, plan IDs, device IDs, license/account numbers, vehicle IDs, URLs, full-face photos/biometrics, and other unique identifiers)
  • State/province removed; country of data collection retained
  • Transcripts of free speech audio removed
  • Original audio waveforms omitted from v1.0; only spectrograms and other derived features are released
  • ID
    sensitive-1
    Name
    Health-related data
    Description
    • Dataset includes clinical conditions and questionnaire responses linked to participants; distributed in de-identified form
  • ID
    confidential-1
    Name
    Restricted content and clinical linkages
    Description
    • Clinical and demographic linkages present; access is restricted via registered access with DUA and required training
  1. ID
    acquisition-1
    Name
    Standardized clinical protocol via custom app
    Description
    • Data collected with standardized protocol including demographics, validated questionnaires, condition-specific items, and voice tasks (e.g., sustained vowel phonation)
    • Recording sessions conducted via custom tablet application; headset used when possible; some participants had multiple sessions
    Was Directly Observed
    yes (voice tasks, recordings)
    Was Reported By Subjects
    yes (questionnaires)
    Was Inferred Derived
    yes (spectrograms, engineered features, ASR transcriptions)
    Was Validated Verified
    Standardized multi-site protocol with IRB/REB oversight
  • ID
    collectmech-1
    Name
    Custom data collection application and headset
    Description
    • Custom tablet application; headset microphone when possible; protocol detailed in referenced documentation and publications
  • ID
    datacollect-1
    Name
    Project investigators at specialty clinics and institutions
    Description
    • Participants screened for inclusion/exclusion; consent obtained prior to standardized data collection
  • ID
    irb-1
    Name
    Institutional Review and Ethics
    Description
    • Data collection and sharing approved by the University of South Florida Institutional Review Board; submitted to the University of Toronto Research Ethics Board
  1. ID
    preprocess-1
    Name
    Audio standardization and feature extraction
    Description
    • Raw audio converted to mono, resampled to 16 kHz with Butterworth anti-aliasing filter
    • Spectrograms via STFT (25 ms window, 10 ms hop, 512-point FFT)
    • Acoustic features via OpenSMILE
    • Phonetic/prosodic features via Parselmouth and Praat (e.g., f0, formants, voice quality)
    • Transcriptions generated using OpenAI Whisper Large
    Used Software
    IDNameURL
    sw-opensmileopenSMILEhttps://audeering.github.io/opensmile/
    sw-praatPraathttps://www.fon.hum.uva.nl/praat/
    sw-parselmouthParselmouth (Python interface to Praat)https://parselmouth.readthedocs.io/
    sw-torchaudiotorchaudiohttps://pytorch.org/audio
    sw-whisperOpenAI Whisper (Large)https://github.com/openai/whisper
    sw-b2aiprepb2aiprep (data preprocessing library)https://github.com/sensein/b2aiprep
    sw-librosalibrosahttps://librosa.org
  • ID
    raw-1
    Name
    Original audio waveforms (not released in v1.0)
    Description
    • Raw audio was collected and preprocessed but is not included in v1.0; future releases aim to include voice data with additional security precautions
ArchivalExternal ResourcesIDNameRestrictions
DOI assigned for dataset releaseshttps://docs.b2ai-voice.orgext-docsDocumentation website
https://doi.org/10.5281/zenodo.14148755ext-redcap-zenodoBridge2AI Voice REDCap (v3.20.0) metadata/toolsIndependent resource; referenced for tooling/context
ID
regulatory-1
Name
Access policy and required training
Description
  • Only Credentialed Users Who Sign The Dua And Complete Tcps 2
    CORE 2022 may access files
  • ID
    maint-1
    Name
    Health Data Nexus hosting
    Description
    • Hosted via Health Data Nexus (Temerty Centre for AI Research and Education in Medicine)
mixed
🚀

Uses

What (other) tasks could the dataset be used for?

  1. ID
    task-1
    Name
    Voice biomarker AI research
    Description
    AI/ML analysis of derived voice representations linked to clinical data to study associations with health conditions.
    Response
    Development and evaluation of AI methods using derived voice representations and linked phenotypes.
ID
license-terms-1
Name
Registered access and DUA
Description
  • Access restricted to credentialed users
  • Data Use Agreement required (Bridge2AI Voice Registered Access Agreement)
  • Required Training
    TCPS 2: CORE 2022
  • ID
    future-impact-1
    Name
    Impact of derived-only release
    Description
    • Exclusion of original audio in v1.0 may limit tasks requiring waveform-level processing or re-annotation; derived features and spectrograms support many AI analyses
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
🔄

Maintenance

How will the dataset be maintained?

1.0
ID
updates-1
Name
Future release plans
Description
  • Future releases aim to include original voice data (audio waveforms) with additional security precautions; pediatric cohort planned for future versions
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema