physionet b2ai-voice 1.1 d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • ID
    purpose-1
    Name
    Purpose
    Response
    Create an ethically sourced, diverse voice dataset linked to health information to enable AI research on voice as a biomarker and support clinical insights across multiple disease areas.
GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health3OT2OD032720-01S1
πŸ“Š

Composition

What do the instances represent?

CountsData TypeIDInstance TypeMissing InformationNameRepresentationSampling Strategies
12523Derived features only: spectrograms (513Γ—N), MFCCs (60Γ—N), static acoustic features (one row per rec...inst-recordingsRecordings (per session and task){'id': 'miss-1', 'name': 'Missing Info', 'missing': ['Original audio waveforms', 'Free speech transcripts'], 'why_missing': ['Privacy protection and de-identification; reduce re-identification risk']}Recording-derived instancesDerived data from voice recordings (e.g., spectrograms, MFCCs, static features){'id': 'samp-1', 'name': 'Sampling Strategy', 'is_sample': ['yes'], 'is_random': ['no'], 'source_data': ['Patients at specialty clinics across five North American sites'], 'is_representative': ['unknown'], 'representative_verification': ['not specified'], 'why_not_representative': ['Participants selected based on membership in predefined disease cohorts rather than population sampling.\n'], 'strategies': ['Deterministic clinical cohort enrollment per standardized protocol']}
306Demographics, clinical variables, validated questionnaires (phenotype.tsv/json)inst-participantsParticipantsParticipant instancesParticipants with linked phenotype and questionnaire data
DistributionIDIdentificationName
Not specifiedsubpop-1Adult participants in specialty clinicsAdult cohort
Not specifiedsubpop-2Voice disorders, Neurological and neurodegenerative disorders, ... (+3 more)Disease cohorts
  • ID
    distfmt-1
    Name
    Distribution Formats
    Description
    • Parquet (.parquet)
      spectrograms.parquet, mfcc.parquet
    • Tsv (.tsv)
      phenotype.tsv, static_features.tsv
    • Json (.json)
      phenotype.json, static_features.json
  • ID
    distdate-1
    Name
    Distribution Date
    Description
    • 2025-01-17
πŸ”

Collection Process

How was the data acquired?

bridge2ai-voice-v1.1
Bridge2AI-Voice v1.1
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a comprehensive, ethically sourced dataset derived from voice recordings linked to clinical information to enable research on voice as a biomarker of health. Version 1.1 provides derived data (e.g., spectrograms, MFCCs, static acoustic features) and rich phenotype information for a cohort initially comprising adults. The initial release (v1.0) included 12,523 recordings from 306 participants collected across five sites in North America, selected from five disease cohort categories (voice disorders, neurological/neurodegenerative disorders, mood/psychiatric disorders, respiratory disorders, and pediatric; v1.1 contains adult cohort data). To reduce re-identification risk, only derived data are shared in this version; raw audio waveforms are excluded. Data collection followed a standardized, multi-site protocol, with de-identification aligned to HIPAA Safe Harbor.
2025-01-17
  • voice
  • Bridge2AI
  • spectrograms
  • MFCC
  • bioacoustics
  • clinical data
  • derived features
  • registered access
  • HIPAA Safe Harbor
  • Parquet
  • phenotype
  • Alistair Johnson
  • Jean-Christophe BΓ©lisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anais Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
  • ID
    gap-1
    Name
    Addressing Gap
    Response
    Addresses the lack of large, high-quality, ethically sourced, multi-institutional, and demographically diverse voice datasets linked to clinical and questionnaire data for reproducible AI research.
  1. ID
    samp-2
    Name
    Sampling Strategy (dataset level)
    Is Sample
    • yes
    Is Random
    • no
    Source Data
    • Specialty clinics across five North American sites
    Is Representative
    • unknown
    Representative Verification
    • not specified
    Strategies
    • Clinical cohort-based enrollment according to inclusion/exclusion criteria
  1. ID
    ext-1
    Name
    External Resources
    External Resources
    Future Guarantees
    • Raw audio via controlled access to protect participant privacy
    Archival
    • DOI-resolved project page on PhysioNet
    Restrictions
    • Registered access license and data use agreement required
  • ID
    conf-1
    Name
    Confidentiality
    Description
    • Contains clinical and demographic data; raw audio excluded to reduce risk; access controlled via registered access and DUA
  • ID
    sens-1
    Name
    Sensitive Elements
    Description
    • Health-related clinical data and demographics may be considered sensitive
ID
deid-1
Name
Deidentification
Description
  • HIPAA Safe Harbor identifiers removed (e.g., names, fine-grained dates, contact numbers, biometric identifiers, etc.)
  • State and province removed; country of data collection retained
  • Free speech transcripts removed
  • Audio waveforms omitted; only derived features are shared in this version
  1. ID
    acq-1
    Name
    Instance Acquisition
    Description
    • Voice recordings collected during clinic visits; phenotype gathered via standardized questionnaires
    Was Directly Observed
    yes
    Was Reported By Subjects
    yes
    Was Inferred Derived
    yes
    Was Validated Verified
    yes
  • ID
    coll-1
    Name
    Collection Mechanisms
    Description
    • Standardized multi-site protocol; custom tablet application for data capture; headset used when possible; data managed via REDCap and exported with open-source tooling
  • ID
    dc-1
    Name
    Data Collectors
    Description
    • Project investigators at specialty clinics across five North American sites
  • ID
    irb-1
    Name
    Ethical Review
    Description
    • Data collection and sharing approved by the University of South Florida Institutional Review Board
  • ID
    dpia-1
    Name
    Data Protection
    Description
    • Privacy protections include HIPAA Safe Harbor de-identification, omission of raw audio and free speech transcripts, and registered/controlled access to sensitive materials
  1. ID
    prep-1
    Name
    Audio preprocessing and feature extraction
    Description
    • Raw audio converted to mono and resampled to 16 kHz with a Butterworth anti-aliasing filter; derived spectrograms computed (25 ms window, 10 ms hop, 512-point FFT); 60 MFCCs extracted from spectrograms; static acoustic and prosodic/phonetic features computed with open-source tools; transcriptions generated using Whisper Large
    Used Software
    IDName
    sw-opensmileopenSMILE
    sw-praatPraat
    sw-parselmouthParselmouth (Python interface to Praat)
    sw-torchaudiotorchaudio
    sw-whisperOpenAI Whisper Large
    sw-b2aiprepb2aiprep (processing library)
  • ID
    clean-1
    Name
    De-identification and redaction
    Description
    • Removal of HIPAA Safe Harbor identifiers; removal of state/province; removal of free speech transcripts; exclusion of raw audio from shared dataset
  • ID
    lab-1
    Name
    Transcription and annotations
    Description
    • Transcriptions generated using OpenAI Whisper Large; free speech transcripts were not shared in this release to protect privacy
  • ID
    raw-1
    Name
    Raw Data Availability
    Description
    • Original raw audio data available via controlled access upon request to DACO@b2ai-voice.org; shared dataset contains derived features only
  • ID
    maint-1
    Name
    Maintainer
    Description
    • Hosted on PhysioNet by the MIT Laboratory for Computational Physiology
Partially (TSV and JSON metadata; Parquet arrays for derived features)
πŸš€

Uses

What (other) tasks could the dataset be used for?

  • ID
    task-1
    Name
    Intended Task
    Response
    Develop and evaluate AI methods for health-related inference from voice, including detection and characterization of conditions associated with acoustic changes (e.g., voice, neurological, mood, and respiratory disorders) using derived features (spectrograms, MFCCs) and phenotype data.
  • ID
    other-1
    Name
    Potential Uses
    Description
    • Benchmarking AI models for acoustic biomarker discovery, condition screening, and monitoring using derived voice features linked with phenotype data
  • ID
    fut-1
    Name
    Future Use Considerations
    Description
    • Use may be limited for tasks requiring raw waveforms or free speech transcripts due to privacy protections; v1.1 includes adults only which may affect generalizability to pediatric populations
ID
lic-terms-1
Name
License and Terms of Use
Description
  • Access Policy
    Only registered users who sign the specified data use agreement can access the files
  • License
    Bridge2AI Voice Registered Access License
  • Data Use Agreement
    Bridge2AI Voice Registered Access Agreement
  • Raw audio disseminated via controlled access to protect participant privacy
πŸ“€

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
ID
veracc-1
Name
Version Access
Description
  • Older versions listed on PhysioNet; files for version 1.1 are no longer available; latest version as of the page is 2.0.1
πŸ”„

Maintenance

How will the dataset be maintained?

1.1
ID
update-1
Name
Update Plan
Description
  • Version History
    v1.0 (initial release), v1.1 (added MFCCs)
  • Future Plan
    aim to include voice data in future releases with additional security precautions
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema