VOICE all combined d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • Name
    primary-purpose
    Response
    Create a large, diverse, ethically sourced, multi-institutional dataset to enable AI research on voice as a biomarker of health, linked to clinical and demographic information.
GrantorGrant NameGrant Number
National Institutes of Health (NIH)Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioacoustic database to understand disease like never before3OT2OD032720-01S1
📊

Composition

What do the instances represent?

  • Name
    distribution-formats
    Description
    • application/x-parquet
    • text/tab-separated-values
    • application/json
DescriptionName
2024-11-27v1.0-release
2025-01-17v1.1-release
  1. Name
    recordings-and-phenotype
    Representation
    Derived representations from voice recordings (spectrograms, MFCCs, acoustic/phonetic/prosodic feature vectors) and participant-level phenotype (demographics, clinical data, validated questionnaires).
    Instance Type
    Participants and their recording sessions (participants, sessions, and recordings; one or more sessions per participant).
    Data Type
    Non-raw, derived audio representations (STFT-based spectrograms; MFCCs; engineered features via OpenSMILE, Parselmouth/Praat; torchaudio-derived features) plus tabular phenotype data and data dictionaries.
    Counts
    12,523
    Label
    Health condition cohorts, demographic attributes, and validated questionnaire responses; task metadata (e.g., task_name), participant_id, and session_id.
    Sampling Strategies
    1. Name
      cohort-based-enrollment
      Strategies
      • Deterministic cohort inclusion based on predefined disease groups at specialty clinics.
      Is Sample
      • True
      Is Random
      • False
      Source Data
      • Specialty clinics across five North American sites.
      Is Representative
      • Not statistically representative of the general population (cohort-based).
      Why Not Representative
      • Focused on predefined disorders to enable targeted biomarker research.
    Missing Information
    • Name
      dataset-level-omissions
      Missing
      • Raw audio waveforms
      • Free-speech transcripts
      Why Missing
      • Privacy, security, and de-identification considerations (HIPAA Safe Harbor; low-risk release design).
  • Name
    disease-cohorts
    Identification
    • Voice disorders
    • Neurological and neurodegenerative disorders
    • Mood and psychiatric disorders
    • Respiratory disorders
    • Pediatric voice and speech disorders (planned; adult cohort only in v1.0/v1.1)
    Distribution
    • Adult cohort only in v1.0 and v1.1; 306 participants across five North American sites.
🔍

Collection Process

How was the data acquired?

bridge2ai-voice
Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
Bridge2AI-Voice is a multi-site, ethically-sourced dataset linking derived data from human voice recordings to detailed demographic, clinical, and validated questionnaire information. The dataset enables research on voice as a biomarker of health across multiple condition cohorts (voice disorders, neurological and neurodegenerative disorders, mood and psychiatric disorders, respiratory disorders, and pediatric voice/speech disorders). The initial releases (v1.0, v1.1) contain low-risk derived data (e.g., spectrograms, MFCCs, engineered acoustic/phonetic/prosodic features) and phenotype tables; raw audio waveforms are not included in public distributions. As of v1.1, the dataset contains 12,523 recordings from 306 adult participants collected across five sites in North America under a standardized protocol, with de-identification following HIPAA Safe Harbor.
en
2025-01-17
  • voice
  • audio
  • bridge2ai
  • biomarker
  • health
  • spectrogram
  • mfcc
  • clinical
  • questionnaires
  • Alistair Johnson
  • Jean-Christophe Bélisle-Pipon
  • David Dorr
  • Satrajit Ghosh
  • Philip Payne
  • Maria Powell
  • Anaïs Rameau
  • Vardit Ravitsky
  • Alexandros Sigaras
  • Olivier Elemento
  • Yael Bensoussan
2024-11-27
Name
ip-restrictions
Description
  • Not specified.
Name
export-control
Description
  • None indicated.
  • Name
    dataset-hosts
    Description
    • Health Data Nexus (v1.0)
    • PhysioNet (v1.1 and later)
  • Name
    unmet-need
    Response
    Address the lack of large, diverse, standardized, multi-institutional voice datasets with linked clinical data and clear ethical frameworks; overcome small sample sizes, limited demographic diversity reporting, and inconsistent collection protocols in prior literature.
RoleNameORCIDAffiliation
ContributorBridge2AI-Voice Project Team-Bridge2AI Program (NIH Common Fund initiative)
  • Name
    participant-session-recording-linkage
    Description
    • Participant_id and session_id link phenotype rows to recording-derived features; multiple sessions may exist per participant.
  • Name
    recommended-splits
    Description
    • No official train/validation/test splits are provided in these releases.
  • Name
    known-issues
    Description
    • None reported in release notes.
  1. Name
    supporting-software-and-protocols
    External Resources
    Archival
    • Public repository for preprocessing code (b2aiprep).
    Restrictions
    • Not applicable to dataset access (software/resources are open; dataset itself is registered/credentialed access).
  • Name
    confidentiality
    Description
    • Dataset releases are designed as low risk; raw audio and free-speech transcripts are withheld; HIPAA Safe Harbor identifiers removed.
  • Name
    content-warnings
    Warnings
    • None stated.
  • Name
    sensitive-health-data
    Description
    • Health-related information (disease cohorts, questionnaires) and voice-derived features are potentially sensitive and handled under registered/credentialed access with DUA.
Name
de-identification
Description
  • HIPAA Safe Harbor identifiers removed (e.g., names, fine-grained dates, contact numbers, emails, IPs, SSNs, MRNs, plan IDs, device IDs, license/account/vehicle identifiers, URLs, full-face photos, biometric identifiers).
  • State and province removed; country of data collection retained.
  • Free-speech transcripts removed.
  • Public releases omit raw audio; only derived representations are distributed.
Mixed (tabular phenotype and dictionaries; array-based spectrograms/MFCCs; tabular engineered features)
  1. Name
    data-acquisition
    Description
    • Audio recording tasks (e.g., sustained vowel phonation) collected during clinic visits; phenotype captured concurrently.
    Was Directly Observed
    True
    Was Reported By Subjects
    True
    Was Inferred Derived
    True
    Was Validated Verified
    True
  • Name
    data-export
    Description
    • REDCap used for data capture and export; conversion and integration performed with open-source tooling (b2aiprep).
  • Name
    collection-teams
    Description
    • Project investigators and clinical teams at specialty clinics across five North American sites.
  • Name
    timeframe
    Description
    • Not explicitly stated; adult cohort collected prior to v1.0 (Nov 2024) and v1.1 (Jan 2025) releases.
  • Name
    hosting-and-access-ethics
    Description
    • Registered/credentialed access, signed DUA, and training (where required) used to ensure ethical data sharing and minimize risks.
  • Name
    data-protection
    Description
    • Dataset release strategy minimizes risk via HIPAA Safe Harbor de-identification and withholding of raw audio and free-speech transcripts; controlled/registered access with DUA.
DescriptionNameUsed Software
Monaural conversion; resampling to 16 kHz; Butterworth anti-aliasing filter.audio-standardization{'name': 'torchaudio', 'url': 'https://pytorch.org/audio'}, {'name': 'b2aiprep', 'url': 'https://github.com/sensein/b2aiprep'}
Short-time FFT spectrograms (25 ms window, 10 ms hop, 512-point FFT); MFCCs (60 coefficients, v1.1).spectral-representations{'name': 'librosa', 'url': 'https://librosa.org'}, {'name': 'torchaudio', 'url': 'https://pytorch.org/audio'}
Acoustic features via OpenSMILE; phonetic/prosodic features via Parselmouth/Praat.feature-engineering{'name': 'openSMILE', 'url': 'https://www.audeering.com/opensmile'}, {'name': 'Parselmouth', 'url': 'https://parselmouth.readthedocs.io'}, {'name': 'Praat', 'url': 'https://www.fon.hum.uva.nl/praat/'}
ASR via OpenAI Whisper Large for task transcriptions (free-speech transcripts not released).transcription{'name': 'OpenAI Whisper', 'url': 'https://github.com/openai/whisper'}
  • Name
    de-id-and-omissions
    Description
    • Removal of HIPAA Safe Harbor identifiers; removal of state/province; removal of free-speech transcripts; omission of raw audio from public releases.
  • Name
    questionnaires-and-task-labels
    Description
    • Validated questionnaires and clinical/phenotypic fields; recording task labels (e.g., task_name) associated with each session/recording.
  • Name
    raw-audio
    Description
    • Original audio waveforms exist but are not publicly distributed; controlled access requests can be directed to DACO@b2ai-voice.org (per PhysioNet notice).
Name
third-party-distribution
Description
Dataset is distributed to third parties under registered/credentialed access with a signed data use agreement; some hosts require documented training (e.g., TCPS 2: CORE 2022 on Health Data Nexus).
DescriptionIDMedia TypeNamePathTitle
Short-time FFT spectrograms (513 x N) derived from 16 kHz monaural audio.spectrograms-parquetapplication/x-parquetspectrograms.parquetspectrograms.parquetSpectrograms
60 Mel-frequency cepstral coefficients (60 x N) derived from spectrograms (added in v1.1).mfcc-parquetapplication/x-parquetmfcc.parquetmfcc.parquetMFCCs
One row per recording with features from OpenSMILE, Parselmouth/Praat, and torchaudio-derived measur...static-features-tsvtext/tab-separated-valuesstatic_features.tsvstatic_features.tsvEngineered acoustic/phonetic/prosodic features
Column-level metadata and descriptions for static_features.tsv.static-features-dictapplication/jsonstatic_features.jsonstatic_features.jsonData dictionary for engineered features
Participant-level demographics, acoustic confounders, clinical data, and validated questionnaire res...phenotype-tsvtext/tab-separated-valuesphenotype.tsvphenotype.tsvPhenotype table
Column-level metadata and descriptions for phenotype.tsv.phenotype-dictapplication/jsonphenotype.jsonphenotype.jsonData dictionary for phenotype
  • Name
    erratum
    Description
    • None specified.
🚀

Uses

What (other) tasks could the dataset be used for?

Name
registered-access-license
Description
  • Files are distributed under the Bridge2AI Voice Registered Access License and require acceptance of the Bridge2AI Voice Registered Access Agreement (DUA).
  • Access Is Restricted To Registered/credentialed Users Who Sign The Dua; On Health Data Nexus, Tcps 2
    CORE 2022 training is required for access.
  • Name
    target-tasks
    Response
    AI model development and evaluation for disease screening, risk stratification, and monitoring using voice-derived representations (classification, regression, and representation learning); methodological research on voice/speech biomarkers.
  • Name
    usage-tracking
    Description
    • Not specified; users are requested to cite the appropriate DOI for each used release (e.g., v1.0 on Health Data Nexus; v1.1 on PhysioNet).
  • Name
    citations
    Description
    • DOI (v1.0): https://doi.org/10.57764/qb6h-em84; DOI (v1.1): https://doi.org/10.13026/249v-w155
  • Name
    potential-uses
    Description
    • Multimodal fusion with clinical data; fair and robust model development; domain shift and cohort generalization studies.
  • Name
    risk-mitigation
    Description
    • Cohort-based design and site effects may impact generalizability; users should consider fairness, demographic diversity, and cohort balance in downstream tasks.
  • Name
    inappropriate-uses
    Description
    • Not specified in source; comply with DUA and ethical standards; avoid attempts at re-identification.
📤

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
Name
version-availability
Description
  • Older versions may be retained for citation purposes; some earlier-version files may no longer be downloadable after newer releases on hosting platforms (e.g., PhysioNet 2.x).
🔄

Maintenance

How will the dataset be maintained?

1.1
2025-01-17
Name
release-notes
Description
  • b2ai-voice v1.1 added Mel-frequency cepstral coefficients (MFCCs).
  • b2ai-voice v1.0 was the first public release with derived spectrograms, engineered features, and phenotype data.
Generated on 2025-11-09 10:17:35 using Bridge2AI Data Sheets Schema