VOICE Dataset Documentation

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

IDResponse
purpose-001Integrate the use of voice as a biomarker of health in clinical care by generating a substantial mul...
purpose-002Create an ethically sourced flagship dataset to enable future research in artificial intelligence an...
purpose-003Establish standards, best practices, and guidelines for voice data collection and analysis to advanc...
DescriptionIDName
Funded through National Institutes of Health grant 3OT2OD032720-01S3 (Bridge2AI: Voice as a Biomarke...funder-001NIH Office of the Director
NIBIB supports PhysioNet managed by MIT Laboratory for Computational Physiology under NIH grant numb...funder-002National Institute of Biomedical Imaging and Bioengineering
πŸ“Š

Composition

What do the instances represent?

  • ID
    instance-001
    Description
    Adult participants presenting at specialty clinics and institutions across five sites in North America. Participants were selected based on membership to five predetermined disease cohort groups: Respiratory disorders, Voice disorders, Neurological disorders, Mood disorders, and Pediatric. As of v1.1, only data from the adult cohort is available. The initial release (v1.0) provides 306 participants with 12,523 recordings collected through standardized protocols.
    Instance Type
    Human participants recruited from specialty clinics at multi-institutional sites. Data collection conducted between 2022 and 2026 through IRB-approved protocols with informed consent.
DescriptionIDName
Participants with laryngeal disorders including laryngeal cancers, vocal fold paralysis, and benign ...subpop-001Voice Disorders cohort
Participants with conditions such as Alzheimer's disease, Parkinson's disease, stroke, and ALS exhib...subpop-002Neurological and Neurodegenerative Disorders cohort
Participants with depression, schizophrenia, bipolar disorders, and anxiety disorders showing vocal ...subpop-003Mood and Psychiatric Disorders cohort
Participants with respiratory conditions including pneumonia, COPD, heart failure, and obstructive s...subpop-004Respiratory Disorders cohort
Pediatric participants with voice and speech disorders including autism spectrum disorder and speech...subpop-005Pediatric cohort
Access UrlsDescriptionIDName
https://physionet.org/content/b2ai-voice/Spectrograms stored in Parquet format (spectrograms.parquet). Each element contains participant_id, ...format-001Parquet for spectrograms
https://physionet.org/content/b2ai-voice/Mel-frequency cepstral coefficients stored in Parquet format (mfcc.parquet). Contains 60xN dimension...format-002Parquet for MFCCs
https://physionet.org/content/b2ai-voice/Tab-delimited phenotype file (phenotype.tsv) with one row per unique participant. Contains demograph...format-003TSV for phenotype data
https://physionet.org/content/b2ai-voice/Tab-delimited static features file (static_features.tsv) with one row per unique recording. Contains...format-004TSV for static features
Contact DACO@b2ai-voice.orgOriginal raw audio waveforms available through controlled access only. Interested users contact DACO...format-005Raw audio (controlled access only)
πŸ”

Collection Process

How was the data acquired?

Bridge2AI-Voice
Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information
The Bridge2AI-Voice project seeks to create an ethically sourced flagship dataset to enable future research in artificial intelligence and support critical insights into the use of voice as a biomarker of health. The human voice contains complex acoustic markers which have been linked to important health conditions including dementia, mood disorders, and cancer. When viewed as a biomarker, voice is a promising characteristic to measure as it is simple to collect, cost-effective, and has broad clinical utility. This comprehensive collection provides voice recordings with corresponding clinical information from participants selected based on known conditions which manifest within the voice waveform including voice disorders, neurological disorders, mood disorders, and respiratory disorders. The dataset is designed to fuel voice AI research, establish data standards, and promote ethical and trustworthy AI/ML development for voice biomarkers of health. Data collection occurs through a multi-institutional collaborative effort using standardized protocols, custom smartphone applications, and rigorous ethical oversight. The initial release (v1.0) provides 12,523 recordings for 306 participants collected across five sites in North America, with derived features such as spectrograms, MFCCs, acoustic features, and clinical phenotype data. Raw audio data is available through controlled access to protect participant privacy.
en
  • voice biomarker
  • acoustic biomarker
  • Bridge2AI
  • voice AI
  • voice disorders
  • neurological disorders
  • neurodegenerative disorders
  • mood disorders
  • psychiatric disorders
  • respiratory disorders
  • pediatric voice disorders
  • speech disorders
  • Parkinson's disease
  • Alzheimer's disease
  • depression
  • schizophrenia
  • bipolar disorder
  • stroke
  • ALS
  • autism
  • speech delay
  • laryngeal cancer
  • vocal fold paralysis
  • pneumonia
  • COPD
  • heart failure
  • obstructive sleep apnea
  • spectrogram
  • MFCC
  • mel-frequency cepstral coefficients
  • OpenSMILE
  • Praat
  • Parselmouth
  • federated learning
  • ethical AI
  • multimodal health data
  • electronic health records
  • EHR
  • radiomics
  • genomics
  • FAIR principles
  • CARE principles
  • PhysioNet
  • Health Data Nexus
IDResponse
gap-001Address the lack of large, high quality, multi-institutional and diverse voice databases linked to m...
gap-002Overcome limitations in existing voice and psychiatric disorder research that has relied on small da...
gap-003Fill the gap in pediatric voice and speech analysis research, which is sparser partly due to ethical...
gap-004Establish missing standards for voice data collection, acoustic analysis, and ethical frameworks for...
RoleNameORCIDAffiliation
ContributorYael Bensoussancreator-001-
ContributorJean-Christophe BΓ©lisle-Piponcreator-002-
ContributorDavid Dorrcreator-003-
ContributorSatrajit Ghoshcreator-004-
ContributorPhilip R.O. Paynecreator-005-
ContributorMaria Ellen Powellcreator-006-
ContributorAnais Rameaucreator-007-
ContributorVardit Ravitskycreator-008-
ContributorAlexandros Sigarascreator-009-
ContributorOlivier Elementocreator-010-
ContributorAlistair Johnsoncreator-011-
ContributorJennifer Siucreator-012-
ContributorBridge2AI-Voice Consortiumcreator-013-
DescriptionIDName
Contains derived features from voice recordings including spectrograms, MFCCs, acoustic features (Op...subset-001Public Access Dataset (PhysioNet Registered Access)
Original raw audio waveforms available through controlled access only to protect participant privacy...subset-002Controlled Access Raw Audio Dataset
  1. ID
    sampling-001
    Description
    Patients presenting at specialty clinics and institutions were screened for inclusion and exclusion criteria prior to their visit by project investigators. Participants were selected based on membership to five predetermined disease cohort groups to ensure representation across conditions affecting voice: (1) Voice Disorders - laryngeal cancers, vocal fold paralysis, benign laryngeal lesions; (2) Neurological and Neurodegenerative Disorders - Alzheimer's, Parkinson's, stroke, ALS; (3) Mood and Psychiatric Disorders - depression, schizophrenia, bipolar disorders; (4) Respiratory disorders - pneumonia, COPD, heart failure, obstructive sleep apnea; (5) Pediatric diseases - autism, speech delay.
    Is Sample
    • True
    Is Random
    • False
    Is Representative
    • False
    Strategies
    • Targeted recruitment from specialty clinics representing five disease cohort categories
    • Screening based on known conditions manifesting within voice waveform
    • Multi-institutional enrollment across five sites in North America
    • Standardized inclusion and exclusion criteria applied by investigators
DescriptionID
Data collection conducted using a custom smartphone application on tablet with headset used when pos...collection-001
Multi-institutional data collection across five specialty clinic sites in North America. Patients pr...collection-002
Data collection protocol involved: (1) demographic information collection, (2) health questionnaires...collection-003
DescriptionID
Voice recording tasks capturing voice, speech, and language data relating to health. Tasks include s...acquisition-001
Self-reported demographic and medical history questionnaires completed by participants who consent. ...acquisition-002
Electronic health record (EHR) access for participants who consent, permitting investigators to acce...acquisition-003
DescriptionIDPreprocessing Details
Raw audio preprocessing by converting to monaural and resampling to 16 kHz with Butterworth anti-ali...preproc-001Conversion to monaural audio, Resampling to 16 kHz sampling rate, ... (+2 more)
Spectrogram extraction - Time-frequency representations computed using short-time Fast Fourier Trans...preproc-002Short-time FFT with 25ms window, 10ms hop length, ... (+2 more)
Mel-frequency cepstral coefficients (MFCC) extraction - 60 MFCCs extracted from spectrograms. MFCCs ...preproc-00360 MFCC coefficients extracted, Derived from spectrograms, ... (+2 more)
Acoustic feature extraction using OpenSMILE (Speech and Music Interpretation by Large-space Extracti...preproc-004OpenSMILE feature extraction, Temporal dynamics captured, ... (+2 more)
Phonetic and prosodic feature computation using Parselmouth and Praat, providing measures of fundame...preproc-005Parselmouth and Praat feature extraction, Fundamental frequency (F0) measurement, ... (+2 more)
Transcription generation using OpenAI's Whisper Large model. Automated speech recognition applied to...preproc-006OpenAI Whisper Large model, Automated transcription of audio, ... (+2 more)
Data export and conversion from REDCap using open source b2aiprep library developed by the team. Phe...preproc-007REDCap data export, b2aiprep library conversion, ... (+2 more)
Cleaning DetailsDescriptionID
HIPAA Safe Harbor compliance, 18 identifier categories removed, ... (+3 more)HIPAA Safe Harbor de-identification applied. Identifiers removed include: names, geographic locators...cleaning-001
Audio waveforms excluded from public release, Derived features only in public dataset, ... (+2 more)Privacy protection measures for public release - Audio waveforms omitted from public dataset, only d...cleaning-002
Standardized collection protocols, Common smartphone application, ... (+2 more)Data standardization across multi-institutional sites through use of standardized protocols, common ...cleaning-003
DescriptionIDMaintainer DetailsName
Multidisciplinary consortium responsible for dataset maintenance including data collection, curation...maintainer-001University of South Florida (lead institution), Multi-institutional data collection sites (five sites in North America), ... (+5 more)Bridge2AI-Voice Consortium
PhysioNet platform managed by MIT Laboratory for Computational Physiology serves as primary distribu...maintainer-002PhysioNet / MIT Laboratory for Computational Physiology
ID
retention-001
Name
Data retention and disposition
Description
Data Transfer and Use Agreement specifies retention requirements. Upon termination or expiration of agreement (two years after start date, project completion, or ethics approval expiration), data shall be destroyed per provider instructions with written certification required within 30 days. Recipient may retain one copy to extent necessary to comply with records retention requirements under law, regulation, institutional policy, and for research integrity and verification purposes. Restrictions apply to archival copies as long as recipient holds data.
Retention Details
  • Two-year agreement term from start date
  • Data destruction required upon termination unless retention justified
  • One archival copy permitted for compliance and verification
  • Written certification of destruction required within 30 days
  • Ongoing restrictions apply to retained copies
  • Provider may unilaterally amend if federal sponsor requires
DescriptionIDSensitive Elements PresentSensitivity Details
Voice recordings contain personally identifiable information and are considered biometric identifier...sensitive-001TrueVoice as biometric identifier, Raw audio waveforms, ... (+2 more)
Electronic health record (EHR) data accessed with participant consent for gold standard validation o...sensitive-002TrueEHR medical information, Diagnoses and symptoms, ... (+2 more)
Demographic information and geographic data collected but de-identified for public release. State an...sensitive-003TrueDemographic data (de-identified), Geographic information (state/province removed), ... (+2 more)
Dataset covered under Certificate of Confidentiality which must be asserted against compulsory legal...sensitive-004TrueCertificate of Confidentiality coverage, Protection against compulsory legal demands, ... (+2 more)
DescriptionExternal ResourcesIDName
Primary distribution platform for public access dataset with registered accesshttps://physionet.org/content/b2ai-voice/resource-001PhysioNet Dataset Landing Page
Comprehensive project documentation and resourceshttps://docs.b2ai-voice.orgresource-002Bridge2AI-Voice Project Documentation
Open source code repository including b2aiprep library and documentation dashboardhttps://github.com/eipm/bridge2ai-docsresource-003Bridge2AI-Voice GitHub Repository
Federal grant information and project detailshttps://reporter.nih.gov/project-details/11376382resource-004NIH RePORTER Project Details
Alternative data repository platformhttps://healthdatanexus.ai/content/b2ai-voice/1.0/resource-005Health Data Nexus
Additional dataset documentation and software releaseshttps://doi.org/10.5281/zenodo.13834653resource-006Zenodo Archive
Research resource for complex physiologic signalshttps://physionet.orgresource-007PhysioNet Platform
Publication describing multi-disorder voice protocol development through team science approach invol...https://doi.org/10.21437/Interspeech.2024-1926resource-008Interspeech 2024 Protocol Publication
Contact for controlled access to raw audio datamailto:DACO@b2ai-voice.orgresource-009Data Access Compliance Office
Parent NIH Common Fund program supporting AI-ready biomedical datasetshttps://bridge2ai.orgresource-010Bridge2AI Program
Open source library for preprocessing raw audio and phenotype datahttps://github.com/sensein/b2aiprepresource-011b2aiprep Software Library
πŸš€

Uses

What (other) tasks could the dataset be used for?

IDResponse
task-001Enable development of AI/ML predictive models for screening, diagnosis, and treatment of voice disor...
task-002Support machine learning models for neurological and neurodegenerative disorders including Alzheimer...
task-003Develop AI algorithms for mood and psychiatric disorder detection including depression, schizophreni...
task-004Create machine learning models for respiratory disorder screening and therapeutic monitoring using r...
task-005Build AI models for pediatric voice and speech disorder detection including autism spectrum disorder...
task-006Promote application of AI/ML for voice research through workforce development, curriculum creation, ...
DescriptionID
Primary intended use is development and validation of AI/ML models for voice as a biomarker of healt...use-001
Research into acoustic biomarkers and development of standards for voice data collection and analysi...use-002
Training and education in voice AI research through workforce development initiatives, curriculum cr...use-003
Multimodal health research combining voice data with EHR information, radiomics, genomics, and other...use-004
Model dataset for ethical AI development in healthcare, demonstrating integration of bioethics guida...use-005
RoleNameORCIDAffiliation
Contributordiscouraged-001-
Contributordiscouraged-002-
Contributordiscouraged-003-
Contributordiscouraged-004-
ID
license-001
Name
Bridge2AI Voice Registered Access License
Description
Public access dataset distributed through PhysioNet under Bridge2AI Voice Registered Access License. Only registered users who sign the specified Data Use Agreement (Bridge2AI Voice Registered Access Agreement) can access files. Data covered under Certificate of Confidentiality which must be asserted against compulsory legal demands. Raw audio data available through controlled access only via Data Access Compliance Office (DACO) requiring distinct application. Recipient must adhere to PhysioNet requirements managed by MIT Laboratory for Computational Physiology, supported by NIBIB under grant R01EB030362.
License Terms
  • Registered access required
  • Data Use Agreement signature mandatory
  • Use restricted to authorized persons listed in agreement
  • No sharing with third parties without prior written consent
  • Appropriate administrative, technical, physical safeguards required
  • Compliance with applicable laws, rules, regulations, professional standards
  • Public disclosure of results encouraged in open-access journals
  • Recognition of data source required in publications
  • Certificate of Confidentiality protections apply
  • Raw audio requires separate controlled access application
  • PhysioNet platform requirements apply
  • Two-year term from start date or project completion
πŸ“€

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
πŸ”„

Maintenance

How will the dataset be maintained?

ID
updates-001
Name
Versioned releases with ongoing data collection
Description
Dataset updated with versioned releases as data collection progresses. Initial release v1.0 published January 17, 2025 with 12,523 recordings from 306 participants. v1.1 released January 17, 2025 adding MFCC features. v2.0.0 released April 16, 2025. v2.0.1 released August 18, 2025. Latest version available at https://doi.org/10.13026/37yb-1t42. Data collection ongoing through November 30, 2026. Version-specific documentation maintained. As of v1.1, only adult cohort data available; pediatric cohort data planned for future releases with additional privacy precautions.
Frequency
Periodic versioned releases during data collection period (2022-2026)
Update Details
  • v1.0 released January 17, 2025 - initial release with 306 participants, 12,523 recordings
  • v1.1 released January 17, 2025 - added MFCC features
  • v2.0.0 released April 16, 2025 - expanded participant cohort
  • v2.0.1 released August 18, 2025 - latest version
  • Ongoing data collection through November 30, 2026
  • Future releases planned with additional participants and pediatric cohort
  • Raw audio data access planned for future releases with additional security precautions
  • Version-specific documentation maintained
  • DOI for latest version vs version-specific DOIs
πŸ‘₯

Human Subjects

Does the dataset relate to people?

ID
hsr-001
Name
Bridge2AI-Voice Human Subjects Research
Description
Data collection and sharing approved by University of South Florida Institutional Review Board. Participants provided written informed consent for data collection initiative and data sharing. Consent process includes authorization for voice data collection, access to medical information through EHR platforms for gold standard validation, and permission to share research data. Bioethics guidance integrated throughout study design and conduct. Ethics module develops new guidelines for consenting to voice data collection, voice data sharing, and utilization in context of voice AI technology. Project addresses ethical and trustworthy issues from voice data generation and AI/ML research through clinical adoption and downstream health decisions.
Involves Human Subjects
True
IRB Approval
  • University of South Florida Institutional Review Board approval
Ethics Review Board
  • University of South Florida Institutional Review Board
Generated on 2025-12-20 19:23:28 using Bridge2AI Data Sheets Schema