VOICE d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

IDResponse
voice:purpose:1
Integrate the use of voice as a biomarker of health in clinical care by generating a substantial multi-institutional, ethically sourced, and diverse voice database linked to multimodal health biomarkers to fuel voice AI research and build predictive models to assist in screening, diagnosis, and treatment of a broad range of diseases.
voice:purpose:2
Create an ethically sourced flagship dataset of 10,000 voices linked to health information to enable future research in artificial intelligence and support critical insights into the use of voice as a biomarker of health, addressing the pressing need for large, high quality, multi-institutional and diverse voice databases linked to other health biomarkers.
voice:purpose:3
Establish standards, best practices, and guidelines for voice data collection and analysis to advance the field of acoustic biomarkers by developing new standards that are AI/ML friendly and enable voice to emerge as a biomarker of health.
voice:purpose:4
Address ethical, legal, and social challenges surrounding voice AI including risks of voice re-identification, vulnerabilities like voice AI hacking, concerns around voice data sharing and privacy, and the influence of gender and racial diversity on the development and application of voice AI technologies.
DescriptionID
National Institutes of Health (NIH) Common Fund Bridge2AI Program. Grant number: 3OT2OD032720-01S3. Opportunity Number: OTA-21-008. Project dates: September 1, 2022 to November 30, 2026. Total funding in 2025: $4,660,942 (Direct Costs: $4,072,321, Indirect Costs: $588,621). Administering Institute: NIH Office of the Director. Study Section: Data Coordination, Mapping, and Modeling (DCMM).
voice:funder:1
National Institute of Biomedical Imaging and Bioengineering (NIBIB). Supports PhysioNet managed by MIT Laboratory for Computational Physiology under NIH grant number R01EB030362, which serves as the primary distribution platform for the Bridge2AI-Voice dataset.
voice:funder:2
📊

Composition

What do the instances represent?

  1. ID
    voice:instance:1
    Description
    Adult participants presenting at specialty clinics (high volume expert clinics) across multiple sites in North America. Participants selected based on membership to five predetermined disease cohort groups: Voice Disorders, Neurological and Neurodegenerative Disorders, Mood and Psychiatric Disorders, Respiratory Disorders, and Pediatric Voice and Speech Disorders. Version 3.0 contains approximately 833 adult participants with ~61,937 voice-derived recordings. Pediatric dataset v1.0 adds 300 participants. Enrollment anticipated to reach 10,000 participants by 2027.
    Instance Type
    Human participants with clinical diagnoses recruited from specialty clinics
    Counts
    833
    Label
    True
    Label Description
    Diagnostic labels assigned by clinical assessment at each site. Clinicians provided diagnoses based on clinical interview and appropriate work-up including laryngoscopy, stroboscopy, MRI, CT, whole genome sequencing, EHR records, and medication prescriptions. Labels include diagnostic categories: vocal pathologies, neurological disorders, psychiatric conditions, respiratory disorders, and pediatric voice/speech disorders. Per Bridge2AI protocols and ICD-10 codes. Single labeler per participant (site clinician).
DescriptionIDName
Participants with laryngeal disorders including laryngeal cancer (T1-T4, biopsy proven), laryngitis (acute, chronic, bacterial, fungal, autoimmune), pre-cancerous lesions (keratosis, leukoplakia), benign vocal cord lesions (nodules, polyps, cysts, Reinke's edema, recurrent laryngeal papilloma), muscle tension dysphonia, spasmodic dysphonia and laryngeal tremor, unilateral vocal fold paralysis, and glottic insufficiency/ presbyphonia. Validated by laryngoscopy images and stroboscopy videos.
voice:subpop:1Voice Disorders cohort
Participants aged 44-85 with clinical diagnoses including mild cognitive impairment, Alzheimer's disease, other dementias (frontotemporal, Lewy body, vascular, mixed, alcohol-induced), ALS (sporadic, familial, spinal/limb-onset, bulbar-onset), and Parkinson's disease (idiopathic PD, multiple system atrophy, progressive supranuclear palsy, corticobasal degeneration). Must be able to read and speak English. Validated by CT Brain, MRI Brain, serum biomarkers (Ptau proteins for AD), and whole genome sequencing. Participants enrolled in deep brain stimulation studies noted.
voice:subpop:2Neurological and Neurodegenerative Disorders cohort
Participants with clinical diagnoses including depression/major depressive disorder, bipolar I/II disorder, anxiety disorder, schizophrenia, ADHD (comorbid), PTSD (comorbid), OCD (comorbid), panic disorder (comorbid), borderline personality disorder (comorbid), eating disorder (comorbid), insomnia/sleep disorder (comorbid), social anxiety disorder (comorbid), autism spectrum disorder (comorbid), alcohol/substance use disorder (comorbid), and other psychiatric disorders. Validated by electronic health summary, clinical diagnosis from psychiatrist, and medication history. Uses GAD-7, PHQ-9, PTSD, ADHD, DSM-5 adult questionnaires among others.
voice:subpop:3Mood and Psychiatric Disorders cohort
Participants with respiratory conditions including airway stenosis (bilateral vocal fold paralysis, supraglottic, glottic, posterior glottic, subglottic, tracheal stenosis, multi-level upper airway stenosis) and chronic cough (bothersome cough >8 weeks). Validated by spirometry, flow volume loops, and CT scan of neck/chest. Dyspnea Index and Leicester Cough Questionnaire administered.
voice:subpop:4Respiratory Disorders cohort
Pediatric participants aged 2-18 with voice and speech disorders recruited exclusively from Hospital for Sick Children (SickKids). Grouped by age: 2-4, 4-6, 6-10, 10+ years. Data collected using reproschema-ui with Bridge2AI-Voice pediatric protocol. Includes 36 pediatric-specific acoustic tasks and specialized questionnaires (C-VHI-10, PVOS, PVRQOL, PHQ-A). Pediatric dataset v1.0 available separately from adult dataset.
voice:subpop:5Pediatric cohort
Healthy volunteer control participants who complete common questionnaires and voice tasks (Voice Handicap Index-10, PHQ-9, GAD-7, Winograd) to provide normative comparisons. voice:subpop:6Control participants
DescriptionIDName
Time-varying features (spectrograms, mel spectrograms, MFCCs, pitch contour, SPARC features, PPGs) stored in Parquet format compatible with Python datasets library. Each element contains participant_id, session_id, task_name, and feature arrays. Excluded for open-response tasks in public release.
voice:format:1Parquet format for spectrograms and derived features
Phenotype data in tab-delimited format (confounders.tsv, demographics.tsv, diagnosis/ *.tsv, enrollment/*.tsv, questionnaire/*.tsv, task/*.tsv) with JSON data dictionaries per file. Static acoustic features (OpenSMILE, Praat, Parselmouth, torchaudio) in static_features.tsv with static_features.json data dictionary. One row per participant or recording as appropriate. BIDS v1.9.0 compliant structure.
voice:format:2TSV/JSON for phenotype and static features
Original raw audio waveforms in WAV format following BIDS structure: sub-{participant_id}/ses-{session_id}/audio/sub{id}_ses{id}_task-{task_name}.wav with companion JSON metadata. Available through controlled access only. Contact DACO@b2ai-voice.org to request access.
voice:format:3WAV audio format (controlled access only)
  • ID
    voice:distdate:1
    Description
    Dataset first published and made available late November 2024 through Health Data Nexus. PhysioNet releases: v1.0 January 17, 2025 (306 participants, 12,523 recordings); v1.1 January 17, 2025 (added MFCC features); v2.0.0 April 16, 2025; v2.0.1 August 18, 2025; v3.0.0 released 2025 (833 adult participants, ~61,937 recordings). Semi-annual releases planned. Pediatric v1.0 available separately.
  • ID
    voice:privacy:1
    Description
    Multiple layered privacy protections: (1) HIPAA Safe Harbor de-identification; (2) removal of all raw audio from public dataset; (3) removal of open-response features from public dataset; (4) Certificate of Confidentiality; (5) registered access requiring DUA signature; (6) controlled access for raw audio requiring institutional DTUA; (7) federated learning technology planned for multi-institutional analysis minimizing data sharing; (8) secure storage requirements in DTUA (administrative, physical, and technical safeguards). No analysis of potential re-identification impact conducted.
RoleNameORCIDAffiliation
Contributorvoice:compensation:1-
🔍

Collection Process

How was the data acquired?

Bridge2AI-Voice
Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information
The Bridge2AI-Voice project seeks to create an ethically sourced flagship dataset to enable future research in artificial intelligence and support critical insights into the use of voice as a biomarker of health. The human voice contains complex acoustic markers which have been linked to important health conditions including dementia, mood disorders, and cancer. When viewed as a biomarker, voice is a promising characteristic to measure as it is simple to collect, cost-effective, and has broad clinical utility. This comprehensive collection provides voice recordings with corresponding clinical information from participants selected based on known conditions which manifest within the voice waveform including voice disorders, neurological disorders, mood disorders, and respiratory disorders. The dataset is designed to fuel voice AI research, establish data standards, and promote ethical and trustworthy AI/ML development for voice biomarkers of health. Data collection occurs through a multi-institutional collaborative effort using standardized protocols, custom smartphone applications, and rigorous ethical oversight. Version 3.0 provides approximately 61,937 voice-derived recordings from 833 adult participants collected across multiple sites in North America, with derived features such as spectrograms, MFCCs, acoustic features, and clinical phenotype data. The pediatric dataset v1.0 is also available with data from 300 participants. Raw audio data is available through controlled access to protect participant privacy.
en
Bensoussan, Yael, et al. "Developing Multi-Disorder Voice Protocols: A team science approach involving clinical expertise, bioethics, standards, and DEI." Proc. Interspeech 2024. 2024. https://doi.org/10.21437/Interspeech.2024-1926
  • voice biomarker
  • acoustic biomarker
  • Bridge2AI
  • voice AI
  • voice disorders
  • neurological disorders
  • neurodegenerative disorders
  • mood disorders
  • psychiatric disorders
  • respiratory disorders
  • pediatric voice disorders
  • speech disorders
  • Parkinson's disease
  • Alzheimer's disease
  • depression
  • schizophrenia
  • bipolar disorder
  • stroke
  • ALS
  • autism spectrum disorder
  • speech delay
  • laryngeal cancer
  • vocal fold paralysis
  • muscle tension dysphonia
  • laryngeal dystonia
  • COPD
  • chronic cough
  • airway stenosis
  • obstructive sleep apnea
  • spectrogram
  • MFCC
  • mel-frequency cepstral coefficients
  • OpenSMILE
  • Praat
  • Parselmouth
  • torchaudio
  • federated learning
  • ethical AI
  • multimodal health data
  • electronic health records
  • EHR
  • radiomics
  • genomics
  • FAIR principles
  • CARE principles
  • PhysioNet
  • Health Data Nexus
  • BIDS
  • Brain Imaging Data Structure
  • b2aiprep
False
IDResponse
voice:gap:1
Address the lack of large, high quality, multi-institutional and diverse voice databases linked to multimodal health biomarkers (demographics, imaging, genomics, risk factors) necessary to fuel voice AI research and answer tangible clinical questions. Previous studies had sample sizes too small or lacked metadata needed for robust, clinically useful models.
voice:gap:2
Overcome limitations in existing voice and psychiatric disorder research that has relied on small datasets with limited demographic diversity reporting, lack of standardized data collection protocols precluding meta-analysis, and possible confounders limiting external validity and clinical usability.
voice:gap:3
Fill the gap in pediatric voice and speech analysis research, which is sparser partly due to ethical concerns and challenges in data acquisition for this cohort, particularly for autism spectrum disorder and speech delay detection.
voice:gap:4
Establish missing standards for voice data collection, acoustic analysis, and ethical frameworks for consenting to voice data collection, sharing, and utilization in the context of voice AI technology development and clinical adoption.
voice:gap:5
Develop software and cloud infrastructure for automated voice data collection through a smartphone application (Bridge2AI-Voice App) that allows non-invasive, user-friendly, high quality voice data collection while minimizing human manipulation and implementing federated learning technology to minimize data sharing while preserving patient privacy.
RoleNameORCIDAffiliation
Contributorvoice:creator:1-
RoleNameORCIDAffiliation
ContributorFeaturized Dataset (Registered Access)voice:subset:1-
ContributorRaw Audio Dataset (Controlled Access)voice:subset:2-
ContributorPediatric Dataset (Registered Access)voice:subset:3-
  1. ID
    voice:sampling:1
    Description
    Non-probability purposive sampling from specialty clinics (high volume expert clinics - outpatient clinics seeing >50 patients per month from same disease category). Patients presenting at clinics screened for eligibility per inclusion/exclusion criteria outlined in protocol Table 1. Not representative of general population due to limited geographic locations and clinic-based recruitment. Current v3.0 dataset is a sample of an ongoing collection targeting 10,000 participants by 2027. English-speaking adult participants (18-120 years); Spanish protocols under development for future releases.
    Sample
    True
    Random Sampling
    False
    Representative Sample
    False
    Why Not Representative
    • Data collected at limited number of geographic locations (five sites)
    • Clinic-based recruitment introduces selection bias toward treatment-seeking populations
    • Remote data collection not included in initial releases
    • Groups with less trust in medical system or less proximal to collection sites underrepresented
    • Current releases contain only English-speaking participants
    • Public releases do not contain equal distribution across disease categories
    Strategies
    • Targeted recruitment from high volume expert specialty clinics
    • Inclusion/exclusion criteria per Bridge2AI-Voice protocol Table 1 for each disease cohort
    • Multi-institutional enrollment across North American sites
    • IRB-approved prospective consent process
    • Pediatric participants exclusively from Hospital for Sick Children (SickKids)
DescriptionID
Voice data collected in clinic using custom Bridge2AI-Voice App on iPad (9th or 10th generation) or iPad Air (5th generation) with Avid AE-36 microphone and Apple dongle connector. App collects breathing sounds and voice, speech, and linguistic tasks along with health information through surveys and validated questionnaires. Research assistant present during collection. Future remote data collection planned but not included in current releases.
voice:collection:1
Pediatric data collected using reproschema-ui with Bridge2AI-Voice pediatric protocol. Same iPad hardware as adult protocol with age-appropriate tasks grouped by age range (2-4, 4-6, 6-10, 10+ years). Questions read to participants when needed.
voice:collection:2
Clinical data including EHR information, imaging (laryngoscopy, stroboscopy, MRI, CT), and genomic data extracted from sites independently and uploaded through REDCap database. No external multimodal data released in current dataset versions.
voice:collection:3
DescriptionID
22 acoustic tasks recorded through Bridge2AI-Voice App for adult participants including: Non-Voice (Respiration, Cough, Breath Sounds, Voluntary Cough); Voice/Non-Speech (Prolonged Vowel /e/, Maximum Phonation Time, Glides, Loudness /Hey/, Diadochokinesis /pa/ta/ka/); Speech (Rainbow Passage, Caterpillar Passage, Cape-V Sentences, Free Speech, Picture Description, Story Recall, Animal Fluency, Open Response Questions, Word-Color Stroop, Productive Vocabulary, Random Item Generation, Cinderella Story). Tasks vary by disease cohort (Part A/Voice/Resp/Mood/Neuro).
voice:acquisition:1
36 pediatric acoustic tasks collected via reproschema-ui including: Speech tasks (ABC's, Ready for School, Favorite Show, Favorite Food, Outside of School, Months, Counting, Naming Animals, Naming Food, Identifying Pictures, Picture Description, Caterpillar Passage, Repeat Words, Role Naming, Repeat Sentences); Voice/Non-Speech tasks (Long Sounds, Noisy Sounds, Silly Sounds /PUH TUH KUH/).
voice:acquisition:2
Self-reported demographic data and medical history questionnaires administered through app. Validated questionnaires integrated for each disease cohort including: VHI-10, PHQ-9, GAD-7, PANAS, Custom Affect Scale, PTSD Adult, ADHD Adult, DSM-5 Adult, Dyspnea Index, Leicester Cough Questionnaire, Winograd Questionnaire, MOCA, C-VHI-10 (peds), PVOS (peds), PVRQOL (peds), PHQ-A (peds). Confounders questionnaire about smoking history, drinking history, and other acoustic confounders.
voice:acquisition:3
Electronic health record (EHR) access for consenting participants permitting investigators to access medical information through EHR platforms to perform gold standard validation of diagnoses and symptoms. Linkage to multimodal health biomarkers including laryngoscopy imaging, radiomics, genomics, respiratory function tests.
voice:acquisition:4
  • ID
    voice:datacollector:1
    Description
    Research teams at each of the five North American collection sites, including medical graduate, and undergraduate students coordinating with site clinicians and doctors. Clinicians and doctors listed under IRB as co-investigators and added to consortium. Participants compensated via electronic gift cards: $40 for sessions under 90 minutes, $80 for sessions over 90 minutes, maximum 3 sessions and $120 total compensation.
  • ID
    voice:timeframe:1
    Description
    Data collection ongoing from September 2022 through November 2026 (project end date). Initial dataset releases began in late 2024. Version 1.0 released January 17, 2025 (306 participants). Version 3.0 released 2025 (833 adult participants). Semi-annual releases planned. Enrollment anticipated to reach 10,000 participants by 2027.
  • ID
    voice:directcoll:1
    Description
    Data collected directly from individuals in clinic settings by trained research assistants using the Bridge2AI-Voice App. Participants notified and consented through IRB-approved process. Clinical diagnoses and EHR data obtained with participant consent from site clinicians. Data collected in USA and Canada.
DescriptionID
Raw audio preprocessing using b2aiprep library: conversion to monaural audio, resampling to 16 kHz with Butterworth anti-aliasing filter. Standardization ensures consistent format across all recordings for downstream feature extraction.
voice:preproc:1
Spectrogram extraction using short-time Fast Fourier Transform (FFT): 25ms window size, 10ms hop length, 512-point FFT. Output spectrograms have 513xN dimensions where N is proportional to audio length. Stored in Parquet format.
voice:preproc:2
Mel-frequency cepstral coefficients (MFCC) extraction: 60 MFCCs extracted from spectrograms capturing perceptually-relevant spectral envelope characteristics. Output dimension 60xN. Stored in Parquet format.
voice:preproc:3
OpenSMILE eGeMaps acoustic feature extraction capturing temporal dynamics and acoustic characteristics. One row per unique recording in static_features.tsv. voice:preproc:4
Parselmouth and Praat phonetic and prosodic feature computation providing fundamental frequency (F0), formants, and voice quality measures. Documented in static_features.json data dictionary. voice:preproc:5
Torchaudio-based feature extraction including pitch contour, spectrograms, mel spectrograms, MFCCs. SPARC-based features including electromagnetic articulography (EMA) estimates, loudness, periodicity, and pitch measures. Phonetic posteriorgrams (PPGs). Stored in two formats: fixed static format and temporal format varying by recording length.
voice:preproc:6
Transcription generation using OpenAI Whisper model applied to audio recordings. Raw audio transcripts reviewed and any recordings containing potentially identifying information or external voices removed. All transcriptions, EMAs, and PPGs from open-response prompts removed from public feature-only dataset.
voice:preproc:7
REDCap and reproschema-ui data exported and converted to BIDS v1.9.0 format using b2aiprep open-source library. Phenotype data organized in tab-delimited files with JSON data dictionaries. Pediatric data first extracted from reproschema-ui to REDCap format then converted to BIDS.
voice:preproc:8
DescriptionID
HIPAA Safe Harbor de-identification for public release: removed direct identifiers (names, civic addresses, social security numbers), indirect identifiers creating significant re-identification risk (select geographic/demographic identifiers, household composition, cultural identity), and sensitive information (household income, mental health status, traumatic life experiences). Geographic data limited; state/province removed, country retained.
voice:cleaning:1
Audio privacy protection for public release: all raw audio waveforms excluded from public dataset. All spectrograms, MFCCs, mel spectrograms, transcriptions, EMAs, and PPGs from open-response prompts removed from feature-only dataset. Free speech transcripts removed. Raw audio available only through controlled access with DACO approval. Raw data stored and retained separately for verified researchers.
voice:cleaning:2
Sensitive field removal based on REDCap data dictionary: all fields encoded as sensitive (column "Identifier?" in REDCap data dictionary CSV) removed from dataset. voice:cleaning:3
Audit protocol applied: missingness tables generated and included with dataset; distribution and outlier checks; categorical responses checked against schema; audio quality control metrics including silence amount, duration, and speech-to-passage accuracy checks. Processing using b2aiprep and SenseLab toolkits.
voice:cleaning:4
  • ID
    voice:labeling:1
    Description
    Diagnostic labels assigned by site clinicians based on clinical assessment and gold standard validation methods per Bridge2AI-Voice protocol Table 1 and ICD-10 codes. For voice disorders: laryngoscopy and stroboscopy. For neurological disorders: CT Brain, MRI Brain, genome sequencing, serum biomarkers. For mood/psychiatric disorders: EHR records, psychiatrist diagnosis, medication history. For respiratory disorders: spirometry, flow volume loops, CT scan. Single labeler (site clinician) per participant. Future labels should include description of exact variables used for determination.
RoleNameORCIDAffiliation
Contributorvoice:thirdparty:1-
DescriptionID
Bridge2AI-Voice Consortium (University of South Florida, lead institution) supported by NIH Bridge2AI program. Contact: [email protected]. Responsible for dataset curation, standards development, ethics oversight, versioning, and updates. Data Access Compliance Office (DACO) manages controlled access applications. Contact: DACO@b2ai-voice.org.
voice:maintainer:1
PhysioNet (MIT Laboratory for Computational Physiology, supported by NIBIB NIH grant R01EB030362) serves as primary distribution platform for registered access dataset. Provides technical infrastructure and access management.
voice:maintainer:2
Health Data Nexus (Temerty Centre for Artificial Intelligence Research and Education in Medicine, T-CAIREM, University of Toronto) maintains alternative platform providing cloud compute alongside dataset. Earlier dataset versions available here. Contact: [email protected].
voice:maintainer:3
ID
voice:retention:1
Name
Data retention and disposition policy
Description
Data Transfer and Use Agreement (DTUA) specifies two-year agreement term from start date. Upon termination or expiration (two years, project completion, ethics approval expiration, or provider termination, whichever occurs first), data shall be destroyed per provider instructions with written certification required within 30 days. Recipient may retain one copy to comply with records retention requirements under law, regulation, institutional policy, and for research integrity and verification. Ongoing restrictions apply to retained copies. Provider may unilaterally amend agreement if federal sponsor requires revision. Health Data Nexus retains dataset as long as useful for research purposes, possibly indefinitely. Version-specific DOIs maintained for all historical versions.
ID
voice:extension:1
Name
Dataset extension and contribution mechanisms
Description
Derivative datasets can be published on Health Data Nexus referencing original source under same access conditions. Open-source repositories (b2aiprep, SenseLab) have discussion forums, issue pages, and pull request mechanisms for contributing improvements to preprocessing code. REDCap data dictionary available at https://github.com/eipm/bridge2ai-redcap with MIT license. Bridge2AI-Voice documentation at https://github.com/eipm/bridge2ai-docs. Future augmentations to voice collection protocol coordinated through consortium.
RoleNameORCIDAffiliation
Contributorvoice:consent:1-
  • ID
    voice:notification:1
    Description
    Participants notified and consented through IRB-approved consent process before data collection. Consent process described eligibility, voluntary participation, data uses, sharing mechanisms, privacy protections, and withdrawal rights. English language used for all communications in current releases.
  • ID
    voice:ethicalreview:1
    Description
    IRB review and approval obtained from University of South Florida Single IRB with subsite IRB approvals. Study reviewed and approved for human subjects research. Bioethics guidance integrated throughout study design and conduct by consortium bioethicists (Belisle-Pipon, Ravitsky). Ethics module develops new guidelines for consenting to voice data collection, voice data sharing, and utilization in context of voice AI technology. Project addresses ethical issues from voice data generation through clinical adoption and downstream health decisions.
RoleNameORCIDAffiliation
Contributorvoice:sensitive:1-
Contributorvoice:sensitive:2-
Contributorvoice:sensitive:3-
Contributorvoice:sensitive:4-
Contributorvoice:sensitive:5-
ID
voice:deident:1
Name
HIPAA Safe Harbor de-identification
Description
HIPAA Safe Harbor de-identification applied. All direct identifiers removed (names, civic addresses, social security numbers). Indirect identifiers removed where creating significant re-identification risk (geographic/demographic identifiers, household composition, cultural identity). Non-identifying sensitive information removed (household income, mental health status, traumatic life experiences). All raw voice data removed from public release. Sensitive fields per REDCap data dictionary removed. Direct identifiers removed: Yes. HIPAA de-identification rules applied: Yes. Dates rebased: Yes. Geographic information removed/generalized: Yes. Narrative text fields removed: Yes. K-anonymization: No.
ID
voice:atrisk:1
Name
Pediatric participants protection
Description
Pediatric participants (aged 2-18) require additional protections. Pediatric data collected exclusively at Hospital for Sick Children (SickKids) with age-appropriate protocols. Distinct IRB considerations for pediatric cohort. Pediatric questionnaires and acoustic tasks specifically designed for different age groups (2-4, 4-6, 6-10, 10+ years). Minimum age for adult dataset is 18 years. Pediatric dataset released separately with additional privacy precautions.
DescriptionExternal ResourcesIDName
Primary registered access distribution platform for adult and pediatric datasetshttps://physionet.org/content/b2ai-voice/voice:resource:1PhysioNet Dataset Landing Page
Comprehensive project documentation, collection methods, governance, healthsheethttps://docs.b2ai-voice.orgvoice:resource:2Bridge2AI-Voice Project Documentation
Source code for docs and dashboard at docs.b2ai-voice.org (MIT license)https://github.com/eipm/bridge2ai-docsvoice:resource:3Bridge2AI-Voice GitHub Documentation Repository
Open source library (Apache-2.0 license) for preprocessing raw audio waveforms into parquet files and merging source data into phenotype files https://github.com/sensein/b2aiprepvoice:resource:4b2aiprep Software Library
REDCap data dictionary and metadata (MIT license) at github.com/eipm/bridge2ai-redcaphttps://github.com/eipm/bridge2ai-redcapvoice:resource:5Bridge2AI REDCap Data Dictionary
Federal grant information for project 3OT2OD032720-01S3https://reporter.nih.gov/project-details/11376382voice:resource:6NIH RePORTER Project Details
Alternative platform providing cloud compute alongside earlier dataset versionshttps://healthdatanexus.ai/content/b2ai-voice/1.0/voice:resource:7Health Data Nexus
Bensoussan Y. et al. (2024). eipm/bridge2ai-redcap. Zenodo. https://zenodo.org/doi/10.5281/zenodo.12760724 https://doi.org/10.5281/zenodo.13834653voice:resource:8Zenodo Archive (REDCap data dictionary)
Bensoussan et al. "Developing Multi-Disorder Voice Protocols: A team science approach involving clinical expertise, bioethics, standards, and DEI." Proc. Interspeech 2024. https://doi.org/10.21437/Interspeech.2024-1926voice:resource:9Interspeech 2024 Protocol Publication
Contact for controlled access to raw audio data and institutional DTUAmailto:DACO@b2ai-voice.orgvoice:resource:10Data Access Compliance Office (DACO)
Parent NIH Common Fund program supporting AI-ready biomedical datasetshttps://bridge2ai.orgvoice:resource:11Bridge2AI Program
Training opportunities for using the datasethttps://www.b2aivoicescholars.org/voice:resource:12Bridge2AI Voice Scholars Training
Software toolkit for processing audio files for research taskshttps://github.com/sensein/senselabvoice:resource:13SenseLab Software
AI-readiness for Biomedical Data - Bridge2AI Recommendations documenthttps://docs.b2ai-voice.orgvoice:resource:14AI-Readiness Recommendations
🚀

Uses

What (other) tasks could the dataset be used for?

IDResponse
voice:task:1
Enable development of AI/ML predictive models for screening, diagnosis, and treatment of voice disorders including laryngeal cancers, vocal fold paralysis, muscle tension dysphonia, laryngeal dystonia, benign laryngeal lesions, and pre-cancerous lesions, leveraging acoustic changes in phonation resulting from changes in vocal fold vibratory function.
voice:task:2
Support machine learning models for neurological and neurodegenerative disorders including Alzheimer's disease, Parkinson's disease, mild cognitive impairment, other dementias, and ALS, detecting voice and speech changes such as slowed speech, low frequency, monotonous speech, vocal tremor, dysarthria, and aphasia. Validation uses CT, MRI, and whole genome sequencing data.
voice:task:3
Develop AI algorithms for mood and psychiatric disorder detection including depression, schizophrenia, bipolar disorder, anxiety disorders, ADHD, PTSD, OCD, and borderline personality disorder, identifying vocal markers through clinical diagnosis, EHR records, and medication history.
voice:task:4
Create machine learning models for respiratory disorder screening and therapeutic monitoring using respiratory sounds, cough sounds, and voice, applicable to conditions such as chronic cough, COPD, airway stenosis, and other respiratory conditions using spirometry, flow volume loops, and CT imaging for validation.
voice:task:5
Build AI models for pediatric voice and speech disorder detection using 36 pediatric-specific acoustic tasks and specialized questionnaires, addressing the relative scarcity of pediatric voice data for participants aged 2-18 from the Hospital for Sick Children (SickKids).
voice:task:6
Promote application of AI/ML for voice research through workforce development, curriculum creation on voice biomarkers of health for FAIR and CARE AI models, and fostering collaborations especially with researchers from underserved communities, building bridges between medical voice research, acoustic engineers, and the AI/ML community.
voice:task:7
Enable AI model pretraining, fine-tuning, benchmarking, and validation using the standardized BIDS-compliant dataset structure with features extracted via b2aiprep, SenseLab, OpenSMILE, Parselmouth, Praat, and torchaudio toolkits.
DescriptionID
Primary intended use: development and validation of AI/ML models for voice as a biomarker of health, supporting screening, diagnosis, and treatment of voice disorders, neurological disorders, mood disorders, respiratory disorders, and pediatric speech disorders through model pretraining, fine-tuning, benchmarking, and validation.
voice:use:1
Research into acoustic biomarkers and development of standards for voice data collection and analysis. Establishing best practices for AI/ML-friendly voice datasets and contributing to the field's maturation as a clinical diagnostic modality.
voice:use:2
Training and education in voice AI research through workforce development initiatives, curriculum creation on FAIR and CARE voice AI model development, and fostering collaborations between medical voice researchers, acoustic engineers, and AI/ML specialists, especially from underserved communities.
voice:use:3
Multimodal health research combining voice data with EHR information, radiomics, genomics, imaging, and other health biomarkers to understand complex disease relationships and improve diagnostic accuracy.
voice:use:4
  • ID
    voice:existinguse:1
    Description
    A restricted version of the dataset containing raw audio has been used in the Bridge2AI Summer School and hackathon for education and research training purposes.
DescriptionID
Non-clinical applications such as hiring decisions, insurance premium adjustments, or any form of surveillance that could lead to discrimination or harm based on health conditions or voice characteristics. These applications could negatively impact individuals.
voice:discouraged:1
Any attempt to re-identify research participants or use data in ways that could foreseeably cause harm or stigmatization to research participants, their families, communities, or specific populations. Dataset covered under Certificate of Confidentiality.
voice:discouraged:2
Development or use of intellectual property protections, database rights, or related rights in ways that would prevent or block access to any element of the dataset or conclusions derived from it. Must respect Fort Lauderdale Agreement and Open Science principles.
voice:discouraged:3
RoleNameORCIDAffiliation
Contributorvoice:prohibited:1-
ID
voice:license:1
Name
Bridge2AI Voice Registered Access License
Description
Public access dataset distributed through PhysioNet under Bridge2AI Voice Registered Access License. Only registered users who sign the specified Data Use Agreement (Bridge2AI Voice Registered Access Agreement) can access files. Data covered under Certificate of Confidentiality which must be asserted against compulsory legal demands. Raw audio available through controlled access only via Data Access Compliance Office (DACO) requiring distinct application and DTUA signed by institutional official. Recipient must adhere to PhysioNet requirements managed by MIT Laboratory for Computational Physiology. Recipient encouraged to publish results in open-access journals. No export controls apply. Commercial and non-commercial research use permitted for Authorized Researchers.
📤

Distribution

How will the dataset be distributed?

10.13026/37yb-1t42
Bridge2AI Voice Registered Access License
ID
voice:versionaccess:1
Name
Version access policy
Description
All dataset versions available through PhysioNet at https://physionet.org/content/b2ai-voice/ with version-specific DOIs. DOI for latest version: https://doi.org/10.13026/37yb-1t42. Earlier versions also available on Health Data Nexus at https://healthdatanexus.ai/content/b2ai-voice/1.0/. By default, older versions continue to be supported, hosted, and made available. Dataset publishers reserve right to remove access to older versions. Each version has unique DOI.
🔄

Maintenance

How will the dataset be maintained?

3.0.0
ID
voice:updates:1
Name
Versioned releases with ongoing data collection
Description
Dataset updated with versioned static releases semi-annually as data collection progresses. Users notified through news items on platforms and standard communication channels. v1.0 released January 17, 2025 (306 participants, 12,523 recordings); v1.1 added MFCC features; v2.0.0 April 16, 2025; v2.0.1 August 18, 2025; v3.0.0 released 2025 (833 adults, ~61,937 recordings). Pediatric v1.0 released separately. Data collection ongoing through November 30, 2026. Target: 10,000 participants by 2027. Future releases will expand Spanish language protocols and add additional multimodal data (imaging, genomics). Version-specific DOIs maintained. Older versions continue to be supported and hosted.
👥

Human Subjects

Does the dataset relate to people?

ID
voice:hsr:1
Name
Bridge2AI-Voice Human Subjects Research
Description
Observational study (cross-sectional design) involving direct collection of voice recordings, questionnaire responses, and EHR linkage from human participants. Prospective informed consent obtained from all participants. IRB approved by University of South Florida Single IRB with subsite IRBs through Single IRB process. Not a drug or medical device study. No data monitoring committee appointed. Study ID: OT2OD032720.
Involves Human Subjects
True
IRB Approval
  • University of South Florida Single IRB approval with subsite IRB approvals through Single IRB process
Ethics Review Board
  • University of South Florida Institutional Review Board (Single IRB)
Special Populations
  • Pediatric participants (aged 2-18 years) at Hospital for Sick Children with age-appropriate protocols
Regulatory Compliance
  • HIPAA Safe Harbor de-identification standards applied
  • Certificate of Confidentiality covering dataset against compulsory legal demands
  • 45 CFR 46 (Common Rule) compliance
  • Data covered as Personally Identifiable Information per OMB Memorandum M-07-16
  • ID
    voice:revocation:1
    Description
    Participants informed they may withdraw from study at any point. If withdrawal occurs during or before voice data collection, data not included in database. Participants informed that research data (including voice recordings) cannot be removed from the database once the voice data collection process is completed. Mechanism: direct communication with research team during collection session.
Generated on 2026-04-15 17:53:11 using Bridge2AI Data Sheets Schema