id: https://doi.org/10.13026/37yb-1t42
name: Bridge2AI-Voice
title: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information
description: 'The Bridge2AI-Voice project seeks to create an ethically sourced flagship dataset to enable future research in artificial intelligence and support critical insights into the use of voice as a biomarker of health. The human voice contains complex acoustic markers which have been linked to important health conditions including dementia, mood disorders, and cancer. When viewed as a biomarker, voice is a promising characteristic to measure as it is simple to collect, cost-effective, and has broad clinical utility. This comprehensive collection provides voice recordings with corresponding clinical information from participants selected based on known conditions which manifest within the voice waveform including voice disorders, neurological disorders, mood disorders, and respiratory disorders. The dataset is designed to fuel voice AI research, establish data standards, and promote ethical and trustworthy AI/ML development for voice biomarkers of health. Data collection occurs through a multi-institutional collaborative effort using standardized protocols, custom smartphone applications, and rigorous ethical oversight. Version 3.0 provides approximately 61,937 voice-derived recordings from 833 adult participants collected across multiple sites in North America, with derived features such as spectrograms, MFCCs, acoustic features, and clinical phenotype data. The pediatric dataset v1.0 is also available with data from 300 participants. Raw audio data is available through controlled access to protect participant privacy. '
doi: 10.13026/37yb-1t42
page: https://docs.b2ai-voice.org
language: en
version: 3.0.0
license: Bridge2AI Voice Registered Access License
is_tabular: false
keywords: - voice biomarker - acoustic biomarker - Bridge2AI - voice AI - voice disorders - neurological disorders - neurodegenerative disorders - mood disorders - psychiatric disorders - respiratory disorders - pediatric voice disorders - speech disorders - Parkinson's disease - Alzheimer's disease - depression - schizophrenia - bipolar disorder - stroke - ALS - autism spectrum disorder - speech delay - laryngeal cancer - vocal fold paralysis - muscle tension dysphonia - laryngeal dystonia - COPD - chronic cough - airway stenosis - obstructive sleep apnea - spectrogram - MFCC - mel-frequency cepstral coefficients - OpenSMILE - Praat - Parselmouth - torchaudio - federated learning - ethical AI - multimodal health data - electronic health records - EHR - radiomics - genomics - FAIR principles - CARE principles - PhysioNet - Health Data Nexus - BIDS - Brain Imaging Data Structure - b2aiprep
publisher: Bridge2AI-Voice Consortium, University of South Florida
download_url: https://physionet.org/content/b2ai-voice/
purposes:
- id: voice:purpose:1
description: 'Integrate the use of voice as a biomarker of health in clinical care by generating a substantial
multi-institutional, ethically sourced, and diverse voice database linked to multimodal health biomarkers
to fuel voice AI research and build predictive models to assist in screening, diagnosis, and treatment
of a broad range of diseases.
'
- id: voice:purpose:2
description: 'Create an ethically sourced flagship dataset of 10,000 voices linked to health information
to enable future research in artificial intelligence and support critical insights into the use of
voice as a biomarker of health, addressing the pressing need for large, high quality, multi-institutional
and diverse voice databases linked to other health biomarkers.
'
- id: voice:purpose:3
description: 'Establish standards, best practices, and guidelines for voice data collection and analysis
to advance the field of acoustic biomarkers by developing new standards that are AI/ML friendly and
enable voice to emerge as a biomarker of health.
'
- id: voice:purpose:4
description: 'Address ethical, legal, and social challenges surrounding voice AI including risks of
voice re-identification, vulnerabilities like voice AI hacking, concerns around voice data sharing
and privacy, and the influence of gender and racial diversity on voice AI.
'
tasks:
- id: voice:task:1
description: 'Enable development of AI/ML predictive models for screening, diagnosis, and treatment
of voice disorders including laryngeal cancers, vocal fold paralysis, muscle tension dysphonia, laryngeal
dystonia, benign laryngeal lesions, and pre-cancerous lesions.
'
- id: voice:task:2
description: 'Support machine learning models for neurological and neurodegenerative disorders including
Alzheimer''s disease, Parkinson''s disease, mild cognitive impairment, other dementias, and ALS, detecting
voice and speech changes.
'
- id: voice:task:3
description: 'Develop AI algorithms for mood and psychiatric disorder detection including depression,
schizophrenia, bipolar disorder, anxiety disorders, ADHD, PTSD, OCD, and borderline personality disorder.
'
- id: voice:task:4
description: 'Create machine learning models for respiratory disorder screening and therapeutic monitoring
using respiratory sounds, cough sounds, and voice, applicable to conditions such as chronic cough,
COPD, airway stenosis, and other respiratory conditions.
'
- id: voice:task:5
description: 'Build AI models for pediatric voice and speech disorder detection using 36 pediatric-specific
acoustic tasks and specialized questionnaires, addressing the relative scarcity of pediatric voice
data for participants aged 2-18.
'
- id: voice:task:6
description: 'Promote application of AI/ML for voice research through workforce development, curriculum
creation on voice biomarkers of health for FAIR and CARE AI models, and fostering collaborations especially
with researchers from underserved communities.
'
- id: voice:task:7
description: 'Enable AI model pretraining, fine-tuning, benchmarking, and validation using the standardized
BIDS-compliant dataset structure with features extracted via b2aiprep, SenseLab, OpenSMILE, Parselmouth,
Praat, and torchaudio toolkits.
'
addressing_gaps:
- id: voice:gap:1
description: 'Address the lack of large, high quality, multi-institutional and diverse voice databases
linked to multimodal health biomarkers (demographics, imaging, genomics, risk factors) necessary to
fuel voice AI research and answer tangible clinical questions.
'
- id: voice:gap:2
description: 'Overcome limitations in existing voice and psychiatric disorder research that has relied
on small datasets with limited demographic diversity reporting, lack of standardized data collection
protocols, and possible confounders limiting external validity.
'
- id: voice:gap:3
description: 'Fill the gap in pediatric voice and speech analysis research, which is sparser partly
due to ethical concerns and challenges in data acquisition for this cohort, particularly for autism
spectrum disorder and speech delay detection.
'
- id: voice:gap:4
description: 'Establish missing standards for voice data collection, acoustic analysis, and ethical
frameworks for consenting to voice data collection, sharing, and utilization in the context of voice
AI technology development and clinical adoption.
'
- id: voice:gap:5
description: 'Develop software and cloud infrastructure for automated voice data collection through
a smartphone application (Bridge2AI-Voice App) that allows non-invasive, user-friendly, high quality
voice data collection while minimizing human manipulation and implementing federated learning technology
to minimize data sharing while preserving patient privacy.
'
creators:
- id: voice:creator:1
description: 'Bridge2AI-Voice Consortium led by Dr. Yael Bensoussan (Contact PI, University of South
Florida, Department of Otolaryngology) and Dr. Olivier Elemento (Co-PI, Weill Cornell Medicine). The
multidisciplinary consortium includes over 50 investigators from 12+ institutions across North America
spanning clinical medicine, biomedical research, machine learning, data science, social science, and
ethics. Key co-investigators include: Alexandros Sigaras (Weill Cornell Medicine), Anais Rameau (Weill
Cornell Medicine), Maria Powell (Vanderbilt University Medical Center), Ruth Bahr (USF), Jennifer
Siu (Hospital for Sick Children), Philip Payne (Washington University in St. Louis), David Dorr (Oregon
Health and Science University), Jean-Christophe Belisle-Pipon (Simon Fraser University), Vardit Ravitsky
(The Hastings Center), Satrajit Ghosh (MIT), Frank Rudzicz (University of Toronto), Jordan Lerner-Ellis
(Sinai Health), Don Bolser (University of Florida), Alistair Johnson (MIT/PhysioNet), and Jennifer
Siu (Hospital for Sick Children).
'
funders:
- id: voice:funder:1
description: 'National Institutes of Health (NIH) Common Fund Bridge2AI Program. Grant number: 3OT2OD032720-01S3.
Opportunity Number: OTA-21-008. Project dates: September 1, 2022 to November 30, 2026. Total funding
in 2025: $4,660,942 (Direct Costs: $4,072,321, Indirect Costs: $588,621). Administering Institute:
NIH Office of the Director. Study Section: Data Coordination, Mapping, and Modeling (DCMM).
'
- id: voice:funder:2
description: 'National Institute of Biomedical Imaging and Bioengineering (NIBIB). Supports PhysioNet
managed by MIT Laboratory for Computational Physiology under NIH grant number R01EB030362, which serves
as the primary distribution platform for the Bridge2AI-Voice dataset.
'
instances:
- id: voice:instance:1
description: 'Adult participants presenting at specialty clinics across multiple sites in North America.
Participants selected based on membership to five predetermined disease cohort groups: Voice Disorders,
Neurological and Neurodegenerative Disorders, Mood and Psychiatric Disorders, Respiratory Disorders,
and Pediatric Voice and Speech Disorders. Version 3.0 contains approximately 833 adult participants
with ~61,937 voice-derived recordings. Pediatric dataset v1.0 adds 300 participants. Enrollment anticipated
to reach 10,000 participants by 2027.
'
instance_type: Human participants with clinical diagnoses recruited from specialty clinics
counts: 833
label: true
label_description: 'Diagnostic labels assigned by clinical assessment at each site. Clinicians provided
diagnoses based on clinical interview and appropriate work-up including laryngoscopy, stroboscopy,
MRI, CT, whole genome sequencing, EHR records, and medication prescriptions. Labels include diagnostic
categories: vocal pathologies, neurological disorders, psychiatric conditions, respiratory disorders,
and pediatric voice/speech disorders. Per Bridge2AI protocols and ICD-10 codes. Single labeler per
participant (site clinician).
'
known_biases:
- id: voice:bias:1
description: 'Clinic-based recruitment introduces selection bias toward treatment-seeking populations.
Participants are recruited from high-volume specialty clinics, which may not represent the full spectrum
of disease severity or the general population with these conditions. Underrepresentation of individuals
with less trust in the medical system or less proximity to collection sites.
'
- id: voice:bias:2
description: 'Current releases contain only English-speaking participants, which may bias acoustic features
and models trained on this data toward English-language phonological patterns. Spanish protocols are
under development for future releases.
'
- id: voice:bias:3
description: 'Limited geographic diversity: data collected at five North American sites only, which
may not capture regional variation in voice characteristics, dialects, or environmental acoustic conditions.
'
- id: voice:bias:4
description: 'Imbalanced distribution across disease categories in public releases; the dataset does
not contain equal representation across all five disease cohort groups. This may affect the performance
and fairness of models trained on the dataset.
'
known_limitations:
- id: voice:limitation:1
description: 'Version 3.0 (833 adult participants) represents an interim dataset from an ongoing collection
targeting 10,000 participants by 2027. Statistical power for some analyses, particularly for less
prevalent conditions, may be limited.
'
- id: voice:limitation:2
description: 'Raw audio waveforms are excluded from public release and available only through controlled
access. This restricts certain types of analyses requiring the original audio signal and adds access
overhead for researchers.
'
- id: voice:limitation:3
description: 'Remote data collection not included in initial releases; all data collected in clinical
settings by research assistants, which may not generalize to naturalistic or home settings.
'
- id: voice:limitation:4
description: 'Multimodal data (imaging, genomics, full EHR data) not included in current public releases;
only derived clinical features and diagnoses are provided.
'
- id: voice:limitation:5
description: 'No analysis of potential re-identification impact has been formally conducted on the de-identified
dataset beyond HIPAA Safe Harbor standards.
'
confidential_elements:
- id: voice:confidential:1
description: 'Raw audio waveforms are considered biometric identifiers under HIPAA and are restricted
to controlled access only. Researchers must apply through the Data Access Compliance Office (DACO)
and sign an institutional Data Use and Transfer Agreement (DTUA).
'
- id: voice:confidential:2
description: 'Open-response audio features (spectrograms, MFCCs, mel spectrograms, transcriptions, EMAs,
and PPGs from open-response prompts) are removed from the public registered access dataset to protect
participant privacy.
'
- id: voice:confidential:3
description: 'The dataset is covered under a Certificate of Confidentiality, which provides legal protection
against compulsory demands such as court orders and subpoenas for identifying information about research
participants.
'
subpopulations:
- id: voice:subpop:1
name: Voice Disorders cohort
description: 'Participants with laryngeal disorders including laryngeal cancer (T1-T4, biopsy proven),
laryngitis (acute, chronic, bacterial, fungal, autoimmune), pre-cancerous lesions (keratosis, leukoplakia),
benign vocal cord lesions (nodules, polyps, cysts, Reinke''s edema, recurrent laryngeal papilloma),
muscle tension dysphonia, spasmodic dysphonia and laryngeal tremor, unilateral vocal fold paralysis,
and glottic insufficiency/ presbyphonia. Validated by laryngoscopy images and stroboscopy videos.
'
- id: voice:subpop:2
name: Neurological and Neurodegenerative Disorders cohort
description: 'Participants aged 44-85 with clinical diagnoses including mild cognitive impairment, Alzheimer''s
disease, other dementias (frontotemporal, Lewy body, vascular, mixed, alcohol-induced), ALS (sporadic,
familial, spinal/limb-onset, bulbar-onset), and Parkinson''s disease (idiopathic PD, multiple system
atrophy, progressive supranuclear palsy, corticobasal degeneration). Validated by CT Brain, MRI Brain,
serum biomarkers, and whole genome sequencing.
'
- id: voice:subpop:3
name: Mood and Psychiatric Disorders cohort
description: 'Participants with clinical diagnoses including depression/major depressive disorder, bipolar
I/II disorder, anxiety disorder, schizophrenia, ADHD, PTSD, OCD, borderline personality disorder,
and other psychiatric disorders. Validated by electronic health summary, clinical diagnosis from psychiatrist,
and medication history.
'
- id: voice:subpop:4
name: Respiratory Disorders cohort
description: 'Participants with respiratory conditions including airway stenosis (bilateral vocal fold
paralysis, supraglottic, glottic, posterior glottic, subglottic, tracheal stenosis, multi-level upper
airway stenosis) and chronic cough (bothersome cough >8 weeks). Validated by spirometry, flow volume
loops, and CT scan of neck/chest.
'
- id: voice:subpop:5
name: Pediatric cohort
description: 'Pediatric participants aged 2-18 with voice and speech disorders recruited exclusively
from Hospital for Sick Children (SickKids). Grouped by age: 2-4, 4-6, 6-10, 10+ years. Data collected
using reproschema-ui with Bridge2AI-Voice pediatric protocol. Includes 36 pediatric-specific acoustic
tasks and specialized questionnaires (C-VHI-10, PVOS, PVRQOL, PHQ-A). Pediatric dataset v1.0 available
separately from adult dataset.
'
- id: voice:subpop:6
name: Control participants
description: 'Healthy volunteer control participants who complete common questionnaires and voice tasks
(Voice Handicap Index-10, PHQ-9, GAD-7, Winograd) to provide normative comparisons.
'
sensitive_elements:
- id: voice:sensitive:1
description: 'Voice recordings are biometric identifiers under HIPAA. Raw audio waveforms are excluded
from public release and available only through controlled access with DACO approval and institutional
DTUA. Risks include voice re-identification, voice AI hacking, and illicit or unauthorized use of
voice data.
'
- id: voice:sensitive:2
description: 'Electronic health record (EHR) data accessed with participant consent for gold standard
validation of diagnoses and symptoms. Sensitive health information including diagnoses, symptoms,
disease-specific clinical data, and multimodal health biomarkers.
'
- id: voice:sensitive:3
description: 'Dataset contains sensitive demographic information including racial and ethnic origins,
sexual orientation, financial and socioeconomic status, and health data. All direct identifiers removed;
indirect identifiers removed where creating significant re-identification risk.
'
- id: voice:sensitive:4
description: 'Free speech task transcriptions may contain potentially identifying information or external
voices. All open-response audio features (spectrograms, MFCCs, mel spectrograms, transcriptions, EMAs,
PPGs) removed from public feature-only dataset.
'
- id: voice:sensitive:5
description: 'Dataset is covered under Certificate of Confidentiality protecting against compulsory
legal demands such as court orders and subpoenas for identifying information or characteristics of
research participants.
'
collection_mechanisms:
- id: voice:collection:1
description: 'Voice data collected in clinic using custom Bridge2AI-Voice App on iPad (9th or 10th generation)
or iPad Air (5th generation) with Avid AE-36 microphone and Apple dongle connector. App collects breathing
sounds and voice, speech, and linguistic tasks along with health information through surveys and validated
questionnaires. Research assistant present during collection. Future remote data collection planned
but not included in current releases.
'
- id: voice:collection:2
description: 'Pediatric data collected using reproschema-ui with Bridge2AI-Voice pediatric protocol.
Same iPad hardware as adult protocol with age-appropriate tasks grouped by age range (2-4, 4-6, 6-10,
10+ years). Questions read to participants when needed.
'
- id: voice:collection:3
description: 'Clinical data including EHR information, imaging (laryngoscopy, stroboscopy, MRI, CT),
and genomic data extracted from sites independently and uploaded through REDCap database. No external
multimodal data released in current dataset versions.
'
acquisition_methods:
- id: voice:acquisition:1
description: '22 acoustic tasks recorded through Bridge2AI-Voice App for adult participants including:
Non-Voice (Respiration, Cough, Breath Sounds, Voluntary Cough); Voice/Non-Speech (Prolonged Vowel
/e/, Maximum Phonation Time, Glides, Loudness /Hey/, Diadochokinesis /pa/ta/ka/); Speech (Rainbow
Passage, Caterpillar Passage, Cape-V Sentences, Free Speech, Picture Description, Story Recall, Animal
Fluency, Open Response Questions, Word-Color Stroop, Productive Vocabulary, Random Item Generation,
Cinderella Story). Tasks vary by disease cohort (Part A/Voice/Resp/Mood/Neuro).
'
- id: voice:acquisition:2
description: '36 pediatric acoustic tasks collected via reproschema-ui including: Speech tasks (ABC''s,
Ready for School, Favorite Show, Favorite Food, Outside of School, Months, Counting, Naming Animals,
Naming Food, Identifying Pictures, Picture Description, Caterpillar Passage, Repeat Words, Role Naming,
Repeat Sentences); Voice/Non-Speech tasks (Long Sounds, Noisy Sounds, Silly Sounds /PUH TUH KUH/).
'
- id: voice:acquisition:3
description: 'Self-reported demographic data and medical history questionnaires administered through
app. Validated questionnaires integrated for each disease cohort including: VHI-10, PHQ-9, GAD-7,
PANAS, Custom Affect Scale, PTSD Adult, ADHD Adult, DSM-5 Adult, Dyspnea Index, Leicester Cough Questionnaire,
Winograd Questionnaire, MOCA, C-VHI-10 (peds), PVOS (peds), PVRQOL (peds), PHQ-A (peds).
'
- id: voice:acquisition:4
description: 'Electronic health record (EHR) access for consenting participants permitting investigators
to access medical information through EHR platforms to perform gold standard validation of diagnoses
and symptoms. Linkage to multimodal health biomarkers including laryngoscopy imaging, radiomics, genomics,
respiratory function tests.
'
collection_timeframes:
- id: voice:timeframe:1
description: 'Data collection ongoing from September 2022 through November 2026 (project end date).
Initial dataset releases began in late 2024. Version 1.0 released January 17, 2025 (306 participants).
Version 3.0 released 2025 (833 adult participants). Semi-annual releases planned. Enrollment anticipated
to reach 10,000 participants by 2027.
'
data_collectors:
- id: voice:datacollector:1
description: 'Research teams at each of the five North American collection sites, including medical
graduate and undergraduate students coordinating with site clinicians and doctors. Clinicians and
doctors listed under IRB as co-investigators and added to consortium. Participants compensated via
electronic gift cards: $40 for sessions under 90 minutes, $80 for sessions over 90 minutes, maximum
3 sessions and $120 total compensation.
'
sampling_strategies:
- id: voice:sampling:1
description: 'Non-probability purposive sampling from specialty clinics (high volume expert clinics,
outpatient clinics seeing >50 patients per month from same disease category). Patients presenting
at clinics screened for eligibility per inclusion/exclusion criteria outlined in protocol Table 1.
Not representative of general population due to limited geographic locations and clinic-based recruitment.
Current v3.0 dataset is a sample of an ongoing collection targeting 10,000 participants by 2027. English-speaking
adult participants (18-120 years); Spanish protocols under development for future releases.
'
is_sample: true
is_random: false
is_representative: false
missing_data_documentation:
- id: voice:missingdata:1
description: 'Missingness tables generated and included with dataset as part of the audit protocol.
Audio quality control metrics applied including silence amount, duration, and speech-to-passage accuracy
checks. Processing using b2aiprep and SenseLab toolkits. Some participants may have incomplete questionnaire
responses or missing sessions.
'
raw_data_sources:
- id: voice:rawsource:1
description: 'Raw audio waveforms in WAV format collected via Bridge2AI-Voice App on iPad devices with
Avid AE-36 microphone. Stored in BIDS v1.9.0 compliant format per participant session and acoustic
task. Available through controlled access only.
'
source_description: 'Raw audio WAV files recorded directly by participants in clinical settings using
the Bridge2AI-Voice App on iPad with Avid AE-36 microphone. Organized in BIDS v1.9.0 directory structure:
sub-{participant_id}/ses-{session_id}/audio/ with companion JSON metadata sidecar files. Available
through DACO-controlled access only.
'
- id: voice:rawsource:2
description: 'REDCap database exports containing EHR-linked clinical data, demographic information,
disease-specific validated questionnaire responses, and diagnostic information. Pediatric data first
extracted from reproschema-ui to REDCap format.
'
source_description: 'REDCap (Research Electronic Data Capture) database exports and reproschema-ui exports
containing clinical phenotype data: demographic information, validated questionnaire responses (VHI-10,
PHQ-9, GAD-7, etc.), diagnostic information, and EHR-linked clinical assessments. Pediatric data first
extracted from reproschema-ui then converted to REDCap format before BIDS conversion.
'
raw_sources:
- id: voice:rawsrc:1
description: 'Raw audio WAV files recorded through the Bridge2AI-Voice App, organized in BIDS v1.9.0
compliant directory structure. Available through controlled access via DACO approval.
'
- id: voice:rawsrc:2
description: 'REDCap database exports and reproschema-ui exports containing clinical phenotype data,
questionnaire responses, and participant demographics.
'
preprocessing_strategies:
- id: voice:preproc:1
description: 'Raw audio preprocessing using b2aiprep library: conversion to monaural audio, resampling
to 16 kHz with Butterworth anti-aliasing filter. Standardization ensures consistent format across
all recordings for downstream feature extraction.
'
- id: voice:preproc:2
description: 'Spectrogram extraction using short-time Fast Fourier Transform (FFT): 25ms window size,
10ms hop length, 512-point FFT. Output spectrograms have 513xN dimensions where N is proportional
to audio length. Stored in Parquet format.
'
- id: voice:preproc:3
description: 'Mel-frequency cepstral coefficients (MFCC) extraction: 60 MFCCs extracted from spectrograms
capturing perceptually-relevant spectral envelope characteristics. Output dimension 60xN. Stored in
Parquet format.
'
- id: voice:preproc:4
description: 'OpenSMILE eGeMaps acoustic feature extraction capturing temporal dynamics and acoustic
characteristics. One row per unique recording in static_features.tsv.
'
- id: voice:preproc:5
description: 'Parselmouth and Praat phonetic and prosodic feature computation providing fundamental
frequency (F0), formants, and voice quality measures. Documented in static_features.json data dictionary.
'
- id: voice:preproc:6
description: 'Torchaudio-based feature extraction including pitch contour, spectrograms, mel spectrograms,
MFCCs. SPARC-based features including electromagnetic articulography (EMA) estimates, loudness, periodicity,
and pitch measures. Phonetic posteriorgrams (PPGs).
'
- id: voice:preproc:7
description: 'Transcription generation using OpenAI Whisper model applied to audio recordings. Raw audio
transcripts reviewed and any recordings containing potentially identifying information or external
voices removed. All transcriptions, EMAs, and PPGs from open-response prompts removed from public
feature-only dataset.
'
- id: voice:preproc:8
description: 'REDCap and reproschema-ui data exported and converted to BIDS v1.9.0 format using b2aiprep
open-source library. Phenotype data organized in tab-delimited files with JSON data dictionaries.
Pediatric data first extracted from reproschema-ui to REDCap format then converted to BIDS.
'
cleaning_strategies:
- id: voice:cleaning:1
description: 'HIPAA Safe Harbor de-identification for public release: removed direct identifiers (names,
civic addresses, social security numbers), indirect identifiers creating significant re-identification
risk (select geographic/demographic identifiers, household composition, cultural identity), and sensitive
information (household income, mental health status, traumatic life experiences). Geographic data
limited; state/province removed, country retained.
'
- id: voice:cleaning:2
description: 'Audio privacy protection for public release: all raw audio waveforms excluded from public
dataset. All spectrograms, MFCCs, mel spectrograms, transcriptions, EMAs, and PPGs from open-response
prompts removed from feature-only dataset. Free speech transcripts removed. Raw audio available only
through controlled access with DACO approval.
'
- id: voice:cleaning:3
description: 'Sensitive field removal based on REDCap data dictionary: all fields encoded as sensitive
(column "Identifier?" in REDCap data dictionary CSV) removed from dataset.
'
- id: voice:cleaning:4
description: 'Audit protocol applied: missingness tables generated and included with dataset; distribution
and outlier checks; categorical responses checked against schema; audio quality control metrics including
silence amount, duration, and speech-to-passage accuracy checks. Processing using b2aiprep and SenseLab
toolkits.
'
labeling_strategies:
- id: voice:labeling:1
description: 'Diagnostic labels assigned by site clinicians based on clinical assessment and gold standard
validation methods per Bridge2AI-Voice protocol Table 1 and ICD-10 codes. For voice disorders: laryngoscopy
and stroboscopy. For neurological disorders: CT Brain, MRI Brain, genome sequencing, serum biomarkers.
For mood/psychiatric disorders: EHR records, psychiatrist diagnosis, medication history. For respiratory
disorders: spirometry, flow volume loops, CT scan. Single labeler (site clinician) per participant.
'
machine_annotation_tools:
- id: voice:machanno:1
description: 'OpenAI Whisper model used for automated transcription of audio recordings. Transcripts
reviewed for quality; recordings with potentially identifying information or external voices removed.
Transcriptions from open-response prompts removed from public release.
'
- id: voice:machanno:2
description: 'b2aiprep open-source library used for automated preprocessing of raw audio, REDCap data
conversion, and BIDS formatting. SenseLab toolkit used for audio quality control and feature extraction.
'
intended_uses:
- id: voice:use:1
description: 'Primary intended use: development and validation of AI/ML models for voice as a biomarker
of health, supporting screening, diagnosis, and treatment of voice disorders, neurological disorders,
mood disorders, respiratory disorders, and pediatric speech disorders through model pretraining, fine-tuning,
benchmarking, and validation.
'
- id: voice:use:2
description: 'Research into acoustic biomarkers and development of standards for voice data collection
and analysis. Establishing best practices for AI/ML-friendly voice datasets and contributing to the
field''s maturation as a clinical diagnostic modality.
'
- id: voice:use:3
description: 'Training and education in voice AI research through workforce development initiatives,
curriculum creation on FAIR and CARE voice AI model development, and fostering collaborations between
medical voice researchers, acoustic engineers, and AI/ML specialists, especially from underserved
communities.
'
- id: voice:use:4
description: 'Multimodal health research combining voice data with EHR information, radiomics, genomics,
imaging, and other health biomarkers to understand complex disease relationships and improve diagnostic
accuracy.
'
existing_uses:
- id: voice:existinguse:1
description: 'A restricted version of the dataset containing raw audio has been used in the Bridge2AI
Summer School and hackathon for education and research training purposes.
'
discouraged_uses:
- id: voice:discouraged:1
description: 'Non-clinical applications such as hiring decisions, insurance premium adjustments, or
any form of surveillance that could lead to discrimination or harm based on health conditions or voice
characteristics. These applications could negatively impact individuals.
'
- id: voice:discouraged:2
description: 'Any attempt to re-identify research participants or use data in ways that could foreseeably
cause harm or stigmatization to research participants, their families, communities, or specific populations.
Dataset covered under Certificate of Confidentiality.
'
- id: voice:discouraged:3
description: 'Development or use of intellectual property protections, database rights, or related rights
in ways that would prevent or block access to any element of the dataset or conclusions derived from
it. Must respect Fort Lauderdale Agreement and Open Science principles.
'
prohibited_uses:
- id: voice:prohibited:1
description: 'Use of the dataset outside of authorized research purposes as defined in the Bridge2AI
Voice Registered Access Agreement. The dataset is intended solely for commercial and non-commercial
research by Authorized Researchers. Sale of all or part of the data on any media is prohibited. Re-identification
attempts are prohibited.
'
future_use_impacts:
- id: voice:futureimpact:1
description: 'Potential risks from future uses include voice re-identification despite de-identification
efforts, discriminatory application of voice biomarker models in clinical or non-clinical settings,
and reinforcement of health disparities if models are trained on non-representative data. The Certificate
of Confidentiality and access controls mitigate some of these risks.
'
- id: voice:futureimpact:2
description: 'Positive anticipated impacts include acceleration of clinical voice AI adoption for screening
and diagnosis of conditions currently lacking cost-effective diagnostic tools, improved health equity
through diverse data collection, and establishment of FAIR and CARE-compliant data standards for the
voice AI field.
'
distribution_formats:
- id: voice:format:1
name: Parquet format for spectrograms and derived features
description: 'Time-varying features (spectrograms, mel spectrograms, MFCCs, pitch contour, SPARC features,
PPGs) stored in Parquet format compatible with Python datasets library. Each element contains participant_id,
session_id, task_name, and feature arrays. Excluded for open-response tasks in public release.
'
- id: voice:format:2
name: TSV/JSON for phenotype and static features
description: 'Phenotype data in tab-delimited format (confounders.tsv, demographics.tsv, diagnosis/
*.tsv, enrollment/*.tsv, questionnaire/*.tsv, task/*.tsv) with JSON data dictionaries per file. Static
acoustic features (OpenSMILE, Praat, Parselmouth, torchaudio) in static_features.tsv with static_features.json
data dictionary. One row per participant or recording as appropriate. BIDS v1.9.0 compliant structure.
'
- id: voice:format:3
name: WAV audio format (controlled access only)
description: 'Original raw audio waveforms in WAV format following BIDS structure: sub-{participant_id}/ses-{session_id}/audio/sub{id}_ses{id}_task-{task_name}.wav
with companion JSON metadata. Available through controlled access only. Contact DACO@b2ai-voice.org
to request access.
'
distribution_dates:
- id: voice:distdate:1
description: 'Dataset first published and made available late November 2024 through Health Data Nexus.
PhysioNet releases: v1.0 January 17, 2025 (306 participants, 12,523 recordings); v1.1 January 17,
2025 (added MFCC features); v2.0.0 April 16, 2025; v2.0.1 August 18, 2025; v3.0.0 released 2025 (833
adult participants, ~61,937 recordings). Semi-annual releases planned. Pediatric v1.0 available separately.
'
distributions:
- id: voice:dist:1
description: 'Featurized dataset (registered access) distributed through PhysioNet. Contains AI-ready
derived features: OpenSMILE eGeMaps, Parselmouth/Praat speech features, torchaudio features, spectrograms,
mel spectrograms, MFCCs, SPARC features, PPGs. Phenotype data in BIDS v1.9.0 compliant TSV/JSON format.
HIPAA Safe Harbor de-identified. Parquet files used for time-varying features (spectrograms, MFCCs,
pitch contour).
'
format: TSV
media_type: text/tab-separated-values
path: https://physionet.org/content/b2ai-voice/
- id: voice:dist:2
description: 'JSON data dictionaries and metadata distributed alongside TSV phenotype files. Each TSV
file has a companion JSON data dictionary describing field names, types, and allowed values. Follows
BIDS v1.9.0 specification.
'
format: JSON
media_type: application/json
path: https://physionet.org/content/b2ai-voice/
- id: voice:dist:3
description: 'Raw audio dataset (controlled access only) distributed through PhysioNet and Health Data
Nexus. Contains original raw audio waveforms compressed in GZ archives following BIDS v1.9.0 structure.
Requires DACO approval and institutional DTUA. Contact DACO@b2ai-voice.org to request access.
'
format: GZ
media_type: application/gzip
path: mailto:DACO@b2ai-voice.org
- id: voice:dist:4
description: 'Pediatric dataset v1.0 (registered access) distributed through PhysioNet as separate download.
Contains data from 300 pediatric participants with 36 pediatric-specific acoustic tasks and specialized
questionnaires. BIDS v1.9.0 compliant TSV/JSON format.
'
format: TSV
media_type: text/tab-separated-values
path: https://physionet.org/content/b2ai-voice/
maintainers:
- id: voice:maintainer:1
description: 'Bridge2AI-Voice Consortium (University of South Florida, lead institution) supported by
NIH Bridge2AI program. Contact: [email protected]. Responsible for dataset curation, standards development,
ethics oversight, versioning, and updates. Data Access Compliance Office (DACO) manages controlled
access applications. Contact: DACO@b2ai-voice.org.
'
- id: voice:maintainer:2
description: 'PhysioNet (MIT Laboratory for Computational Physiology, supported by NIBIB NIH grant R01EB030362)
serves as primary distribution platform for registered access dataset. Provides technical infrastructure
and access management.
'
- id: voice:maintainer:3
description: 'Health Data Nexus (Temerty Centre for Artificial Intelligence Research and Education in
Medicine, T-CAIREM, University of Toronto) maintains alternative platform providing cloud compute
alongside dataset. Earlier dataset versions available here. Contact: [email protected].
'
updates:
id: voice:updates:1
name: Versioned releases with ongoing data collection
description: 'Dataset updated with versioned static releases semi-annually as data collection progresses.
Users notified through news items on platforms and standard communication channels. v1.0 released
January 17, 2025 (306 participants, 12,523 recordings); v1.1 added MFCC features; v2.0.0 April 16,
2025; v2.0.1 August 18, 2025; v3.0.0 released 2025 (833 adults, ~61,937 recordings). Pediatric v1.0
released separately. Data collection ongoing through November 30, 2026. Target: 10,000 participants
by 2027. Future releases will expand Spanish language protocols and add additional multimodal data
(imaging, genomics). Version-specific DOIs maintained. Older versions continue to be supported and
hosted.
'
retention_limit:
id: voice:retention:1
name: Data retention and disposition policy
description: 'Data Transfer and Use Agreement (DTUA) specifies two-year agreement term from start date.
Upon termination or expiration, data shall be destroyed per provider instructions with written certification
required within 30 days. Recipient may retain one copy to comply with records retention requirements
under law, regulation, institutional policy, and for research integrity and verification. Ongoing
restrictions apply to retained copies. Provider may unilaterally amend agreement if federal sponsor
requires revision. Health Data Nexus retains dataset as long as useful for research purposes, possibly
indefinitely. Version-specific DOIs maintained for all historical versions.
'
version_access:
id: voice:versionaccess:1
name: Version access policy
description: 'All dataset versions available through PhysioNet at https://physionet.org/content/b2ai-voice/
with version-specific DOIs. DOI for latest version: https://doi.org/10.13026/37yb-1t42. Earlier versions
also available on Health Data Nexus at https://healthdatanexus.ai/content/b2ai-voice/1.0/. By default,
older versions continue to be supported, hosted, and made available. Dataset publishers reserve right
to remove access to older versions. Each version has unique DOI.
'
extension_mechanism:
id: voice:extension:1
name: Dataset extension and contribution mechanisms
description: 'Derivative datasets can be published on Health Data Nexus referencing original source
under same access conditions. Open-source repositories (b2aiprep, SenseLab) have discussion forums,
issue pages, and pull request mechanisms for contributing improvements to preprocessing code. REDCap
data dictionary available at https://github.com/eipm/bridge2ai-redcap with MIT license. Bridge2AI-Voice
documentation at https://github.com/eipm/bridge2ai-docs. Future augmentations to voice collection
protocol coordinated through consortium.
'
ethical_reviews:
- id: voice:ethicalreview:1
description: 'IRB review and approval obtained from University of South Florida Single IRB with subsite
IRB approvals through Single IRB process. Study reviewed and approved for human subjects research.
Bioethics guidance integrated throughout study design and conduct by consortium bioethicists (Belisle-Pipon,
Ravitsky). Ethics module develops new guidelines for consenting to voice data collection, voice data
sharing, and utilization in context of voice AI technology. Project addresses ethical issues from
voice data generation through clinical adoption and downstream health decisions.
'
human_subject_research:
id: voice:hsr:1
name: Bridge2AI-Voice Human Subjects Research
description: 'Observational study (cross-sectional design) involving direct collection of voice recordings,
questionnaire responses, and EHR linkage from human participants. Prospective informed consent obtained
from all participants. IRB approved by University of South Florida Single IRB with subsite IRBs through
Single IRB process. Not a drug or medical device study. No data monitoring committee appointed. Study
ID: OT2OD032720. Involves human subjects: Yes. IRB approval: University of South Florida Institutional
Review Board (Single IRB). Special populations: Pediatric participants (aged 2-18 years) at Hospital
for Sick Children. Regulatory compliance: HIPAA Safe Harbor, 45 CFR 46 (Common Rule), Certificate
of Confidentiality, OMB Memorandum M-07-16.
'
involves_human_subjects: true
informed_consent:
- id: voice:consent:1
description: 'All participants duly informed and provided prospective informed consent for data collection
and use via IRB-approved consent process. Consent includes: authorization for voice data collection
and speaking tasks, demographic and medical history questionnaires, disease- specific validated questionnaires,
access to medical information through EHR platforms, and permission to share research data. Data that
poses low risk of re-identification shared in registered access; heightened re-identification risk
data shared through controlled access mechanism. No restriction on commercial vs. non-commercial use;
no geographic restriction; no restriction to specific research type. Participants may withdraw at
any point before voice data collection is completed.
'
at_risk_populations:
id: voice:atrisk:1
name: Pediatric participants protection
description: 'Pediatric participants (aged 2-18) require additional protections. Pediatric data collected
exclusively at Hospital for Sick Children (SickKids) with age-appropriate protocols. Distinct IRB
considerations for pediatric cohort. Pediatric questionnaires and acoustic tasks specifically designed
for different age groups (2-4, 4-6, 6-10, 10+ years). Minimum age for adult dataset is 18 years. Pediatric
dataset released separately with additional privacy precautions.
'
license_and_use_terms:
id: voice:license:1
name: Bridge2AI Voice Registered Access License
description: 'Public access dataset distributed through PhysioNet under Bridge2AI Voice Registered Access
License. Only registered users who sign the specified Data Use Agreement (Bridge2AI Voice Registered
Access Agreement) can access files. Data covered under Certificate of Confidentiality which must be
asserted against compulsory legal demands. Raw audio available through controlled access only via
Data Access Compliance Office (DACO) requiring distinct application and DTUA signed by institutional
official. Recipient must adhere to PhysioNet requirements managed by MIT Laboratory for Computational
Physiology. Recipient encouraged to publish results in open-access journals. No export controls apply.
Commercial and non-commercial research use permitted for Authorized Researchers.
'
ip_restrictions:
id: voice:ip:1
name: Intellectual property restrictions
description: 'Recipient shall not develop or use any intellectual property protections, database rights,
or related rights in any element of the dataset or in conclusions derived from it that would prevent
or block access to any element of the dataset or those conclusions. Fort Lauderdale Agreement principles
apply. Recipients must respect Open Science principles. Recipient shall not disclose, release, sell,
rent, or lease data to third parties without prior written consent of Provider.
'
regulatory_restrictions:
id: voice:regulatory:1
name: Regulatory restrictions and compliance requirements
description: 'Dataset covered under Certificate of Confidentiality per 45 CFR 46 (Common Rule), which
must be asserted against compulsory legal demands such as court orders and subpoenas. HIPAA Safe Harbor
de-identification standards applied. Data covered as Personally Identifiable Information per OMB Memorandum
M-07-16. No export control restrictions apply. Authorized Researchers must comply with all applicable
laws and regulations. DTUA specifies two-year term with data destruction requirements.
'
is_deidentified:
id: voice:deident:1
name: HIPAA Safe Harbor de-identification
description: 'HIPAA Safe Harbor de-identification applied. All direct identifiers removed (names, civic
addresses, social security numbers). Indirect identifiers removed where creating significant re-identification
risk (geographic/demographic identifiers, household composition, cultural identity). Non-identifying
sensitive information removed (household income, mental health status, traumatic life experiences).
All raw voice data removed from public release. Sensitive fields per REDCap data dictionary removed.
Direct identifiers removed: Yes. HIPAA de-identification rules applied: Yes. Dates rebased: Yes. Geographic
information removed/generalized: Yes. Narrative text fields removed: Yes. K-anonymization: No.
'
external_resources:
- id: voice:resource:1
name: PhysioNet Dataset Landing Page
description: Primary registered access distribution platform for adult and pediatric datasets
external_resources:
- https://physionet.org/content/b2ai-voice/
- id: voice:resource:2
name: Bridge2AI-Voice Project Documentation
description: Comprehensive project documentation, collection methods, governance, healthsheet
external_resources:
- https://docs.b2ai-voice.org
- id: voice:resource:3
name: Bridge2AI-Voice GitHub Documentation Repository
description: Source code for docs and dashboard at docs.b2ai-voice.org (MIT license)
external_resources:
- https://github.com/eipm/bridge2ai-docs
- id: voice:resource:4
name: b2aiprep Software Library
description: 'Open source library (Apache-2.0 license) for preprocessing raw audio waveforms into parquet
files and merging source data into phenotype files
'
external_resources:
- https://github.com/sensein/b2aiprep
- id: voice:resource:5
name: Bridge2AI REDCap Data Dictionary
description: REDCap data dictionary and metadata (MIT license)
external_resources:
- https://github.com/eipm/bridge2ai-redcap
- id: voice:resource:6
name: NIH RePORTER Project Details
description: Federal grant information for project 3OT2OD032720-01S3
external_resources:
- https://reporter.nih.gov/project-details/11376382
- id: voice:resource:7
name: Health Data Nexus
description: Alternative platform providing cloud compute alongside earlier dataset versions
external_resources:
- https://healthdatanexus.ai/content/b2ai-voice/1.0/
- id: voice:resource:8
name: Zenodo Archive (REDCap data dictionary)
description: 'Bensoussan Y. et al. (2024). eipm/bridge2ai-redcap. Zenodo archive.
'
external_resources:
- https://doi.org/10.5281/zenodo.13834653
- id: voice:resource:9
name: Interspeech 2024 Protocol Publication
description: 'Bensoussan et al. "Developing Multi-Disorder Voice Protocols: A team science approach
involving clinical expertise, bioethics, standards, and DEI." Proc. Interspeech 2024.
'
external_resources:
- https://doi.org/10.21437/Interspeech.2024-1926
- id: voice:resource:10
name: Data Access Compliance Office (DACO)
description: Contact for controlled access to raw audio data and institutional DTUA
external_resources:
- mailto:DACO@b2ai-voice.org
- id: voice:resource:11
name: Bridge2AI Program
description: Parent NIH Common Fund program supporting AI-ready biomedical datasets
external_resources:
- https://bridge2ai.org
- id: voice:resource:12
name: Bridge2AI Voice Scholars Training
description: Training opportunities for using the dataset
external_resources:
- https://www.b2aivoicescholars.org/