VOICE Dataset Documentation

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

DescriptionIDName
Integrate the use of voice as a biomarker of health in clinical care by generating a substantial multi-institutional, ethically sourced, and diverse voice database linked to multimodal health biomarkers to fuel voice AI research and build predictive models to assist in screening, diagnosis, and treatment of a broad range of diseases.
purpose-001Integrate voice as biomarker in clinical care
Create an ethically sourced flagship dataset to enable future research in artificial intelligence and support critical insights into the use of voice as a biomarker of health, addressing the pressing need for large, high quality, multi-institutional and diverse voice databases linked to other health biomarkers.
purpose-002Create ethically sourced flagship dataset
Establish standards, best practices, and guidelines for voice data collection and analysis to advance the field of acoustic biomarkers by developing new standards that are AI/ML friendly and enable voice to emerge as a biomarker of health.
purpose-003Establish standards for voice data
Ensure patient protection through ethical and fairness principles, create safe and innovative infrastructures to disseminate ethically sourced data, and guide the development of voice AI by addressing important issues related to patient privacy protection, ethical and fair representation of populations, and clinical accuracy.
purpose-004Promote ethical AI development
DescriptionIDName
Funded through National Institutes of Health grant 3OT2OD032720-01S3 (Bridge2AI: Voice as a Biomarker of Health - Building an ethically sourced, bioaccoustic database to understand disease like never before). Opportunity Number: OTA-21-008. Project dates: September 1, 2022 to November 30, 2026. Total funding in 2025: $4,660,942 (Direct Costs: $4,072,321, Indirect Costs: $588,621). Administered by NIH Office of the Director through the Bridge2AI Program. Study Section: Data Coordination, Mapping, and Modeling [DCMM]. Congressional District: 15, Tampa, FL.
funder-001NIH Office of the Director
NIBIB supports PhysioNet managed by MIT Laboratory for Computational Physiology under NIH grant number R01EB030362, which serves as a distribution platform for the Bridge2AI-Voice dataset.
funder-002National Institute of Biomedical Imaging and Bioengineering
πŸ“Š

Composition

What do the instances represent?

  1. ID
    instance-001
    Name
    Adult participants with voice-affecting conditions
    Description
    Adult participants presenting at specialty clinics and institutions across five sites in North America. Participants were selected based on membership to five predetermined disease cohort groups: Respiratory disorders, Voice disorders, Neurological disorders, Mood disorders, and Pediatric. As of v1.1, only data from the adult cohort is available. The initial release (v1.0) provides 306 participants with 12,523 recordings collected through standardized protocols.
    Instance Type
    Human participants recruited from specialty clinics at multi-institutional sites. Data collection conducted between September 1, 2022 and November 30, 2026 through IRB-approved protocols with informed consent.
DescriptionIDName
Participants with laryngeal disorders including laryngeal cancers, vocal fold paralysis, and benign laryngeal lesions that affect vocal fold shape, mass, density, and tension resulting in changes in vibratory function and phonation.
subpop-001Voice Disorders cohort
Participants with conditions such as Alzheimer's disease, Parkinson's disease, stroke, and ALS exhibiting voice and speech changes including slowed speech, low frequency, monotonous speech, vocal tremor, dysarthria, and aphasia.
subpop-002Neurological and Neurodegenerative Disorders cohort
Participants with depression, schizophrenia, bipolar disorders, and anxiety disorders showing vocal changes such as decreased fundamental frequency (f0), monotonous speech, and anxiety-related increases in F0.
subpop-003Mood and Psychiatric Disorders cohort
Participants with respiratory conditions including pneumonia, COPD, heart failure, and obstructive sleep apnea where respiratory sounds, cough sounds, and voice are used for diagnostic and monitoring purposes.
subpop-004Respiratory Disorders cohort
Pediatric participants with voice and speech disorders including autism spectrum disorder and speech delays. Data collection for this cohort addresses ethical concerns and acquisition challenges specific to pediatric populations. Note: As of v1.1, pediatric data not yet released.
subpop-005Pediatric cohort
Access UrlsDescriptionIDName
https://physionet.org/content/b2ai-voice/
Spectrograms stored in Parquet format (spectrograms.parquet). Each element contains participant_id, session_id, task_name, and 513xN dimension spectrogram array. Compatible with Python datasets library and common data science tools. Can be loaded with Dataset.from_parquet() and plotted using librosa.power_to_db().
format-001Parquet for spectrograms
https://physionet.org/content/b2ai-voice/
Mel-frequency cepstral coefficients stored in Parquet format (mfcc.parquet). Contains 60xN dimension MFCC arrays derived from spectrograms. Compatible with Python datasets library and common data science tools. Added in v1.1 release.
format-002Parquet for MFCCs
https://physionet.org/content/b2ai-voice/
Tab-delimited phenotype file (phenotype.tsv) with one row per unique participant. Contains demographics, acoustic confounders, and validated questionnaire responses. Accompanied by JSON data dictionary (phenotype.json) with column descriptions. Can be loaded with pd.read_csv(sep="\t").
format-003TSV for phenotype data
https://physionet.org/content/b2ai-voice/
Tab-delimited static features file (static_features.tsv) with one row per unique recording. Contains OpenSMILE, Praat, Parselmouth, and torchaudio features. Accompanied by JSON data dictionary (static_features.json) with feature descriptions.
format-004TSV for static features
Contact DACO@b2ai-voice.org
Original raw audio waveforms available through controlled access only. Interested users contact DACO@b2ai-voice.org for application process. Disseminated through Data Access Compliance Office with formal vetting and approval process. Covered under Certificate of Confidentiality.
format-005Raw audio (controlled access only)
πŸ”

Collection Process

How was the data acquired?

Bridge2AI-Voice
Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information
The Bridge2AI-Voice project seeks to create an ethically sourced flagship dataset to enable future research in artificial intelligence and support critical insights into the use of voice as a biomarker of health. The human voice contains complex acoustic markers which have been linked to important health conditions including dementia, mood disorders, and cancer. When viewed as a biomarker, voice is a promising characteristic to measure as it is simple to collect, cost-effective, and has broad clinical utility. This comprehensive collection provides voice recordings with corresponding clinical information from participants selected based on known conditions which manifest within the voice waveform including voice disorders, neurological disorders, mood disorders, and respiratory disorders. The dataset is designed to fuel voice AI research, establish data standards, and promote ethical and trustworthy AI/ML development for voice biomarkers of health. Data collection occurs through a multi-institutional collaborative effort using standardized protocols, custom smartphone applications, and rigorous ethical oversight. The initial release (v1.0) provides 12,523 recordings for 306 participants collected across five sites in North America, with derived features such as spectrograms, MFCCs, acoustic features, and clinical phenotype data. Raw audio data is available through controlled access to protect participant privacy.
en
  • voice biomarker
  • acoustic biomarker
  • Bridge2AI
  • voice AI
  • voice disorders
  • neurological disorders
  • neurodegenerative disorders
  • mood disorders
  • psychiatric disorders
  • respiratory disorders
  • pediatric voice disorders
  • speech disorders
  • Parkinson's disease
  • Alzheimer's disease
  • depression
  • schizophrenia
  • bipolar disorder
  • stroke
  • ALS
  • autism
  • speech delay
  • laryngeal cancer
  • vocal fold paralysis
  • pneumonia
  • COPD
  • heart failure
  • obstructive sleep apnea
  • spectrogram
  • MFCC
  • mel-frequency cepstral coefficients
  • OpenSMILE
  • Praat
  • Parselmouth
  • federated learning
  • ethical AI
  • multimodal health data
  • electronic health records
  • EHR
  • radiomics
  • genomics
  • FAIR principles
  • CARE principles
  • PhysioNet
  • Health Data Nexus
DescriptionIDName
Address the lack of large, high quality, multi-institutional and diverse voice databases linked to multimodal health biomarkers (demographics, imaging, genomics, risk factors) necessary to fuel voice AI research and answer tangible clinical questions.
gap-001Lack of diverse voice databases
Overcome limitations in existing voice and psychiatric disorder research that has relied on small datasets with limited demographic diversity reporting, lack of standardized data collection protocols precluding meta-analysis, and possible confounders limiting external validity and clinical usability.
gap-002Limited psychiatric disorder datasets
Fill the gap in pediatric voice and speech analysis research, which is sparser partly due to ethical concerns and challenges in data acquisition for this cohort, particularly for autism and speech delay detection.
gap-003Pediatric data scarcity
Establish missing standards for voice data collection, acoustic analysis, and ethical frameworks for consenting to voice data collection, sharing, and utilization in the context of voice AI technology development and clinical adoption.
gap-004Missing voice data standards
Address important issues related to patient privacy protection, ethical and fair representation of populations, and clinical accuracy that are arising as voice AI gains attention from multi-nationals and the tech world.
gap-005Privacy and ethics concerns
RoleNameORCIDAffiliation
ContributorYael Emilie Bensoussancreator-001-
ContributorJean-Christophe BΓ©lisle-Piponcreator-002-
ContributorDavid A. Dorrcreator-003-
ContributorSatrajit Sujit Ghoshcreator-004-
ContributorPhilip R.O. Paynecreator-005-
ContributorMaria Ellen Powellcreator-006-
ContributorAnais Rameaucreator-007-
ContributorVardit Ravitskycreator-008-
ContributorAlexandros Sigarascreator-009-
ContributorOlivier Elementocreator-010-
ContributorAlistair Johnsoncreator-011-
ContributorJennifer Siucreator-012-
ContributorBridge2AI-Voice Consortiumcreator-013-
DescriptionIDName
Contains derived features from voice recordings including spectrograms (513xN dimensions), MFCCs (60xN dimensions), acoustic features (OpenSMILE), phonetic and prosodic features (Parselmouth and Praat), and transcriptions (OpenAI Whisper). Also includes phenotype data with demographics, acoustic confounders, and responses to validated questionnaires. Available through PhysioNet with registered access requiring data use agreement. HIPAA Safe Harbor identifiers removed, state/province removed, country retained. Audio waveforms omitted, only derived features available. Free speech transcripts removed to protect privacy.
subset-001Public Access Dataset (PhysioNet Registered Access)
Original raw audio waveforms available through controlled access only to protect participant privacy. Interested users can request access by contacting DACO@b2ai-voice.org. Raw audio data disseminated through Data Access Compliance Office (DACO) requiring distinct application and formal vetting. Covered under Certificate of Confidentiality which must be asserted against compulsory legal demands.
subset-002Controlled Access Raw Audio Dataset
Time-frequency representations computed using short-time FFT with 25ms window, 10ms hop length, 512-point FFT. Each element contains participant_id, session_id, task_name, and 513xN dimension spectrogram. Available in spectrograms.parquet file.
subset-003Spectrograms Parquet file
60 Mel-frequency cepstral coefficients extracted from spectrograms. Each element has 60xN dimension where N is proportional to audio length. Available in mfcc.parquet file (added in v1.1).
subset-004MFCC Parquet file
Tab-delimited file with one row per unique participant containing demographics, acoustic confounders, and validated questionnaire responses. Accompanied by phenotype.json data dictionary. Available in phenotype.tsv.
subset-005Phenotype TSV file
Tab-delimited file with one row per unique recording containing OpenSMILE, Praat, Parselmouth, and torchaudio features. Accompanied by static_features.json data dictionary. Available in static_features.tsv.
subset-006Static Features TSV file
  1. ID
    sampling-001
    Name
    Targeted disease cohort recruitment
    Description
    Patients presenting at specialty clinics and institutions were screened for inclusion and exclusion criteria prior to their visit by project investigators. Participants were selected based on membership to five predetermined disease cohort groups to ensure representation across conditions affecting voice: (1) Voice Disorders - laryngeal cancers, vocal fold paralysis, benign laryngeal lesions; (2) Neurological and Neurodegenerative Disorders - Alzheimer's, Parkinson's, stroke, ALS; (3) Mood and Psychiatric Disorders - depression, schizophrenia, bipolar disorders; (4) Respiratory disorders - pneumonia, COPD, heart failure, obstructive sleep apnea; (5) Pediatric diseases - autism, speech delay.
    Sample
    True
    Random Sampling
    False
    Representative Sample
    False
    Strategies
    • Targeted recruitment from specialty clinics representing five disease cohort categories
    • Screening based on known conditions manifesting within voice waveform
    • Multi-institutional enrollment across five sites in North America
    • Standardized inclusion and exclusion criteria applied by investigators
    • Protocols developed through team science approach involving clinical expertise, bioethics, standards, and DEI
DescriptionIDName
Data collection conducted using a custom smartphone application on tablet with headset used when possible. Standardized protocol for data collection adopted across all sites. Single session sufficient for most participants, though subset required multiple sessions resulting in more than one session per participant in dataset. Software enables non-invasive, user-friendly, high quality voice data collection while minimizing human manipulation and includes integrated acoustic amplifiers and acoustic quality standardization.
collection-001Smartphone application data collection
Multi-institutional data collection across five specialty clinic sites in North America. Patients presenting at clinics screened for eligibility, consented for data collection initiative and data sharing. Enrollment occurred between September 1, 2022 and November 30, 2026 under IRB-approved protocols with informed consent processes developed specifically for voice data collection.
collection-002Multi-institutional clinic recruitment
Data collection protocol involved: (1) demographic information collection, (2) health questionnaires, (3) targeted questionnaires about known voice confounders, (4) disease- specific information, (5) voice recording tasks such as sustained phonation of vowel sounds, (6) conventional acoustic tasks including respiratory sounds, cough sounds, and free speech prompts. Data exported and converted from REDCap using open source b2aiprep library.
collection-003Standardized data collection protocol
Cloud infrastructure for automated voice data collection developed to allow analysis of multi-institutional data while minimizing data sharing and preserving patient privacy. Federated learning technology implemented to protect data privacy during collaborative analysis.
collection-004Federated learning infrastructure
DescriptionIDName
Voice recording tasks capturing voice, speech, and language data relating to health. Tasks include sustained phonation of vowel sounds (e.g., prolonged /e/), conventional acoustic tasks including respiratory sounds, cough sounds, and free speech prompts. Recordings performed using custom smartphone application with headset when possible to standardize acoustic quality.
acquisition-001Voice recording tasks
Self-reported demographic and medical history questionnaires completed by participants who consent. Disease-specific validated questionnaires administered. Targeted questionnaires inquiring about known confounders for voice including acoustic confounders.
acquisition-002Questionnaires and surveys
Electronic health record (EHR) access for participants who consent, permitting investigators to access medical information through EHR platforms to perform gold standard validation of diagnoses and symptoms. Linkage to multimodal health biomarkers including radiomics and genomics when available.
acquisition-003Electronic health record access
DescriptionIDNamePreprocessing Details
Raw audio preprocessing by converting to monaural and resampling to 16 kHz with Butterworth anti-aliasing filter applied. Standardization ensures consistent format across all recordings for downstream feature extraction.
preproc-001Audio standardizationConversion to monaural audio, Resampling to 16 kHz sampling rate, Butterworth anti-aliasing filter applied, Standardized format enables consistent feature extraction
Spectrogram extraction - Time-frequency representations computed using short-time Fast Fourier Transform (FFT) with 25ms window size, 10ms hop length, and 512-point FFT. Output spectrograms have 513xN dimensions where N is proportional to audio length.
preproc-002Spectrogram generationShort-time FFT with 25ms window, 10ms hop length, 512-point FFT, Output dimension 513xN, Can be plotted in decibels by converting from power representation
Mel-frequency cepstral coefficients (MFCC) extraction - 60 MFCCs extracted from spectrograms. MFCCs capture perceptually-relevant spectral envelope characteristics important for voice analysis. Output dimension 60xN. Added in v1.1 release.
preproc-003MFCC extraction60 MFCC coefficients extracted, Derived from spectrograms, Output dimension 60xN, Captures spectral envelope characteristics
Acoustic feature extraction using OpenSMILE (Speech and Music Interpretation by Large-space Extraction), capturing temporal dynamics and acoustic characteristics. Features provided in static_features.tsv with one row per unique recording.
preproc-004OpenSMILE feature extractionOpenSMILE feature extraction, Temporal dynamics captured, Acoustic characteristics quantified, Static features per recording
Phonetic and prosodic feature computation using Parselmouth and Praat, providing measures of fundamental frequency (f0), formants, and voice quality. Features documented in static_features.json data dictionary.
preproc-005Praat and Parselmouth featuresParselmouth and Praat feature extraction, Fundamental frequency (F0) measurement, Formant analysis, Voice quality metrics
Transcription generation using OpenAI's Whisper Large model. Automated speech recognition applied to audio recordings. Free speech transcripts subsequently removed from public release to protect participant privacy.
preproc-006Automated transcriptionOpenAI Whisper Large model, Automated transcription of audio, Free speech transcripts removed for privacy, Structured task transcripts may be retained
Data export and conversion from REDCap using open source b2aiprep library developed by the team. Phenotype data merged into tab-delimited format with data dictionary (phenotype.json) providing column descriptions.
preproc-007REDCap export and conversionREDCap data export, b2aiprep library conversion, Tab-delimited phenotype file generation, JSON data dictionary creation
Cleaning DetailsDescriptionIDName
HIPAA Safe Harbor compliance, 18 identifier categories removed, Geographic data limited to country level, Date precision limited to year, Biometric identifiers removed
HIPAA Safe Harbor de-identification applied. Identifiers removed include: names, geographic locators (state/province removed, country retained), dates at resolution finer than years, phone/fax numbers, email addresses, IP addresses, Social Security Numbers, medical record numbers, health plan beneficiary numbers, device identifiers, license numbers, account numbers, vehicle identifiers, website URLs, full face photos, biometric identifiers, and any unique identifiers.
cleaning-001HIPAA Safe Harbor de-identification
Audio waveforms excluded from public release, Derived features only in public dataset, Free speech transcripts removed, Raw audio requires controlled access
Privacy protection measures for public release - Audio waveforms omitted from public dataset, only derived features (spectrograms, MFCCs, acoustic features) made available. Free speech transcripts removed. Raw audio available only through controlled access with DACO approval.
cleaning-002Privacy protection for public release
Standardized collection protocols, Common smartphone application, REDCap data management, Multi-site harmonization
Data standardization across multi-institutional sites through use of standardized protocols, common data collection application, and REDCap data management system. Ensures consistency and quality across five collection sites.
cleaning-003Multi-site data standardization
DescriptionIDMaintainer DetailsName
Multidisciplinary consortium responsible for dataset maintenance including data collection, curation, standards development, ethics oversight, and distribution. Led by University of South Florida (Yael Bensoussan, Contact PI) with multi-institutional partnerships.
maintainer-001University of South Florida (lead institution, Tampa, FL), Multi-institutional data collection sites (five sites in North America), MIT Laboratory for Computational Physiology (PhysioNet distribution), Data Access Compliance Office (DACO) for controlled access, Bioethics and social science teams, Standards and tool development teams (acoustic amplifiers, quality standardization), Workforce development and education teams, Cloud infrastructure and federated learning teamsBridge2AI-Voice Consortium
PhysioNet platform managed by MIT Laboratory for Computational Physiology serves as primary distribution mechanism for public access dataset. Supported by National Institute of Biomedical Imaging and Bioengineering (NIBIB) under NIH grant R01EB030362.
maintainer-002MIT Laboratory for Computational Physiology, PhysioNet platform infrastructure, NIBIB grant R01EB030362 support, Registered access system management, Data use agreement administrationPhysioNet / MIT Laboratory for Computational Physiology
ID
retention-001
Name
Data retention and disposition
Description
Data Transfer and Use Agreement specifies retention requirements. Upon termination or expiration of agreement (two years after start date, project completion, ethics approval expiration, or provider termination), data shall be destroyed per provider instructions with written certification required within 30 days. Recipient may retain one copy to extent necessary to comply with records retention requirements under law, regulation, institutional policy, and for research integrity and verification purposes. Restrictions apply to archival copies as long as recipient holds data. Provider may unilaterally amend agreement if federal sponsor requires revision.
Retention Details
  • Two-year agreement term from start date
  • Data destruction required upon termination unless retention justified
  • One archival copy permitted for compliance and verification
  • Written certification of destruction required within 30 days
  • Ongoing restrictions apply to retained copies
  • Provider may unilaterally amend if federal sponsor requires
  • Termination if recipient objects to amendments
  • Disposition instructions in DTUA Attachment 1
DescriptionIDNameSensitive Elements PresentSensitivity Details
Voice recordings contain personally identifiable information and are considered biometric identifiers under HIPAA. Voice is increasingly recognized as a biomarker by tech companies (Google, Amazon, Mozilla, Apple) raising privacy concerns. Raw audio waveforms omitted from public release to protect privacy. Available only through controlled access with DACO approval and formal vetting process.
sensitive-001Voice as biometric identifierTrueVoice as biometric identifier, Raw audio waveforms, Speech patterns and characteristics, Controlled access required for raw audio, Privacy protection through omission from public dataset
Electronic health record (EHR) data accessed with participant consent for gold standard validation of diagnoses and symptoms. Medical information linked to voice data provides sensitive health information including disease-specific clinical data, multimodal health biomarkers, radiomics, and genomics data.
sensitive-002Health information from EHRTrueEHR medical information, Diagnoses and symptoms (validated), Disease-specific clinical data across five cohort categories, Multimodal health biomarkers (radiomics, genomics), Linked to voice recordings
Demographic information and geographic data collected but de-identified for public release. State and province removed, only country retained. Protected by HIPAA Safe Harbor de-identification standards. Dates limited to year precision.
sensitive-003Demographic and geographic dataTrueDemographic data (de-identified), Geographic information (state/province removed, country retained), Medical history questionnaires, Disease-specific validated questionnaires, Acoustic confounders questionnaires
Dataset covered under Certificate of Confidentiality which must be asserted against compulsory legal demands such as court orders and subpoenas for identifying information or characteristics of research participants. Provides additional legal protections beyond standard de-identification. Mentioned explicitly in Data Transfer and Use Agreement.
sensitive-004Certificate of Confidentiality protectionTrueCertificate of Confidentiality coverage, Protection against compulsory legal demands, Court order and subpoena protection, Participant identification protection, Must be asserted by data holders
Data is Personally Identifiable Information as defined in OMB Memorandum M-07-16, not covered under HIPAA, FERPA, or similar laws requiring special terms. Security controls adequate to protect PII required including administrative, physical, and technical safeguards to secure electronic protected health information.
sensitive-005Personally Identifiable InformationTruePersonally Identifiable Information (OMB M-07-16), Not covered under HIPAA/FERPA, Administrative, physical, technical safeguards required, Electronic protected health information security requirements, Secure storage and destruction protocols
DescriptionExternal ResourcesIDName
Primary distribution platform for public access dataset with registered access
https://physionet.org/content/b2ai-voice/resource-001PhysioNet Dataset Landing Page
Comprehensive project documentation and resources
https://docs.b2ai-voice.orgresource-002Bridge2AI-Voice Project Documentation
Open source code repository including b2aiprep library and documentation dashboard
https://github.com/eipm/bridge2ai-docsresource-003Bridge2AI-Voice GitHub Repository
Federal grant information and project details for grant 3OT2OD032720-01S3
https://reporter.nih.gov/project-details/11376382resource-004NIH RePORTER Project Details
Alternative data repository platform hosting v1.0 dataset
https://healthdatanexus.ai/content/b2ai-voice/1.0/resource-005Health Data Nexus
Additional dataset documentation, software releases, and Bridge2AI Voice REDCap v3.20.0
https://doi.org/10.5281/zenodo.13834653, https://doi.org/10.5281/zenodo.14148755resource-006Zenodo Archive
Research resource for complex physiologic signals (Goldberger et al. 2000, Circulation)
https://physionet.orgresource-007PhysioNet Platform
Publication describing multi-disorder voice protocol development through team science approach involving clinical expertise, bioethics, standards, and DEI (Rameau et al. 2024)
https://doi.org/10.21437/Interspeech.2024-1926resource-008Interspeech 2024 Protocol Publication
Contact for controlled access to raw audio data
mailto:DACO@b2ai-voice.orgresource-009Data Access Compliance Office
Parent NIH Common Fund program supporting AI-ready biomedical datasets
https://bridge2ai.orgresource-010Bridge2AI Program
Open source library for preprocessing raw audio and phenotype data (Bevers et al.)
https://github.com/sensein/b2aiprepresource-011b2aiprep Software Library
DOI for latest version of dataset (v2.0.1 as of August 2025)
https://doi.org/10.13026/37yb-1t42resource-012Latest Version DOI
Version-specific DOI for v1.1 release
https://doi.org/10.13026/249v-w155resource-013Version 1.1 DOI
πŸš€

Uses

What (other) tasks could the dataset be used for?

DescriptionIDName
Enable development of AI/ML predictive models for screening, diagnosis, and treatment of voice disorders including laryngeal cancers, vocal fold paralysis, and benign laryngeal lesions, leveraging acoustic changes in phonation resulting from changes in vocal fold vibratory function.
task-001Voice disorder diagnosis and screening
Support machine learning models for neurological and neurodegenerative disorders including Alzheimer's disease, Parkinson's disease, stroke, and ALS, detecting voice and speech changes such as slowed speech, low frequency, monotonous speech, vocal tremor, dysarthria, and aphasia.
task-002Neurological disorder detection
Develop AI algorithms for mood and psychiatric disorder detection including depression, schizophrenia, and bipolar disorders, identifying vocal markers such as decreased fundamental frequency, monotonous speech patterns, and anxiety-related increases in F0.
task-003Mood and psychiatric disorder screening
Create machine learning models for respiratory disorder screening and therapeutic monitoring using respiratory sounds, cough sounds, and voice, applicable to conditions such as pneumonia, COPD, heart failure, and obstructive sleep apnea.
task-004Respiratory disorder monitoring
Build AI models for pediatric voice and speech disorder detection including autism spectrum disorder and speech delays, addressing the relative scarcity of pediatric voice data and associated ethical challenges.
task-005Pediatric disorder detection
Promote application of AI/ML for voice research through workforce development, curriculum creation, and fostering collaborations especially with researchers from underserved communities, building bridges between medical voice research, acoustic engineers, and the AI/ML community.
task-006Workforce development
DescriptionIDName
Primary intended use is development and validation of AI/ML models for voice as a biomarker of health, supporting screening, diagnosis, and treatment of voice disorders, neurological disorders, mood disorders, respiratory disorders, and pediatric speech disorders.
use-001Voice AI model development
Research into acoustic biomarkers and development of standards for voice data collection and analysis. Establishing best practices for AI/ML-friendly voice datasets and contributing to the field's maturation as a clinical diagnostic modality.
use-002Acoustic biomarker research
Training and education in voice AI research through workforce development initiatives, curriculum creation on voice biomarkers of health and development/validation/implementation of AI models that are FAIR and CARE, fostering collaborations between medical voice researchers, acoustic engineers, and AI/ML specialists, especially from underserved communities.
use-003Workforce development and training
Multimodal health research combining voice data with EHR information, radiomics, genomics, and other health biomarkers to understand complex disease relationships and improve diagnostic accuracy through integrated analysis of diverse data types.
use-004Multimodal health research
Model dataset for ethical AI development in healthcare, demonstrating integration of bioethics guidance, ethical data collection practices, informed consent processes, privacy protection through federated learning, and trustworthy AI/ML development from data generation through clinical adoption and downstream health decisions.
use-005Ethical AI development model
RoleNameORCIDAffiliation
ContributorDirect clinical decision-makingdiscouraged-001-
ContributorRe-identification attemptsdiscouraged-002-
ContributorData use agreement violationsdiscouraged-003-
ContributorSurveillance or discriminationdiscouraged-004-
ID
license-001
Name
Bridge2AI Voice Registered Access License
Description
Public access dataset distributed through PhysioNet under Bridge2AI Voice Registered Access License. Only registered users who sign the specified Data Use Agreement (Bridge2AI Voice Registered Access Agreement) can access files. Data covered under Certificate of Confidentiality which must be asserted against compulsory legal demands such as court orders and subpoenas. Raw audio data available through controlled access only via Data Access Compliance Office (DACO) requiring distinct application. Agreement term is two years after start date, upon completion of project, upon termination, or upon expiration of applicable ethics approval, whichever occurs first. Recipient must adhere to PhysioNet requirements managed by MIT Laboratory for Computational Physiology, supported by NIBIB under grant R01EB030362. Data shall be destroyed upon termination per provider instructions with written certification required within 30 days.
License Terms
  • Registered access required with institutional email
  • Data Use Agreement signature mandatory
  • Use restricted to authorized persons listed in agreement
  • No sharing with third parties without prior written consent
  • Appropriate administrative, technical, physical safeguards required
  • Compliance with applicable laws, rules, regulations, professional standards
  • Public disclosure of results encouraged in open-access journals
  • Recognition of data source required in publications
  • Certificate of Confidentiality protections apply
  • Raw audio requires separate controlled access application
  • PhysioNet platform requirements apply
  • Two-year term from start date or project completion
  • Data destruction required upon termination with written certification
  • No use or disclosure other than permitted by agreement
  • Unauthorized use must be reported within 5 business days
  • IRB approval required for recipient use
πŸ“€

Distribution

How will the dataset be distributed?

Bridge2AI Voice Registered Access License
πŸ”„

Maintenance

How will the dataset be maintained?

ID
updates-001
Name
Versioned releases with ongoing data collection
Description
Dataset updated with versioned releases as data collection progresses. Initial release v1.0 published January 17, 2025 with 12,523 recordings from 306 participants (adult cohort only). v1.1 released January 17, 2025 adding MFCC features. v2.0.0 released April 16, 2025. v2.0.1 released August 18, 2025 (latest version). Latest version DOI: https://doi.org/10.13026/37yb-1t42. Version-specific DOIs also available. Data collection ongoing through November 30, 2026. As of v1.1, only adult cohort data available; pediatric cohort data planned for future releases with additional privacy precautions. Raw audio data access planned for future releases with enhanced security measures.
Frequency
Periodic versioned releases during data collection period (September 2022 - November 2026)
Update Details
  • v1.0 released January 17, 2025 - initial release with 306 participants, 12,523 recordings, adult cohort only
  • v1.1 released January 17, 2025 - added MFCC features (60xN dimensions)
  • v2.0.0 released April 16, 2025 - expanded participant cohort
  • v2.0.1 released August 18, 2025 - latest version
  • Ongoing data collection through November 30, 2026
  • Future releases planned with additional participants and pediatric cohort
  • Raw audio data access planned with additional security precautions
  • Version-specific documentation maintained
  • DOI for latest version vs version-specific DOIs
  • Files for older versions may be removed when superseded
πŸ‘₯

Human Subjects

Does the dataset relate to people?

ID
hsr-001
Name
Bridge2AI-Voice Human Subjects Research
Description
Data collection and sharing approved by University of South Florida Institutional Review Board. Participants provided written informed consent for data collection initiative and data sharing. Consent process includes authorization for voice data collection, access to medical information through EHR platforms for gold standard validation, and permission to share research data. New guidelines developed for consenting to voice data collection, voice data sharing, and utilization in context of voice AI technology. Bioethics guidance integrated throughout study design and conduct through dedicated Ethics Module. Project integrates existing scholarship, tools, and guidance with development of new standards and normative insights for identifying, anticipating, addressing, and providing guidance on ethical and trustworthy issues from voice data generation and AI/ML research through clinical adoption and downstream health decisions. Expertise drawn from law, ethics, health services, biomedical science, engineering, and scientific publications.
Involves Human Subjects
True
IRB Approval
  • University of South Florida Institutional Review Board approval
Ethics Review Board
  • University of South Florida Institutional Review Board
  • Ethics Module team providing ongoing ethical oversight
  • Bioethics and social science advisory teams
Generated on 2026-04-21 15:34:03 using Bridge2AI Data Sheets Schema