Bridge2AI-Voice
Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information
The human voice contains complex acoustic markers which have been linked to important health conditions including dementia, mood disorders, and cancer. When viewed as a biomarker, voice is a promising characteristic to measure as it is simple to collect, cost-effective, and has broad clinical utility. Recent advances in artificial intelligence have provided techniques to extract previously unknown prognostically useful information from dense data elements such as images. The Bridge2AI-Voice project seeks to create an ethically sourced flagship dataset to enable future research in artificial intelligence and support critical insights into the use of voice as a biomarker of health. Bridge2AI-Voice provides a comprehensive collection of derived data from voice recordings with corresponding clinical information, demographic data, and validated questionnaires collected across multiple North American sites. Initial releases focus on low-risk, de-identified derived data (e.g., spectrograms, acoustic features) and phenotype data to reduce re-identification risk while enabling broad research utility.
- voice
- bridge2ai
- PhysioNet
- Health Data Nexus
- Alistair Johnson
- Jean-Christophe Bélisle-Pipon
- David Dorr
- Satrajit Ghosh
- Philip Payne
- Maria Powell
- Anais Rameau
- Vardit Ravitsky
- Alexandros Sigaras
- Olivier Elemento
- Yael Bensoussan
- Bridge2AI-Voice Consortium
- PhysioNet
- ID
- physionet-b2ai-voice-1.1
- Name
- Bridge2AI-Voice v1.1 (PhysioNet)
- Title
- Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information (version 1.1)
- Description
- Derived audio representations (e.g., spectrograms, MFCCs, acoustic and phonetic/prosodic features) and associated phenotype and questionnaire data from adult participants recruited at specialty clinics across five North American sites. Raw audio is not included in this release to reduce re-identification risk. Common identifiers include participant_id, session_id, and task_name.
- DOI
- doi:10.13026/249v-w155
- Issued
- 2025-01-17
- Version
- 1.1
- Keywords
- voice
- bridge2ai
- PhysioNet
- RRID:SCR_007345
- License
- Bridge2AI Voice Registered Access License
- Created By
- PhysioNet
- Bridge2AI-Voice Consortium
- Purposes
- Name
- primary-purpose
- Response
- Enable ethically sourced, large-scale research on voice as a biomarker of health by linking derived voice representations to demographic, clinical, and questionnaire data.
- Tasks
Name Response example-task-1 Development and benchmarking of models associating voice-derived features with health conditions. example-task-2 Exploration of acoustic, phonetic, and prosodic correlates of disease using de-identified derived da... - Addressing Gaps
- Name
- gap-addressed
- Response
- Lack of an ethically sourced, clinically linked, multi-site voice dataset with robust de-identification suitable for AI research.
- Funders
Grantor Grant Name Grant Number National Institutes of Health Bridge2AI: Voice as a Biomarker of Health 3OT2OD032720-01S1 National Institute of Biomedical Imaging and Bioengineering (NIBIB), NIH PhysioNet infrastructure support R01EB030362 - Instances
- Name
- dataset-instances
- Representation
- Derived representations of voice recordings linked to participant-level phenotype and questionnaire data.
- Instance Type
- Participants and their recordings; per-recording static feature rows and per-participant phenotype rows.
- Data Type
- De-identified derived features (spectrograms, MFCCs, acoustic and phonetic/prosodic features) and structured phenotype/questionnaire data.
- Counts
- 12,523
- Label
- Clinical and demographic attributes (e.g., condition groups, validated questionnaires) are available as metadata; no explicit machine-learning labels are defined in this release.
- Sampling Strategies
- Name
- purposive-sampling
- Is Sample
- True
- Is Random
- False
- Source Data
- Adult patients recruited at specialty clinics across five North American sites
- Is Representative
- False
- Why Not Representative
- Participants were selected based on conditions known to manifest in voice, which may affect generalizability.
- Strategies
- Purposive sampling by predefined condition groups
- Missing Information
- Name
- privacy-removals
- Missing
- Free speech transcripts
- Raw audio waveforms
- Why Missing
- Removed to reduce re-identification risk and comply with HIPAA Safe Harbor.
- Subpopulations
- Name
- adult-cohort
- Identification
- Adults only in v1.1
- Distribution
- 306 participants; 12,523 recordings across predefined clinical condition groups
- Sensitive Elements
- Name
- health-data
- Description
- Contains de-identified health-related data (clinical and questionnaire information).
- Is Deidentified
- Name
- hipaa-safe-harbor
- Description
- HIPAA Safe Harbor de-identification applied; removal of identifiers (e.g., names, fine-grained dates, contact details, geographic locators below country), removal of free speech transcripts, and omission of raw audio in v1.1.
- Acquisition Methods
- Name
- acquisition-overview
- Description
- Voice recordings directly observed; phenotype/questionnaire data reported by participants; derived features computed from recordings.
- Was Directly Observed
- True
- Was Reported By Subjects
- True
- Was Inferred Derived
- True
- Was Validated Verified
- Standardized protocols and multi-site QA were used; derived features computed via established toolkits.
- Collection Mechanisms
- Name
- data-capture
- Description
- Custom tablet application and headset microphones used when possible; data exported from REDCap and converted using an open-source library.
- Data Collectors
- Name
- site-teams
- Description
- Researchers and clinicians at five North American specialty-clinic sites.
- Collection Timeframes
- Ethical Reviews
- Name
- irb-approval
- Description
- Data collection and sharing approved by the University of South Florida Institutional Review Board.
- Preprocessing Strategies
- Name
- waveform-prep-and-feature-extraction
- Description
- Waveforms converted to mono and resampled to 16 kHz with anti-aliasing; spectrograms via STFT (25 ms window, 10 ms hop, 512-point FFT, power); 60-coefficient MFCCs computed; static acoustic features extracted (e.g., openSMILE) and phonetic/prosodic features via Parselmouth/Praat; transcripts generated by OpenAI Whisper Large (free speech transcripts removed prior to release).
- Used Software
Description Name URL Open-source library used to preprocess waveforms and merge phenotype data. b2aiprep https://github.com/sensein/b2aiprep Acoustic feature extraction toolkit. openSMILE Speech analysis software for phonetics. Praat Python interface to Praat. Parselmouth Audio processing components for PyTorch. torchaudio ASR model used to generate transcripts (free speech transcripts removed before release). OpenAI Whisper Large
- Cleaning Strategies
- Name
- de-identification-and-redactions
- Description
- Removal of HIPAA Safe Harbor identifiers; removal of state/province with retention of country; removal of free speech transcripts; omission of raw audio waveforms in v1.1.
- Labeling Strategies
- Name
- transcription
- Description
- Automatic speech recognition using OpenAI Whisper Large; transcripts of free speech removed prior to release.
- Raw Sources
- Name
- raw-audio-availability
- Description
- Raw audio waveforms were collected but are not distributed in v1.1; only derived representations are provided.
- Existing Uses
Description Name {'Johnson, A., Bélisle-Pipon, J., Dorr, D., Ghosh, S., Payne, P., Powell, M., Rameau, A., Ravitsky, V., Sigaras, A., Elemento, O., & Bensoussan, Y. (2025). Bridge2AI-Voice': 'An ethically-sourced, diverse voice dataset linked to health information (version 1.1). PhysioNet. RRID:SCR_007345. https://doi.org/10.13026/249v-w155'} dataset-citation Goldberger, A., et al. (2000). PhysioBank, PhysioToolkit, and PhysioNet. Circulation. RRID:SCR_007345. platform-citation - Use Repository
- Name
- project-and-platform-links
- Description
- Project Documentation Site
- https://docs.b2ai-voice.org
- DOI Landing For V1.1 On Physionet
- https://doi.org/10.13026/249v-w155
- Other Tasks
- Future Use Impacts
- Name
- derived-only-constraints
- Description
- The absence of raw audio may limit certain analyses (e.g., new feature extraction requiring original waveforms) but reduces re-identification risk.
- Discouraged Uses
- Distribution Formats
- Name
- formats
- Description
- Parquet
- TSV
- JSON
- Distribution Dates
- Name
- initial-release
- Description
- 2025-01-17
- License And Use Terms
- Name
- access-and-licensing
- Description
- Platform
- PhysioNet
- Access
- Registered/Restricted Access; only registered users who sign the Bridge2AI Voice Registered Access Agreement may access files.
- License
- Bridge2AI Voice Registered Access License; applicable Data Use Agreement required.
- Ip Restrictions
- Regulatory Restrictions
- Maintainers
- Name
- maintainers
- Description
- PhysioNet platform team
- Bridge2AI-Voice consortium
- Errata
- Updates
- Name
- version-history
- Description
- V1.0 (2024)
- Initial release.
- V1.1 (2025 01 17)
- Added MFCCs; files for v1.1 are no longer available on the platform.
- Newer Versions Available
- 2.0.0 (2025-04-16), 2.0.1 (2025-08-18).
- Retention Limit
- Version Access
- Name
- older-version-availability
- Description
- Older version 1.1 files are no longer available; latest version on the platform is 2.0.1.
- Extension Mechanism
- Is Tabular
- yes
- ID
- healthdatanexus-voice-1.0
- Name
- Health Data Nexus VOICE 1.0
- Title
- VOICE 1.0
- Description
- Resource page for the VOICE 1.0 project on Health Data Nexus.
- Page
- https://healthdatanexus.ai/content/b2ai-voice/1.0/
- Version
- 1.0
- DOI
- doi:10.57764/qb6h-em84
- Keywords
- VOICE
- Health Data Nexus
- b2ai-voice
- healthdatanexus.ai
- Created By
- Health Data Nexus
- Distribution Formats
- Distribution Dates
- License And Use Terms
- Name
- unspecified-license
- Description
- Resource/landing page for the VOICE 1.0 project; licensing for downloadable data not specified on this page.