Interleaved Semantic Evaluation

Project: VOICE · Method: claudecode_agent
YAML: data/d4d_concatenated/claudecode_agent/VOICE_d4d.yaml
R10 JSON: data/evaluation_llm/rubric10_semantic/concatenated/VOICE_claudecode_agent_evaluation.json
R20 JSON: data/evaluation_llm/rubric20_semantic/concatenated/VOICE_claudecode_agent_evaluation.json
Model: claude-sonnet-4-5-20250929
Rubric10 (semantic)
46/50 (92.0%)
Rubric20 (semantic)
79.0/84 (94.0%)
Consistency checks (R10/R20)
38 pass · 0 fail · 3 warn
Mapped feedback / fields
140 across 49 fields
R10 sub-element R20 question Semantic issue

Strengths

  • Outstanding structural completeness with all mandatory fields populated at exceptional depth (1,575-char description, 86 keywords)
  • Exemplary ethical documentation covering all human subjects protections: IRB approval, HIPAA Safe Harbor, Certificate of Confidentiality, 8-layer privacy protection, consent/revocation mechanisms, pediatric safeguards
  • Exceptional technical transparency with 8 preprocessing strategies, specific software tools (b2aiprep, OpenSMILE, Whisper), processing parameters (16kHz, 60 MFCCs), and open-source repository links
  • Comprehensive collection protocol documentation including hardware specs (iPad models, Avid AE-36 microphone), 22 adult acoustic tasks, 36 pediatric tasks, validated questionnaires per cohort
  • Excellent version tracking with 6 releases documented (v1.0-v3.0.0) including dates, participant counts, and feature additions; version-specific DOIs maintained
  • Strong FAIR compliance: PhysioNet DOI (10.13026/37yb-1t42), BIDS v1.9.0 conformance, 14 external resource links across platforms, standard formats (Parquet/TSV/JSON/WAV)
  • Detailed subpopulation characterization with clinical validation methods per cohort (laryngoscopy for voice disorders, MRI/CT for neuro, EHR/meds for psychiatric)
  • Clear two-tier access model semantically appropriate for risk level: registered access for deidentified features, controlled access via DACO for raw audio
  • Exceptional funding documentation with specific amounts ($4.66M), grant number, administering institute, and 50+ investigators with affiliations
  • Strong semantic consistency across human subjects research → IRB → consent → deidentification → privacy → vulnerable populations

Weaknesses

  • Limited publication documentation: single Interspeech 2024 citation; no dataset descriptor paper or additional methodological publications despite multi-year project scope
  • Provenance tracking lacks formal errata documentation or structured change log beyond version history
  • No RRID identifier present for software tools despite extensive use of research resources (OpenSMILE, Praat, b2aiprep)
  • Minor: grant number uses supplemental award notation (3...S3) which may cause confusion without explanation

Field-by-field

id
id: https://doi.org/10.13026/37yb-1t42
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information, description: 400+ chars comprehensive, keywords: 86 keywords, license_and_use_terms: detailed with access mechanisms
qualityAll mandatory fields present with exceptional content quality
correctnessAll fields semantically appropriate and correctly formatted
consistencyFields align with dataset scope and access model
name
name: Bridge2AI-Voice
no field-level feedback matched
title
title: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information
✓ 1/1 R10 1.Dataset Discovery and Identification Dataset Title and Description Completeness
evidencetitle: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information; description: 1,240 characters covering purpose, conditions, data types, collection methods, version details
qualityComprehensive description (>200 chars) with specific participant counts (833 adults, 300 pediatric), recording counts (~61,937), collection sites (North America, multi-institutional), data modalities (voice recordings, clinical info, derived features), and versioning (v3.0)
semanticcompleteness: excellent; specificity: high
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information, description: 400+ chars comprehensive, keywords: 86 keywords, license_and_use_terms: detailed with access mechanisms
qualityAll mandatory fields present with exceptional content quality
correctnessAll fields semantically appropriate and correctly formatted
consistencyFields align with dataset scope and access model
description
description: 'The Bridge2AI-Voice project seeks to create an ethically sourced flagship dataset to enable
  future research in artificial intelligence and support critical insights into the use of voice as a
  biomarker of health. The human voice contains complex acoustic markers which have been linked to important
  health conditions including dementia, mood disorders, and cancer. When viewed as a biomarker, voice
  is a promising characteristic to measure as it is simple to collect, cost-effective, and has broad clinical
  utility. This comprehensive collection provides voice recordings with corresponding clinical information
  from participants selected based on known conditions which manifest within the voice waveform including
  voice disorders, neurological disorders, mood disorders, and respiratory disorders. The dataset is designed
  to fuel voice AI research, establish data standards, and promote ethical and trustworthy AI/ML development
  for voice biomarkers of health. Data collection occurs through a multi-institutional collaborative effort
  using standardized protocols, custom smartphone applications, and rigorous ethical oversight. Version
  3.0 provides approximately 61,937 voice-derived recordings from 833 adult participants collected across
  multiple sites in North America, with derived features such as spectrograms, MFCCs, acoustic features,
  and clinical phenotype data. The pediatric dataset v1.0 is also available with data from 300 participants.
  Raw audio data is available through controlled access to protect participant privacy.

  '
✓ 1/1 R10 1.Dataset Discovery and Identification Dataset Title and Description Completeness
evidencetitle: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information; description: 1,240 characters covering purpose, conditions, data types, collection methods, version details
qualityComprehensive description (>200 chars) with specific participant counts (833 adults, 300 pediatric), recording counts (~61,937), collection sites (North America, multi-institutional), data modalities (voice recordings, clinical info, derived features), and versioning (v3.0)
semanticcompleteness: excellent; specificity: high
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information, description: 400+ chars comprehensive, keywords: 86 keywords, license_and_use_terms: detailed with access mechanisms
qualityAll mandatory fields present with exceptional content quality
correctnessAll fields semantically appropriate and correctly formatted
consistencyFields align with dataset scope and access model
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1,575 chars; purposes: 4 entries averaging 350+ chars each; addressing_gaps: 5 entries with detailed rationale
qualityExceptional narrative depth with comprehensive context and justification
correctnessNarratives accurately describe dataset scope, methods, and scientific context
consistencyPurpose statements align with tasks, gaps, and collection design
doi
doi: 10.13026/37yb-1t42
⚠ low R10 · consistency
issuePage field points to docs.b2ai-voice.org while PhysioNet is the primary distribution platform
fieldspage, doi
fixConsider pointing page to PhysioNet landing page matching DOI, or clarify that docs site is intentional
✓ 1/1 R10 1.Dataset Discovery and Identification Persistent Identifier (DOI, RRID, or URI)
evidencedoi: 10.13026/37yb-1t42
qualityValid DOI format with PhysioNet-specific prefix (10.13026) indicating proper registration with DataCite through PhysioNet
semanticformat_check: pass; prefix_plausibility: pass_physionet
✓ 1/1 R10 10.Cross-Platform and Community Integration Citation and DOI for Cross-referencing
evidencecitation: Bensoussan, Yael, et al. 'Developing Multi-Disorder Voice Protocols...' Proc. Interspeech 2024. https://doi.org/10.21437/Interspeech.2024-1926; doi: 10.13026/37yb-1t42 (PhysioNet DOI); version-specific DOIs maintained; Zenodo DOI for REDCap dictionary: 10.5281/zenodo.13834653
qualityComprehensive citation infrastructure: recommended citation (Bensoussan et al. Interspeech 2024 with DOI), dataset DOI (10.13026/37yb-1t42), version-specific DOIs for each release, Zenodo DOI for data dictionary (10.5281/zenodo.13834653), NIH RePORTER link for funding cross-reference
semanticcitation_completeness: excellent; doi_validity: valid_multiple_dois; cross_reference_enabled: yes
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information, description: 400+ chars comprehensive, keywords: 86 keywords, license_and_use_terms: detailed with access mechanisms
qualityAll mandatory fields present with exceptional content quality
correctnessAll fields semantically appropriate and correctly formatted
consistencyFields align with dataset scope and access model
4/5 R20 Q14 (Technical Documentation) Associated Publications
levelMultiple references and dataset citation
evidencecitation: Bensoussan et al. Interspeech 2024 (https://doi.org/10.21437/Interspeech.2024-1926); external_resources: Zenodo archive (10.5281/zenodo.13834653), NIH RePORTER project link; missing: additional methodological papers or dataset descriptor publications
qualityGood publication documentation with formal citation and DOIs; could benefit from additional methodological papers given project scope
correctnessDOIs valid format for Interspeech conference and Zenodo archive
consistencyCitation describes protocol development aligning with collection methods; publication date (2024) consistent with data collection timeline
1/1 R20 Q20 (FAIRness & Accessibility) Interlinking Across Platforms
levelPass - Cross-platform links verified
evidenceexternal_resources: 14 resources linking PhysioNet (primary distribution), Health Data Nexus (alternative platform with cloud compute), GitHub (bridge2ai-docs, b2aiprep, bridge2ai-redcap, senselab), Zenodo (archive), NIH RePORTER (grant info), doi.org (publications), Bridge2AI program site, training site (b2aivoicescholars.org)
qualityExceptional cross-platform interlinking connecting data repositories, code, documentation, publications, and training resources
correctnessPlatform links appropriate for roles: PhysioNet/Nexus for data, GitHub for code, Zenodo for archival, NIH RePORTER for funding
consistencyCross-platform references align with described distribution (PhysioNet primary), software (GitHub repos), and infrastructure (Health Data Nexus for compute)
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass - DOI and persistent URLs present
evidencedoi: 10.13026/37yb-1t42 (PhysioNet), page: https://docs.b2ai-voice.org, external_resources: 14 persistent URLs including PhysioNet, GitHub repos, Zenodo
qualityExcellent identifier coverage with PhysioNet DOI and multiple persistent access points
correctnessDOI prefix 10.13026 correctly identifies PhysioNet registrar
consistencyDOI matches page URL domain and distribution platform
page
page: https://docs.b2ai-voice.org
⚠ low R10 · consistency
issuePage field points to docs.b2ai-voice.org while PhysioNet is the primary distribution platform
fieldspage, doi
fixConsider pointing page to PhysioNet landing page matching DOI, or clarify that docs site is intentional
✓ 1/1 R10 1.Dataset Discovery and Identification Landing Page and Resources
evidencepage: https://docs.b2ai-voice.org; external_resources includes PhysioNet landing page, GitHub repos, documentation sites
qualityMultiple accessible landing pages documented: project docs (docs.b2ai-voice.org), PhysioNet distribution (physionet.org/content/b2ai-voice/), Health Data Nexus platform, plus 14 external resource entries
semanticaccessibility: multiple_platforms; url_validity: valid; issues: page field points to docs site rather than primary DOI landing page (PhysioNet) - minor inconsistency but both are valid
✓ 1/1 R10 10.Cross-Platform and Community Integration Outreach Materials and Documentation Links
evidenceexternal_resources: Project documentation site (docs.b2ai-voice.org), GitHub documentation repository, training program (B2AI Voice Scholars at b2aivoicescholars.org), software libraries (b2aiprep, SenseLab on GitHub), REDCap data dictionary (GitHub + Zenodo), PhysioNet landing page, publications (Interspeech 2024), NIH RePORTER project details
qualityComprehensive outreach and documentation: dedicated documentation site, training program (B2AI Voice Scholars), GitHub repositories (4 repos for software and documentation), published protocols (Interspeech 2024), data dictionary (public on GitHub and Zenodo), platform landing pages (PhysioNet, Health Data Nexus), funding transparency (NIH RePORTER)
semanticoutreach_completeness: excellent; educational_resources: yes; documentation_accessibility: public
1/1 R20 Q16 (FAIRness & Accessibility) Findability (Persistent Links)
levelPass - Multiple persistent URLs present
evidencepage: https://docs.b2ai-voice.org; external_resources: 14 URLs including https://physionet.org/content/b2ai-voice/, https://healthdatanexus.ai/content/b2ai-voice/1.0/, GitHub repos, DOIs (10.5281/zenodo.13834653, 10.21437/Interspeech.2024-1926), NIH RePORTER
qualityExcellent findability with multiple persistent access points across platforms
correctnessAll URLs follow standard patterns and include appropriate domains (.org, .ai, github.com, doi.org)
consistencyURLs align with described platforms (PhysioNet, Health Data Nexus, GitHub, Zenodo)
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass - DOI and persistent URLs present
evidencedoi: 10.13026/37yb-1t42 (PhysioNet), page: https://docs.b2ai-voice.org, external_resources: 14 persistent URLs including PhysioNet, GitHub repos, Zenodo
qualityExcellent identifier coverage with PhysioNet DOI and multiple persistent access points
correctnessDOI prefix 10.13026 correctly identifies PhysioNet registrar
consistencyDOI matches page URL domain and distribution platform
language
language: en
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Systematic Biases Identified and Described
evidencesampling_strategies.why_not_representative includes clinic-based recruitment bias (treatment-seeking populations), geographic bias (5 sites only), language bias (English-only), access bias (groups with less trust or proximity underrepresented); non-probability purposive sampling acknowledged
qualityMultiple systematic biases identified: selection bias (clinic-based recruitment favoring treatment-seeking), geographic bias (North America, 5 sites), linguistic bias (English-only in current releases), accessibility bias (underrepresentation of low-trust/low-proximity groups), purposive non-random sampling acknowledged
semanticbias_documentation: excellent; bias_types_identified: 4
version
version: 3.0.0
✓ 1/1 R10 1.Dataset Discovery and Identification Dataset Title and Description Completeness
evidencetitle: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information; description: 1,240 characters covering purpose, conditions, data types, collection methods, version details
qualityComprehensive description (>200 chars) with specific participant counts (833 adults, 300 pediatric), recording counts (~61,937), collection sites (North America, multi-institutional), data modalities (voice recordings, clinical info, derived features), and versioning (v3.0)
semanticcompleteness: excellent; specificity: high
✓ 1/1 R10 1.Dataset Discovery and Identification Hierarchical Structure (parent datasets, relationships)
evidenceMultiple subsets documented: Featurized Dataset (Registered Access), Raw Audio Dataset (Controlled Access), Pediatric Dataset (Registered Access) with separate versioning; version progression v1.0→v1.1→v2.0.0→v2.0.1→v3.0.0
qualityClear hierarchical structure with adult/pediatric datasets, access tiers (registered vs controlled), subset relationships, and version lineage documented throughout file
semanticstructure_clarity: excellent; relationships_defined: yes
✓ 1/1 R10 10.Cross-Platform and Community Integration Citation and DOI for Cross-referencing
evidencecitation: Bensoussan, Yael, et al. 'Developing Multi-Disorder Voice Protocols...' Proc. Interspeech 2024. https://doi.org/10.21437/Interspeech.2024-1926; doi: 10.13026/37yb-1t42 (PhysioNet DOI); version-specific DOIs maintained; Zenodo DOI for REDCap dictionary: 10.5281/zenodo.13834653
qualityComprehensive citation infrastructure: recommended citation (Bensoussan et al. Interspeech 2024 with DOI), dataset DOI (10.13026/37yb-1t42), version-specific DOIs for each release, Zenodo DOI for data dictionary (10.5281/zenodo.13834653), NIH RePORTER link for funding cross-reference
semanticcitation_completeness: excellent; doi_validity: valid_multiple_dois; cross_reference_enabled: yes
✓ 1/1 R10 10.Cross-Platform and Community Integration Related Datasets with Typed Relationships
evidencesubsets: 3 related subsets with defined relationships (Featurized Dataset as public registered access derived from raw, Raw Audio Dataset as controlled access source, Pediatric Dataset as separate cohort with distinct protocols); version relationships documented (v1.0→v1.1→v2.0.0→v2.0.1→v3.0.0); adult vs pediatric dataset separation with cross-references
qualityClear dataset relationships: 3 subsets with typed access relationships (featurized ← derives from → raw audio; pediatric as separate release), version lineage documented (5 versions with progression), subset access tier relationships (registered vs controlled), adult/pediatric split with cross-references, version-specific DOIs enabling precise referencing
semanticrelationship_clarity: excellent; relationship_types: derivation_access_versioning
✓ 1/1 R10 6.Data Provenance and Version Tracking Dataset Version Number Provided
evidenceversion: 3.0.0
qualityExplicit version number provided (3.0.0) using semantic versioning format
semanticversion_format: semantic_versioning
✓ 1/1 R10 6.Data Provenance and Version Tracking Version Access Methods Documented
evidenceversion_access: All dataset versions available through PhysioNet at https://physionet.org/content/b2ai-voice/ with version-specific DOIs; earlier versions on Health Data Nexus at https://healthdatanexus.ai/content/b2ai-voice/1.0/; older versions continue to be supported, hosted, and made available; each version has unique DOI
qualityComprehensive version access documentation: all versions available on PhysioNet with version-specific DOIs, alternative platform (Health Data Nexus) for earlier versions, commitment to maintaining older versions, unique DOI per version
semanticversion_access_clarity: excellent; historical_preservation: yes
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidencecollection_mechanisms: 3 detailed entries with hardware specs (iPad 9th/10th gen, iPad Air 5th gen, Avid AE-36 microphone, Apple dongle), custom app, reproschema-ui for pediatric; acquisition_methods: 4 entries detailing 22 adult tasks, 36 pediatric tasks, validated questionnaires, EHR linkage; data_collectors: research teams at 5 sites with compensation details; collection_timeframes: September 2022 - November 2026 with version-specific release dates; direct_collection: clinic-based with IRB consent
qualityExceptionally comprehensive collection protocol with specific hardware, software, tasks per cohort, site details, and timeline
correctnessCollection methods appropriate for clinical voice research; hardware choices (Avid AE-36 microphone) professional-grade
consistencyCollection mechanisms align with BIDS output, acoustic tasks match described features, timeframe matches funding period and version releases
5/5 R20 Q13 (Technical Documentation) Version History Documentation
levelComprehensive versioning with errata, updates, and release notes
evidenceversion: 3.0.0; updates: detailed with v1.0 (Jan 17, 2025, 306 participants, 12,523 recordings), v1.1 (MFCC features added), v2.0.0 (Apr 16, 2025), v2.0.1 (Aug 18, 2025), v3.0.0 (833 adults, 61,937 recordings), pediatric v1.0 (300 participants), semi-annual release plan, target 10,000 by 2027; version_access: all versions hosted with version-specific DOIs on PhysioNet and Health Data Nexus; distribution_dates: release timeline documented
qualityExceptional version tracking with specific dates, participant/recording counts per version, feature additions, and perpetual access policy
correctnessVersion progression logical (increasing participant counts, feature additions); dates align with project timeline
consistencyVersion information consistent across updates, distribution_dates, and version_access sections
4/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelStructured version control with timestamps
evidenceupdates: comprehensive version history with dates and participant counts per version (v1.0 Jan 17 2025, v1.1, v2.0.0 Apr 16 2025, v2.0.1 Aug 18 2025, v3.0.0); version_access: version-specific DOIs maintained; missing: explicit errata documentation or formal change log structure
qualityStrong provenance through versioned releases with dates and metrics; could benefit from formal errata/change log documentation
correctnessVersion dates follow logical temporal progression aligned with collection timeline
consistencyVersion history consistent across updates, distribution_dates, and version_access sections; participant counts increase monotonically
citation
citation: 'Bensoussan, Yael, et al. "Developing Multi-Disorder Voice Protocols: A team science approach
  involving clinical expertise, bioethics, standards, and DEI." Proc. Interspeech 2024. 2024. https://doi.org/10.21437/Interspeech.2024-1926

  '
⚠ low R20 · consistency
issueCitation provided but no additional external publications referenced despite multi-year project
fieldscitation, external_resources
fixConsider adding additional project publications if available
✓ 1/1 R10 10.Cross-Platform and Community Integration Citation and DOI for Cross-referencing
evidencecitation: Bensoussan, Yael, et al. 'Developing Multi-Disorder Voice Protocols...' Proc. Interspeech 2024. https://doi.org/10.21437/Interspeech.2024-1926; doi: 10.13026/37yb-1t42 (PhysioNet DOI); version-specific DOIs maintained; Zenodo DOI for REDCap dictionary: 10.5281/zenodo.13834653
qualityComprehensive citation infrastructure: recommended citation (Bensoussan et al. Interspeech 2024 with DOI), dataset DOI (10.13026/37yb-1t42), version-specific DOIs for each release, Zenodo DOI for data dictionary (10.5281/zenodo.13834653), NIH RePORTER link for funding cross-reference
semanticcitation_completeness: excellent; doi_validity: valid_multiple_dois; cross_reference_enabled: yes
4/5 R20 Q14 (Technical Documentation) Associated Publications
levelMultiple references and dataset citation
evidencecitation: Bensoussan et al. Interspeech 2024 (https://doi.org/10.21437/Interspeech.2024-1926); external_resources: Zenodo archive (10.5281/zenodo.13834653), NIH RePORTER project link; missing: additional methodological papers or dataset descriptor publications
qualityGood publication documentation with formal citation and DOIs; could benefit from additional methodological papers given project scope
correctnessDOIs valid format for Interspeech conference and Zenodo archive
consistencyCitation describes protocol development aligning with collection methods; publication date (2024) consistent with data collection timeline
keywords
keywords:
- voice biomarker
- acoustic biomarker
- Bridge2AI
- voice AI
- voice disorders
- neurological disorders
- neurodegenerative disorders
- mood disorders
- psychiatric disorders
- respiratory disorders
- pediatric voice disorders
- speech disorders
- Parkinson's disease
- Alzheimer's disease
- depression
- schizophrenia
- bipolar disorder
- stroke
- ALS
- autism spectrum disorder
- speech delay
- laryngeal cancer
- vocal fold paralysis
- muscle tension dysphonia
- laryngeal dystonia
- COPD
- chronic cough
- airway stenosis
- obstructive sleep apnea
- spectrogram
- MFCC
- mel-frequency cepstral coefficients
- OpenSMILE
- Praat
- Parselmouth
- torchaudio
- federated learning
- ethical AI
- multimodal health data
- electronic health records
- EHR
- radiomics
- genomics
- FAIR principles
- CARE principles
- PhysioNet
- Health Data Nexus
- BIDS
- Brain Imaging Data Structure
- b2aiprep
✓ 1/1 R10 1.Dataset Discovery and Identification Keywords or Tags for Searchability
evidencekeywords: 51 terms including voice biomarker, Bridge2AI, specific disorders (Parkinson's, Alzheimer's, depression), tools (OpenSMILE, Praat), standards (FAIR, CARE, BIDS), platforms (PhysioNet)
qualityExtensive keyword set (51 keywords) covering medical domains, technical methods, software tools, data standards, and platforms - far exceeds minimum threshold of 5
semanticcount: 51; coverage: comprehensive
✓ 1/1 R10 10.Cross-Platform and Community Integration Community Standards or Schema Conformance
evidenceBIDS v1.9.0 (Brain Imaging Data Structure) compliance; ICD-10 diagnostic coding; HIPAA Safe Harbor de-identification standards; FAIR principles and CARE principles mentioned in keywords; 45 CFR 46 (Common Rule) compliance; Fort Lauderdale Agreement and Open Science principles
qualityExtensive standards conformance: BIDS v1.9.0 (neuroimaging/biomedical data structure standard), ICD-10 (diagnostic coding), HIPAA Safe Harbor (de-identification), FAIR and CARE principles (data ethics), Common Rule (human subjects), Fort Lauderdale Agreement (data sharing ethics), Open Science principles
semanticstandard_count: 7; community_alignment: excellent
✓ 1/1 R10 3.Data Reuse and Interoperability Schema or Ontology Conformance Stated
evidenceBIDS v1.9.0 (Brain Imaging Data Structure) compliant; ICD-10 codes for diagnoses; HIPAA Safe Harbor standards; keywords mention FAIR principles, CARE principles
qualityMultiple schema/standard conformances: BIDS v1.9.0 (data structure), ICD-10 (diagnostic coding), HIPAA Safe Harbor (de-identification), FAIR and CARE principles (data sharing ethics)
semanticschema_conformance: multiple_standards; ontology_use: ICD-10
✓ 1/1 R10 5.Data Composition and Structure Data Topics or Conditions Represented
evidenceinstances: 5 predetermined disease cohort groups documented; subpopulations detail specific conditions (laryngeal cancer, Alzheimer's, Parkinson's, depression, schizophrenia, bipolar, COPD, chronic cough, airway stenosis, pediatric voice/speech disorders); keywords include 20+ specific conditions
qualityExtensive disease/condition coverage: 5 primary cohorts (voice, neuro, mood, respiratory, pediatric) with granular conditions documented (e.g., T1-T4 laryngeal cancer, idiopathic PD vs MSA vs PSP, bipolar I/II, autism spectrum disorder, speech delay)
semanticcondition_coverage: comprehensive; specificity: excellent
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information, description: 400+ chars comprehensive, keywords: 86 keywords, license_and_use_terms: detailed with access mechanisms
qualityAll mandatory fields present with exceptional content quality
correctnessAll fields semantically appropriate and correctly formatted
consistencyFields align with dataset scope and access model
5/5 R20 Q3 (Structural Completeness) Keyword Diversity
level≥8 keywords
evidencekeywords: 86 unique keywords covering clinical domains (voice/neurological/mood/respiratory disorders, specific diseases), technical terms (spectrogram, MFCC, OpenSMILE), standards (FAIR, CARE, BIDS), platforms (PhysioNet, Health Data Nexus)
qualityExceptionally comprehensive keyword coverage spanning clinical, technical, ethical, and infrastructure domains
correctnessAll keywords semantically appropriate for voice biomarker dataset
consistencyKeywords align with described clinical cohorts, software tools, and data standards
is_tabular
is_tabular: false
✓ 1/1 R10 5.Data Composition and Structure Variable-Level Metadata and Tabular Flag
evidenceis_tabular: false; variables documented via static_features.tsv with static_features.json data dictionary; phenotype files (demographics.tsv, confounders.tsv, diagnosis/*.tsv, questionnaire/*.tsv) each with JSON data dictionaries; REDCap data dictionary on GitHub; 22 adult acoustic tasks, 36 pediatric tasks described
qualityComprehensive variable metadata: tabular flag correctly set to false (multimodal audio+phenotype data), JSON data dictionaries for all TSV files, acoustic task descriptions (22 adult, 36 pediatric), feature extraction tools documented (OpenSMILE, Praat, Parselmouth)
semantictabular_flag_accuracy: correct_false_for_audio; metadata_availability: excellent
purposes
purposes:
- id: voice:purpose:1
  response: 'Integrate the use of voice as a biomarker of health in clinical care by generating a substantial
    multi-institutional, ethically sourced, and diverse voice database linked to multimodal health biomarkers
    to fuel voice AI research and build predictive models to assist in screening, diagnosis, and treatment
    of a broad range of diseases.

    '
- id: voice:purpose:2
  response: 'Create an ethically sourced flagship dataset of 10,000 voices linked to health information
    to enable future research in artificial intelligence and support critical insights into the use of
    voice as a biomarker of health, addressing the pressing need for large, high quality, multi-institutional
    and diverse voice databases linked to other health biomarkers.

    '
- id: voice:purpose:3
  response: 'Establish standards, best practices, and guidelines for voice data collection and analysis
    to advance the field of acoustic biomarkers by developing new standards that are AI/ML friendly and
    enable voice to emerge as a biomarker of health.

    '
- id: voice:purpose:4
  response: 'Address ethical, legal, and social challenges surrounding voice AI including risks of voice
    re-identification, vulnerabilities like voice AI hacking, concerns around voice data sharing and privacy,
    and the influence of gender and racial diversity on the development and application of voice AI technologies.

    '
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Motivation or Purpose for Dataset Creation
evidencepurposes: 4 detailed purposes covering voice biomarker integration in clinical care, ethically sourced flagship dataset creation, standards/best practices establishment, and addressing ethical/legal/social challenges of voice AI
qualityComprehensive motivation documentation: 4 distinct purposes addressing scientific needs (voice biomarker research), dataset creation goals (10,000 voices with multimodal health data), standards development (AI/ML-friendly guidelines), and ethical challenges (re-identification, privacy, diversity)
semanticpurpose_clarity: excellent; scientific_rationale: well_defined
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1,575 chars; purposes: 4 entries averaging 350+ chars each; addressing_gaps: 5 entries with detailed rationale
qualityExceptional narrative depth with comprehensive context and justification
correctnessNarratives accurately describe dataset scope, methods, and scientific context
consistencyPurpose statements align with tasks, gaps, and collection design
tasks
tasks:
- id: voice:task:1
  response: 'Enable development of AI/ML predictive models for screening, diagnosis, and treatment of
    voice disorders including laryngeal cancers, vocal fold paralysis, muscle tension dysphonia, laryngeal
    dystonia, benign laryngeal lesions, and pre-cancerous lesions, leveraging acoustic changes in phonation
    resulting from changes in vocal fold vibratory function.

    '
- id: voice:task:2
  response: 'Support machine learning models for neurological and neurodegenerative disorders including
    Alzheimer''s disease, Parkinson''s disease, mild cognitive impairment, other dementias, and ALS, detecting
    voice and speech changes such as slowed speech, low frequency, monotonous speech, vocal tremor, dysarthria,
    and aphasia. Validation uses CT, MRI, and whole genome sequencing data.

    '
- id: voice:task:3
  response: 'Develop AI algorithms for mood and psychiatric disorder detection including depression, schizophrenia,
    bipolar disorder, anxiety disorders, ADHD, PTSD, OCD, and borderline personality disorder, identifying
    vocal markers through clinical diagnosis, EHR records, and medication history.

    '
- id: voice:task:4
  response: 'Create machine learning models for respiratory disorder screening and therapeutic monitoring
    using respiratory sounds, cough sounds, and voice, applicable to conditions such as chronic cough,
    COPD, airway stenosis, and other respiratory conditions using spirometry, flow volume loops, and CT
    imaging for validation.

    '
- id: voice:task:5
  response: 'Build AI models for pediatric voice and speech disorder detection using 36 pediatric-specific
    acoustic tasks and specialized questionnaires, addressing the relative scarcity of pediatric voice
    data for participants aged 2-18 from the Hospital for Sick Children (SickKids).

    '
- id: voice:task:6
  response: 'Promote application of AI/ML for voice research through workforce development, curriculum
    creation on voice biomarkers of health for FAIR and CARE AI models, and fostering collaborations especially
    with researchers from underserved communities, building bridges between medical voice research, acoustic
    engineers, and the AI/ML community.

    '
- id: voice:task:7
  response: 'Enable AI model pretraining, fine-tuning, benchmarking, and validation using the standardized
    BIDS-compliant dataset structure with features extracted via b2aiprep, SenseLab, OpenSMILE, Parselmouth,
    Praat, and torchaudio toolkits.

    '
✓ 1/1 R10 5.Data Composition and Structure Variable-Level Metadata and Tabular Flag
evidenceis_tabular: false; variables documented via static_features.tsv with static_features.json data dictionary; phenotype files (demographics.tsv, confounders.tsv, diagnosis/*.tsv, questionnaire/*.tsv) each with JSON data dictionaries; REDCap data dictionary on GitHub; 22 adult acoustic tasks, 36 pediatric tasks described
qualityComprehensive variable metadata: tabular flag correctly set to false (multimodal audio+phenotype data), JSON data dictionaries for all TSV files, acoustic task descriptions (22 adult, 36 pediatric), feature extraction tools documented (OpenSMILE, Praat, Parselmouth)
semantictabular_flag_accuracy: correct_false_for_audio; metadata_availability: excellent
✓ 1/1 R10 6.Data Provenance and Version Tracking Provenance and Source Derivation Documented
evidenceMultiple provenance indicators: collection_mechanisms (Bridge2AI-Voice App on iPad with Avid AE-36 microphone), acquisition_methods (22 adult tasks, 36 pediatric tasks via reproschema-ui), data_collectors (research teams at 5 North American sites), preprocessing tools (b2aiprep, OpenSMILE, Parselmouth, Whisper), source cohort descriptions, IRB study ID (OT2OD032720)
qualityComprehensive provenance: data sources (5 collection sites), collection instruments (iPad hardware, microphone specs), software (Bridge2AI-Voice App, reproschema-ui), preprocessing pipeline (b2aiprep library), research teams (site clinicians, research assistants), IRB study identifier
semanticprovenance_completeness: excellent; source_traceability: yes
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Primary Research Objectives or Tasks
evidencetasks: 7 detailed tasks including AI/ML model development for voice disorders, neurological disorders, mood/psychiatric disorders, respiratory disorders, pediatric disorders, workforce development, and AI model pretraining/benchmarking; addressing_gaps: 5 specific research gaps
qualityExceptional task documentation: 7 specific research tasks (voice disorders, neuro, mood, respiratory, pediatric, workforce development, model pretraining) with disease-specific details and validation methods; 5 research gaps addressed (data scarcity, pediatric limitations, standards gaps, infrastructure needs)
semantictask_specificity: excellent; research_gaps_identified: yes
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Data Acquisition Methods Listed
evidenceacquisition_methods: 22 adult acoustic tasks (Non-Voice: Respiration, Cough; Voice/Non-Speech: Prolonged Vowel, MPT, Glides, DDK; Speech: Rainbow Passage, Free Speech, Stroop, etc.); 36 pediatric tasks; validated questionnaires (VHI-10, PHQ-9, GAD-7, MOCA, etc.); EHR access for multimodal validation
qualityComprehensive acquisition documentation: 22 adult acoustic tasks with specific names (Rainbow Passage, Caterpillar Passage, DDK /pa/ta/ka/), 36 pediatric tasks, disease-specific validated questionnaires (15+ named instruments), EHR linkage for validation (imaging, genomics, respiratory function)
semanticacquisition_detail: excellent; task_specificity: high
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidencecollection_mechanisms: 3 detailed entries with hardware specs (iPad 9th/10th gen, iPad Air 5th gen, Avid AE-36 microphone, Apple dongle), custom app, reproschema-ui for pediatric; acquisition_methods: 4 entries detailing 22 adult tasks, 36 pediatric tasks, validated questionnaires, EHR linkage; data_collectors: research teams at 5 sites with compensation details; collection_timeframes: September 2022 - November 2026 with version-specific release dates; direct_collection: clinic-based with IRB consent
qualityExceptionally comprehensive collection protocol with specific hardware, software, tasks per cohort, site details, and timeline
correctnessCollection methods appropriate for clinical voice research; hardware choices (Avid AE-36 microphone) professional-grade
consistencyCollection mechanisms align with BIDS output, acoustic tasks match described features, timeframe matches funding period and version releases
addressing_gaps
addressing_gaps:
- id: voice:gap:1
  response: 'Address the lack of large, high quality, multi-institutional and diverse voice databases
    linked to multimodal health biomarkers (demographics, imaging, genomics, risk factors) necessary to
    fuel voice AI research and answer tangible clinical questions. Previous studies had sample sizes too
    small or lacked metadata needed for robust, clinically useful models.

    '
- id: voice:gap:2
  response: 'Overcome limitations in existing voice and psychiatric disorder research that has relied
    on small datasets with limited demographic diversity reporting, lack of standardized data collection
    protocols precluding meta-analysis, and possible confounders limiting external validity and clinical
    usability.

    '
- id: voice:gap:3
  response: 'Fill the gap in pediatric voice and speech analysis research, which is sparser partly due
    to ethical concerns and challenges in data acquisition for this cohort, particularly for autism spectrum
    disorder and speech delay detection.

    '
- id: voice:gap:4
  response: 'Establish missing standards for voice data collection, acoustic analysis, and ethical frameworks
    for consenting to voice data collection, sharing, and utilization in the context of voice AI technology
    development and clinical adoption.

    '
- id: voice:gap:5
  response: 'Develop software and cloud infrastructure for automated voice data collection through a smartphone
    application (Bridge2AI-Voice App) that allows non-invasive, user-friendly, high quality voice data
    collection while minimizing human manipulation and implementing federated learning technology to minimize
    data sharing while preserving patient privacy.

    '
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Primary Research Objectives or Tasks
evidencetasks: 7 detailed tasks including AI/ML model development for voice disorders, neurological disorders, mood/psychiatric disorders, respiratory disorders, pediatric disorders, workforce development, and AI model pretraining/benchmarking; addressing_gaps: 5 specific research gaps
qualityExceptional task documentation: 7 specific research tasks (voice disorders, neuro, mood, respiratory, pediatric, workforce development, model pretraining) with disease-specific details and validation methods; 5 research gaps addressed (data scarcity, pediatric limitations, standards gaps, infrastructure needs)
semantictask_specificity: excellent; research_gaps_identified: yes
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1,575 chars; purposes: 4 entries averaging 350+ chars each; addressing_gaps: 5 entries with detailed rationale
qualityExceptional narrative depth with comprehensive context and justification
correctnessNarratives accurately describe dataset scope, methods, and scientific context
consistencyPurpose statements align with tasks, gaps, and collection design
creators
creators:
- id: voice:creator:1
  description: 'Bridge2AI-Voice Consortium led by Dr. Yael Bensoussan (Contact PI, University of South
    Florida, Department of Otolaryngology) and Dr. Olivier Elemento (Co-PI, Weill Cornell Medicine). The
    multidisciplinary consortium includes over 50 investigators from 12+ institutions across North America
    spanning clinical medicine, biomedical research, machine learning, data science, social science, and
    ethics. Key co-investigators include: Alexandros Sigaras (Weill Cornell Medicine), Anais Rameau (Weill
    Cornell Medicine), Maria Powell (Vanderbilt University Medical Center), Ruth Bahr (USF), Jennifer
    Siu (Hospital for Sick Children), Philip Payne (Washington University in St. Louis), David Dorr (Oregon
    Health & Science University), Jean-Christophe Belisle-Pipon (Simon Fraser University), Vardit Ravitsky
    (The Hastings Center), Satrajit Ghosh (MIT), Frank Rudzicz (University of Toronto), Jordan Lerner-Ellis
    (Sinai Health), Don Bolser (University of Florida), Alistair Johnson (MIT/PhysioNet), and Jennifer
    Siu (Hospital for Sick Children).

    '
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Creators and Acknowledgements Documented
evidencecreators: Bridge2AI-Voice Consortium with 50+ investigators from 12+ institutions; Contact PI: Dr. Yael Bensoussan (USF); Co-PI: Dr. Olivier Elemento (Weill Cornell); 14 key co-investigators named with institutional affiliations; participant_compensation documented; research staff authorship acknowledgment
qualityComprehensive creator documentation: consortium structure (50+ investigators, 12+ institutions), leadership (Contact PI, Co-PI), 14 key co-investigators with names and affiliations, participant compensation ($40-120), research staff acknowledgment (authorship)
semanticcreator_completeness: excellent; institutional_attribution: yes
5/5 R20 Q7 (Metadata Quality & Content) Funding and Acknowledgements Completeness
levelFunders with grants + creators with affiliations
evidencefunders: 2 entries with NIH grant 3OT2OD032720-01S3 ($4,660,942 total funding with direct/indirect breakdown), NIBIB grant R01EB030362; creators: Bridge2AI-Voice Consortium with 50+ investigators from 12+ institutions, PIs and co-investigators listed with affiliations
qualityExceptional funding documentation with specific amounts, grant numbers, administering institutes, and comprehensive creator attribution
correctnessGrant number 3OT2OD032720-01S3 follows NIH supplement notation format (valid)
consistencyFunding aligns with project scope, timeline (2022-2026), and multi-institutional design
funders
funders:
- id: voice:funder:1
  description: 'National Institutes of Health (NIH) Common Fund Bridge2AI Program. Grant number: 3OT2OD032720-01S3.
    Opportunity Number: OTA-21-008. Project dates: September 1, 2022 to November 30, 2026. Total funding
    in 2025: $4,660,942 (Direct Costs: $4,072,321, Indirect Costs: $588,621). Administering Institute:
    NIH Office of the Director. Study Section: Data Coordination, Mapping, and Modeling (DCMM).

    '
- id: voice:funder:2
  description: 'National Institute of Biomedical Imaging and Bioengineering (NIBIB). Supports PhysioNet
    managed by MIT Laboratory for Computational Physiology under NIH grant number R01EB030362, which serves
    as the primary distribution platform for the Bridge2AI-Voice dataset.

    '
⚠ low R10 · correctness
issueGrant number format '3OT2OD032720-01S3' includes supplement suffix which is non-standard but valid NIH format
fieldsfunders
fixFormat is valid NIH supplemental grant notation - no action required
⚠ low R20 · correctness
issueGrant number format 3OT2OD032720-01S3 uses unusual prefix '3' instead of standard NIH OT2
fieldsfunders
fixVerify grant number format - appears to be supplemental award with '3' prefix and 'S3' suffix which is valid NIH notation
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Funding Sources and Mechanisms Listed
evidencefunders: NIH Common Fund Bridge2AI Program, NIBIB (supporting PhysioNet via R01EB030362); NIH Office of the Director administering; Study Section: Data Coordination, Mapping, and Modeling (DCMM); Total funding 2025: $4,660,942
qualityDetailed funding documentation: NIH Common Fund Bridge2AI Program (primary), NIBIB (PhysioNet support), administering institute (NIH Office of the Director), study section (DCMM), funding amounts ($4,660,942 total in 2025, with direct/indirect cost breakdown)
semanticfunding_completeness: excellent; agency_specificity: yes
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Grant IDs or Award Numbers Present
evidencefunders: Grant number: 3OT2OD032720-01S3 (Bridge2AI-Voice primary grant); R01EB030362 (NIBIB supporting PhysioNet); Opportunity Number: OTA-21-008; Study ID: OT2OD032720
qualityMultiple grant identifiers: 3OT2OD032720-01S3 (primary grant with supplement notation), R01EB030362 (PhysioNet infrastructure), OTA-21-008 (opportunity number), OT2OD032720 (study ID); NIH grant formats validated
semanticgrant_number_format: valid_nih_supplement; grant_number_count: 2; pattern_match: NIH_format_OT2_and_R01; issues: Supplement suffix -01S3 is valid NIH notation but less common in citations
5/5 R20 Q7 (Metadata Quality & Content) Funding and Acknowledgements Completeness
levelFunders with grants + creators with affiliations
evidencefunders: 2 entries with NIH grant 3OT2OD032720-01S3 ($4,660,942 total funding with direct/indirect breakdown), NIBIB grant R01EB030362; creators: Bridge2AI-Voice Consortium with 50+ investigators from 12+ institutions, PIs and co-investigators listed with affiliations
qualityExceptional funding documentation with specific amounts, grant numbers, administering institutes, and comprehensive creator attribution
correctnessGrant number 3OT2OD032720-01S3 follows NIH supplement notation format (valid)
consistencyFunding aligns with project scope, timeline (2022-2026), and multi-institutional design
instances
instances:
- id: voice:instance:1
  description: 'Adult participants presenting at specialty clinics (high volume expert clinics) across
    multiple sites in North America. Participants selected based on membership to five predetermined disease
    cohort groups: Voice Disorders, Neurological and Neurodegenerative Disorders, Mood and Psychiatric
    Disorders, Respiratory Disorders, and Pediatric Voice and Speech Disorders. Version 3.0 contains approximately
    833 adult participants with ~61,937 voice-derived recordings. Pediatric dataset v1.0 adds 300 participants.
    Enrollment anticipated to reach 10,000 participants by 2027.

    '
  instance_type: Human participants with clinical diagnoses recruited from specialty clinics
  counts: 833
  label: true
  label_description: 'Diagnostic labels assigned by clinical assessment at each site. Clinicians provided
    diagnoses based on clinical interview and appropriate work-up including laryngoscopy, stroboscopy,
    MRI, CT, whole genome sequencing, EHR records, and medication prescriptions. Labels include diagnostic
    categories: vocal pathologies, neurological disorders, psychiatric conditions, respiratory disorders,
    and pediatric voice/speech disorders. Per Bridge2AI protocols and ICD-10 codes. Single labeler per
    participant (site clinician).

    '
✓ 1/1 R10 5.Data Composition and Structure Number of Instances or Samples Reported
evidenceinstances.counts: 833 adult participants; ~61,937 voice-derived recordings in v3.0; pediatric v1.0 adds 300 participants; v1.0 had 306 participants with 12,523 recordings; target 10,000 participants by 2027
qualitySpecific counts throughout: 833 adults (v3.0), ~61,937 recordings, 300 pediatric participants, version-specific counts (v1.0: 306 participants, 12,523 recordings), enrollment target (10,000 by 2027)
semanticcount_specificity: excellent; version_progression_logical: yes
✓ 1/1 R10 5.Data Composition and Structure Data Topics or Conditions Represented
evidenceinstances: 5 predetermined disease cohort groups documented; subpopulations detail specific conditions (laryngeal cancer, Alzheimer's, Parkinson's, depression, schizophrenia, bipolar, COPD, chronic cough, airway stenosis, pediatric voice/speech disorders); keywords include 20+ specific conditions
qualityExtensive disease/condition coverage: 5 primary cohorts (voice, neuro, mood, respiratory, pediatric) with granular conditions documented (e.g., T1-T4 laryngeal cancer, idiopathic PD vs MSA vs PSP, bipolar I/II, autism spectrum disorder, speech delay)
semanticcondition_coverage: comprehensive; specificity: excellent
5/5 R20 Q15 (Technical Documentation) Human Subject Representation
levelDetailed demographics and inclusion/exclusion criteria
evidenceinstances: 833 adult participants from specialty clinics, 300 pediatric; subpopulations: 6 detailed cohorts (Voice Disorders, Neuro/Neurodegenerative, Mood/Psychiatric, Respiratory, Pediatric, Control) with clinical validation methods, age ranges, diagnostic criteria; sampling_strategies: inclusion/exclusion per protocol Table 1, non-representative acknowledged with specific reasons (geographic limitation, clinic-based, English-only); demographics in phenotype files mentioned
qualityExceptional human subject characterization with detailed subpopulations, clinical validation per cohort, and explicit discussion of representativeness limitations
correctnessSubpopulation descriptions clinically accurate with appropriate validation methods per disease category
consistencyCohort descriptions align with tasks (22 adult, 36 pediatric), purposes, and addressing_gaps sections
1/1 R20 Q5 (Structural Completeness) Data File Size Availability
levelPass - Instance counts provided
evidenceinstances: 833 adult participants, 300 pediatric participants; ~61,937 recordings in v3.0; v1.0: 306 participants, 12,523 recordings
qualityDetailed instance counts with version-specific breakdowns
correctnessInstance counts internally consistent across versions (increasing progression)
consistencyCounts align with version history and collection timeframe descriptions
subsets
subsets:
- id: voice:subset:1
  name: Featurized Dataset (Registered Access)
  description: 'Contains AI-ready derived features from voice recordings distributed through PhysioNet
    under registered access. Features include: OpenSMILE eGeMaps features per audio file, Parselmouth/Praat
    speech features, speech intelligibility metrics, torchaudio-based pitch contour, spectrograms, mel
    spectrograms, MFCCs, SPARC-based features including EMA estimates, phonetic posteriorgrams (PPGs).
    Phenotype data includes demographics, acoustic confounders, validated questionnaires, and clinical
    diagnoses in BIDS-compliant format. HIPAA Safe Harbor identifiers removed. Free speech transcripts
    and open-response audio features removed to protect privacy. Spectrograms, MFCCs, mel spectrograms,
    transcriptions, EMAs, and PPGs from open-response prompts removed.

    '
- id: voice:subset:2
  name: Raw Audio Dataset (Controlled Access)
  description: 'Original raw audio waveforms available through controlled access only to protect participant
    privacy. Users must email DACO@b2ai-voice.org and complete a Data Access Request Form (DARF), Data
    Use Agreement (DUA), and institutional Data Use and Transfer Agreement (DTUA) signed by authorized
    official. Raw audio has been used in Bridge2AI Summer School and hackathon. Data stored in BIDS v1.9.0
    compliant format with WAV audio files and JSON metadata per participant session and acoustic task.

    '
- id: voice:subset:3
  name: Pediatric Dataset (Registered Access)
  description: 'Pediatric dataset v1.0 containing data from 300 pediatric participants. Collected exclusively
    at Hospital for Sick Children (SickKids) using reproschema-ui with the Bridge2AI-Voice pediatric protocol.
    Participants grouped by age (2-4, 4-6, 6-10, 10+ years). Contains 36 pediatric-specific acoustic tasks
    and pediatric validated questionnaires (C-VHI-10, PVOS, PVRQOL, PHQ-A). Available through PhysioNet
    under separate registered access link.

    '
✓ 1/1 R10 1.Dataset Discovery and Identification Hierarchical Structure (parent datasets, relationships)
evidenceMultiple subsets documented: Featurized Dataset (Registered Access), Raw Audio Dataset (Controlled Access), Pediatric Dataset (Registered Access) with separate versioning; version progression v1.0→v1.1→v2.0.0→v2.0.1→v3.0.0
qualityClear hierarchical structure with adult/pediatric datasets, access tiers (registered vs controlled), subset relationships, and version lineage documented throughout file
semanticstructure_clarity: excellent; relationships_defined: yes
✓ 1/1 R10 10.Cross-Platform and Community Integration Related Datasets with Typed Relationships
evidencesubsets: 3 related subsets with defined relationships (Featurized Dataset as public registered access derived from raw, Raw Audio Dataset as controlled access source, Pediatric Dataset as separate cohort with distinct protocols); version relationships documented (v1.0→v1.1→v2.0.0→v2.0.1→v3.0.0); adult vs pediatric dataset separation with cross-references
qualityClear dataset relationships: 3 subsets with typed access relationships (featurized ← derives from → raw audio; pediatric as separate release), version lineage documented (5 versions with progression), subset access tier relationships (registered vs controlled), adult/pediatric split with cross-references, version-specific DOIs enabling precise referencing
semanticrelationship_clarity: excellent; relationship_types: derivation_access_versioning
sampling_strategies
sampling_strategies:
- id: voice:sampling:1
  description: 'Non-probability purposive sampling from specialty clinics (high volume expert clinics
    - outpatient clinics seeing >50 patients per month from same disease category). Patients presenting
    at clinics screened for eligibility per inclusion/exclusion criteria outlined in protocol Table 1.
    Not representative of general population due to limited geographic locations and clinic-based recruitment.
    Current v3.0 dataset is a sample of an ongoing collection targeting 10,000 participants by 2027. English-speaking
    adult participants (18-120 years); Spanish protocols under development for future releases.

    '
  is_sample: true
  is_random: false
  is_representative: false
  why_not_representative:
  - Data collected at limited number of geographic locations (five sites)
  - Clinic-based recruitment introduces selection bias toward treatment-seeking populations
  - Remote data collection not included in initial releases
  - Groups with less trust in medical system or less proximal to collection sites underrepresented
  - Current releases contain only English-speaking participants
  - Public releases do not contain equal distribution across disease categories
  strategies:
  - Targeted recruitment from high volume expert specialty clinics
  - Inclusion/exclusion criteria per Bridge2AI-Voice protocol Table 1 for each disease cohort
  - Multi-institutional enrollment across North American sites
  - IRB-approved prospective consent process
  - Pediatric participants exclusively from Hospital for Sick Children (SickKids)
✓ 1/1 R10 5.Data Composition and Structure Data Quality Issues and Anomalies Documented
evidencesampling_strategies: Non-probability purposive sampling, not representative of general population, selection bias acknowledged; cleaning_strategies: audit protocol with missingness tables, distribution/outlier checks, categorical response validation, audio quality control (silence, duration, speech-to-passage accuracy)
qualityData quality documentation includes: sampling bias acknowledgment (clinic-based, limited geographic sites), representativeness limitations (English-only, treatment-seeking populations), audit protocol (missingness tables, outlier checks, quality metrics), audio QC metrics
semanticquality_awareness: excellent; bias_acknowledgment: yes
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Known Limitations Documented
evidencesampling_strategies.why_not_representative: 5 specific limitations listed (limited geographic locations - 5 sites, clinic-based selection bias, no remote collection in initial releases, underrepresentation of groups with less medical system trust or proximal access, English-only in current releases, unequal disease category distribution); is_sample: true, is_random: false, is_representative: false
qualityExplicit limitations documentation: 5 specific representativeness limitations identified (geographic constraints, clinic-based bias, language restrictions, trust/access barriers, disease distribution imbalance), sampling flags correctly set (sample=true, random=false, representative=false), acknowledgment of treatment-seeking population bias
semanticlimitation_specificity: excellent; bias_awareness: yes
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Systematic Biases Identified and Described
evidencesampling_strategies.why_not_representative includes clinic-based recruitment bias (treatment-seeking populations), geographic bias (5 sites only), language bias (English-only), access bias (groups with less trust or proximity underrepresented); non-probability purposive sampling acknowledged
qualityMultiple systematic biases identified: selection bias (clinic-based recruitment favoring treatment-seeking), geographic bias (North America, 5 sites), linguistic bias (English-only in current releases), accessibility bias (underrepresentation of low-trust/low-proximity groups), purposive non-random sampling acknowledged
semanticbias_documentation: excellent; bias_types_identified: 4
5/5 R20 Q15 (Technical Documentation) Human Subject Representation
levelDetailed demographics and inclusion/exclusion criteria
evidenceinstances: 833 adult participants from specialty clinics, 300 pediatric; subpopulations: 6 detailed cohorts (Voice Disorders, Neuro/Neurodegenerative, Mood/Psychiatric, Respiratory, Pediatric, Control) with clinical validation methods, age ranges, diagnostic criteria; sampling_strategies: inclusion/exclusion per protocol Table 1, non-representative acknowledged with specific reasons (geographic limitation, clinic-based, English-only); demographics in phenotype files mentioned
qualityExceptional human subject characterization with detailed subpopulations, clinical validation per cohort, and explicit discussion of representativeness limitations
correctnessSubpopulation descriptions clinically accurate with appropriate validation methods per disease category
consistencyCohort descriptions align with tasks (22 adult, 36 pediatric), purposes, and addressing_gaps sections
subpopulations
subpopulations:
- id: voice:subpop:1
  name: Voice Disorders cohort
  description: 'Participants with laryngeal disorders including laryngeal cancer (T1-T4, biopsy proven),
    laryngitis (acute, chronic, bacterial, fungal, autoimmune), pre-cancerous lesions (keratosis, leukoplakia),
    benign vocal cord lesions (nodules, polyps, cysts, Reinke''s edema, recurrent laryngeal papilloma),
    muscle tension dysphonia, spasmodic dysphonia and laryngeal tremor, unilateral vocal fold paralysis,
    and glottic insufficiency/ presbyphonia. Validated by laryngoscopy images and stroboscopy videos.

    '
- id: voice:subpop:2
  name: Neurological and Neurodegenerative Disorders cohort
  description: 'Participants aged 44-85 with clinical diagnoses including mild cognitive impairment, Alzheimer''s
    disease, other dementias (frontotemporal, Lewy body, vascular, mixed, alcohol-induced), ALS (sporadic,
    familial, spinal/limb-onset, bulbar-onset), and Parkinson''s disease (idiopathic PD, multiple system
    atrophy, progressive supranuclear palsy, corticobasal degeneration). Must be able to read and speak
    English. Validated by CT Brain, MRI Brain, serum biomarkers (Ptau proteins for AD), and whole genome
    sequencing. Participants enrolled in deep brain stimulation studies noted.

    '
- id: voice:subpop:3
  name: Mood and Psychiatric Disorders cohort
  description: 'Participants with clinical diagnoses including depression/major depressive disorder, bipolar
    I/II disorder, anxiety disorder, schizophrenia, ADHD (comorbid), PTSD (comorbid), OCD (comorbid),
    panic disorder (comorbid), borderline personality disorder (comorbid), eating disorder (comorbid),
    insomnia/sleep disorder (comorbid), social anxiety disorder (comorbid), autism spectrum disorder (comorbid),
    alcohol/substance use disorder (comorbid), and other psychiatric disorders. Validated by electronic
    health summary, clinical diagnosis from psychiatrist, and medication history. Uses GAD-7, PHQ-9, PTSD,
    ADHD, DSM-5 adult questionnaires among others.

    '
- id: voice:subpop:4
  name: Respiratory Disorders cohort
  description: 'Participants with respiratory conditions including airway stenosis (bilateral vocal fold
    paralysis, supraglottic, glottic, posterior glottic, subglottic, tracheal stenosis, multi-level upper
    airway stenosis) and chronic cough (bothersome cough >8 weeks). Validated by spirometry, flow volume
    loops, and CT scan of neck/chest. Dyspnea Index and Leicester Cough Questionnaire administered.

    '
- id: voice:subpop:5
  name: Pediatric cohort
  description: 'Pediatric participants aged 2-18 with voice and speech disorders recruited exclusively
    from Hospital for Sick Children (SickKids). Grouped by age: 2-4, 4-6, 6-10, 10+ years. Data collected
    using reproschema-ui with Bridge2AI-Voice pediatric protocol. Includes 36 pediatric-specific acoustic
    tasks and specialized questionnaires (C-VHI-10, PVOS, PVRQOL, PHQ-A). Pediatric dataset v1.0 available
    separately from adult dataset.

    '
- id: voice:subpop:6
  name: Control participants
  description: 'Healthy volunteer control participants who complete common questionnaires and voice tasks
    (Voice Handicap Index-10, PHQ-9, GAD-7, Winograd) to provide normative comparisons.

    '
✓ 1/1 R10 5.Data Composition and Structure Cohort or Subpopulations Characteristics Described
evidencesubpopulations: 6 detailed cohorts (Voice Disorders, Neurological/Neurodegenerative, Mood/Psychiatric, Respiratory, Pediatric, Control) with specific diagnoses, age ranges, validation methods, and questionnaires for each
qualityExceptional subpopulation documentation: 6 cohorts with disease-specific details, inclusion criteria, age ranges (e.g., 44-85 for neuro, 2-18 for peds), validation methods (imaging, genomics, clinical diagnosis), and specialized questionnaires per cohort
semanticsubpopulation_detail: exceptional; cohort_count: 6
✓ 1/1 R10 5.Data Composition and Structure Data Topics or Conditions Represented
evidenceinstances: 5 predetermined disease cohort groups documented; subpopulations detail specific conditions (laryngeal cancer, Alzheimer's, Parkinson's, depression, schizophrenia, bipolar, COPD, chronic cough, airway stenosis, pediatric voice/speech disorders); keywords include 20+ specific conditions
qualityExtensive disease/condition coverage: 5 primary cohorts (voice, neuro, mood, respiratory, pediatric) with granular conditions documented (e.g., T1-T4 laryngeal cancer, idiopathic PD vs MSA vs PSP, bipolar I/II, autism spectrum disorder, speech delay)
semanticcondition_coverage: comprehensive; specificity: excellent
5/5 R20 Q15 (Technical Documentation) Human Subject Representation
levelDetailed demographics and inclusion/exclusion criteria
evidenceinstances: 833 adult participants from specialty clinics, 300 pediatric; subpopulations: 6 detailed cohorts (Voice Disorders, Neuro/Neurodegenerative, Mood/Psychiatric, Respiratory, Pediatric, Control) with clinical validation methods, age ranges, diagnostic criteria; sampling_strategies: inclusion/exclusion per protocol Table 1, non-representative acknowledged with specific reasons (geographic limitation, clinic-based, English-only); demographics in phenotype files mentioned
qualityExceptional human subject characterization with detailed subpopulations, clinical validation per cohort, and explicit discussion of representativeness limitations
correctnessSubpopulation descriptions clinically accurate with appropriate validation methods per disease category
consistencyCohort descriptions align with tasks (22 adult, 36 pediatric), purposes, and addressing_gaps sections
collection_mechanisms
collection_mechanisms:
- id: voice:collection:1
  description: 'Voice data collected in clinic using custom Bridge2AI-Voice App on iPad (9th or 10th generation)
    or iPad Air (5th generation) with Avid AE-36 microphone and Apple dongle connector. App collects breathing
    sounds and voice, speech, and linguistic tasks along with health information through surveys and validated
    questionnaires. Research assistant present during collection. Future remote data collection planned
    but not included in current releases.

    '
- id: voice:collection:2
  description: 'Pediatric data collected using reproschema-ui with Bridge2AI-Voice pediatric protocol.
    Same iPad hardware as adult protocol with age-appropriate tasks grouped by age range (2-4, 4-6, 6-10,
    10+ years). Questions read to participants when needed.

    '
- id: voice:collection:3
  description: 'Clinical data including EHR information, imaging (laryngoscopy, stroboscopy, MRI, CT),
    and genomic data extracted from sites independently and uploaded through REDCap database. No external
    multimodal data released in current dataset versions.

    '
✓ 1/1 R10 6.Data Provenance and Version Tracking Provenance and Source Derivation Documented
evidenceMultiple provenance indicators: collection_mechanisms (Bridge2AI-Voice App on iPad with Avid AE-36 microphone), acquisition_methods (22 adult tasks, 36 pediatric tasks via reproschema-ui), data_collectors (research teams at 5 North American sites), preprocessing tools (b2aiprep, OpenSMILE, Parselmouth, Whisper), source cohort descriptions, IRB study ID (OT2OD032720)
qualityComprehensive provenance: data sources (5 collection sites), collection instruments (iPad hardware, microphone specs), software (Bridge2AI-Voice App, reproschema-ui), preprocessing pipeline (b2aiprep library), research teams (site clinicians, research assistants), IRB study identifier
semanticprovenance_completeness: excellent; source_traceability: yes
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Collection Mechanisms and Settings Described
evidencecollection_mechanisms: Voice data collected in clinic using custom Bridge2AI-Voice App on iPad (9th/10th gen or iPad Air 5th gen) with Avid AE-36 microphone and Apple dongle connector; research assistant present; collection_timeframes: September 2022 - November 2026; semi-annual releases planned
qualityDetailed collection documentation: specific hardware (iPad models, Avid AE-36 microphone, Apple dongle), software (Bridge2AI-Voice App, reproschema-ui for pediatric), setting (in-clinic with research assistant), timeframe (Sep 2022 - Nov 2026), sites (5 North American specialty clinics)
semanticmechanism_specificity: excellent; replicability: high
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidencecollection_mechanisms: 3 detailed entries with hardware specs (iPad 9th/10th gen, iPad Air 5th gen, Avid AE-36 microphone, Apple dongle), custom app, reproschema-ui for pediatric; acquisition_methods: 4 entries detailing 22 adult tasks, 36 pediatric tasks, validated questionnaires, EHR linkage; data_collectors: research teams at 5 sites with compensation details; collection_timeframes: September 2022 - November 2026 with version-specific release dates; direct_collection: clinic-based with IRB consent
qualityExceptionally comprehensive collection protocol with specific hardware, software, tasks per cohort, site details, and timeline
correctnessCollection methods appropriate for clinical voice research; hardware choices (Avid AE-36 microphone) professional-grade
consistencyCollection mechanisms align with BIDS output, acoustic tasks match described features, timeframe matches funding period and version releases
acquisition_methods
acquisition_methods:
- id: voice:acquisition:1
  description: '22 acoustic tasks recorded through Bridge2AI-Voice App for adult participants including:
    Non-Voice (Respiration, Cough, Breath Sounds, Voluntary Cough); Voice/Non-Speech (Prolonged Vowel
    /e/, Maximum Phonation Time, Glides, Loudness /Hey/, Diadochokinesis /pa/ta/ka/); Speech (Rainbow
    Passage, Caterpillar Passage, Cape-V Sentences, Free Speech, Picture Description, Story Recall, Animal
    Fluency, Open Response Questions, Word-Color Stroop, Productive Vocabulary, Random Item Generation,
    Cinderella Story). Tasks vary by disease cohort (Part A/Voice/Resp/Mood/Neuro).

    '
- id: voice:acquisition:2
  description: '36 pediatric acoustic tasks collected via reproschema-ui including: Speech tasks (ABC''s,
    Ready for School, Favorite Show, Favorite Food, Outside of School, Months, Counting, Naming Animals,
    Naming Food, Identifying Pictures, Picture Description, Caterpillar Passage, Repeat Words, Role Naming,
    Repeat Sentences); Voice/Non-Speech tasks (Long Sounds, Noisy Sounds, Silly Sounds /PUH TUH KUH/).

    '
- id: voice:acquisition:3
  description: 'Self-reported demographic data and medical history questionnaires administered through
    app. Validated questionnaires integrated for each disease cohort including: VHI-10, PHQ-9, GAD-7,
    PANAS, Custom Affect Scale, PTSD Adult, ADHD Adult, DSM-5 Adult, Dyspnea Index, Leicester Cough Questionnaire,
    Winograd Questionnaire, MOCA, C-VHI-10 (peds), PVOS (peds), PVRQOL (peds), PHQ-A (peds). Confounders
    questionnaire about smoking history, drinking history, and other acoustic confounders.

    '
- id: voice:acquisition:4
  description: 'Electronic health record (EHR) access for consenting participants permitting investigators
    to access medical information through EHR platforms to perform gold standard validation of diagnoses
    and symptoms. Linkage to multimodal health biomarkers including laryngoscopy imaging, radiomics, genomics,
    respiratory function tests.

    '
✓ 1/1 R10 6.Data Provenance and Version Tracking Provenance and Source Derivation Documented
evidenceMultiple provenance indicators: collection_mechanisms (Bridge2AI-Voice App on iPad with Avid AE-36 microphone), acquisition_methods (22 adult tasks, 36 pediatric tasks via reproschema-ui), data_collectors (research teams at 5 North American sites), preprocessing tools (b2aiprep, OpenSMILE, Parselmouth, Whisper), source cohort descriptions, IRB study ID (OT2OD032720)
qualityComprehensive provenance: data sources (5 collection sites), collection instruments (iPad hardware, microphone specs), software (Bridge2AI-Voice App, reproschema-ui), preprocessing pipeline (b2aiprep library), research teams (site clinicians, research assistants), IRB study identifier
semanticprovenance_completeness: excellent; source_traceability: yes
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Data Acquisition Methods Listed
evidenceacquisition_methods: 22 adult acoustic tasks (Non-Voice: Respiration, Cough; Voice/Non-Speech: Prolonged Vowel, MPT, Glides, DDK; Speech: Rainbow Passage, Free Speech, Stroop, etc.); 36 pediatric tasks; validated questionnaires (VHI-10, PHQ-9, GAD-7, MOCA, etc.); EHR access for multimodal validation
qualityComprehensive acquisition documentation: 22 adult acoustic tasks with specific names (Rainbow Passage, Caterpillar Passage, DDK /pa/ta/ka/), 36 pediatric tasks, disease-specific validated questionnaires (15+ named instruments), EHR linkage for validation (imaging, genomics, respiratory function)
semanticacquisition_detail: excellent; task_specificity: high
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidencecollection_mechanisms: 3 detailed entries with hardware specs (iPad 9th/10th gen, iPad Air 5th gen, Avid AE-36 microphone, Apple dongle), custom app, reproschema-ui for pediatric; acquisition_methods: 4 entries detailing 22 adult tasks, 36 pediatric tasks, validated questionnaires, EHR linkage; data_collectors: research teams at 5 sites with compensation details; collection_timeframes: September 2022 - November 2026 with version-specific release dates; direct_collection: clinic-based with IRB consent
qualityExceptionally comprehensive collection protocol with specific hardware, software, tasks per cohort, site details, and timeline
correctnessCollection methods appropriate for clinical voice research; hardware choices (Avid AE-36 microphone) professional-grade
consistencyCollection mechanisms align with BIDS output, acoustic tasks match described features, timeframe matches funding period and version releases
data_collectors
data_collectors:
- id: voice:datacollector:1
  description: 'Research teams at each of the five North American collection sites, including medical
    graduate, and undergraduate students coordinating with site clinicians and doctors. Clinicians and
    doctors listed under IRB as co-investigators and added to consortium. Participants compensated via
    electronic gift cards: $40 for sessions under 90 minutes, $80 for sessions over 90 minutes, maximum
    3 sessions and $120 total compensation.

    '
✓ 1/1 R10 6.Data Provenance and Version Tracking Provenance and Source Derivation Documented
evidenceMultiple provenance indicators: collection_mechanisms (Bridge2AI-Voice App on iPad with Avid AE-36 microphone), acquisition_methods (22 adult tasks, 36 pediatric tasks via reproschema-ui), data_collectors (research teams at 5 North American sites), preprocessing tools (b2aiprep, OpenSMILE, Parselmouth, Whisper), source cohort descriptions, IRB study ID (OT2OD032720)
qualityComprehensive provenance: data sources (5 collection sites), collection instruments (iPad hardware, microphone specs), software (Bridge2AI-Voice App, reproschema-ui), preprocessing pipeline (b2aiprep library), research teams (site clinicians, research assistants), IRB study identifier
semanticprovenance_completeness: excellent; source_traceability: yes
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidencecollection_mechanisms: 3 detailed entries with hardware specs (iPad 9th/10th gen, iPad Air 5th gen, Avid AE-36 microphone, Apple dongle), custom app, reproschema-ui for pediatric; acquisition_methods: 4 entries detailing 22 adult tasks, 36 pediatric tasks, validated questionnaires, EHR linkage; data_collectors: research teams at 5 sites with compensation details; collection_timeframes: September 2022 - November 2026 with version-specific release dates; direct_collection: clinic-based with IRB consent
qualityExceptionally comprehensive collection protocol with specific hardware, software, tasks per cohort, site details, and timeline
correctnessCollection methods appropriate for clinical voice research; hardware choices (Avid AE-36 microphone) professional-grade
consistencyCollection mechanisms align with BIDS output, acoustic tasks match described features, timeframe matches funding period and version releases
collection_timeframes
collection_timeframes:
- id: voice:timeframe:1
  description: 'Data collection ongoing from September 2022 through November 2026 (project end date).
    Initial dataset releases began in late 2024. Version 1.0 released January 17, 2025 (306 participants).
    Version 3.0 released 2025 (833 adult participants). Semi-annual releases planned. Enrollment anticipated
    to reach 10,000 participants by 2027.

    '
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Collection Mechanisms and Settings Described
evidencecollection_mechanisms: Voice data collected in clinic using custom Bridge2AI-Voice App on iPad (9th/10th gen or iPad Air 5th gen) with Avid AE-36 microphone and Apple dongle connector; research assistant present; collection_timeframes: September 2022 - November 2026; semi-annual releases planned
qualityDetailed collection documentation: specific hardware (iPad models, Avid AE-36 microphone, Apple dongle), software (Bridge2AI-Voice App, reproschema-ui for pediatric), setting (in-clinic with research assistant), timeframe (Sep 2022 - Nov 2026), sites (5 North American specialty clinics)
semanticmechanism_specificity: excellent; replicability: high
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidencecollection_mechanisms: 3 detailed entries with hardware specs (iPad 9th/10th gen, iPad Air 5th gen, Avid AE-36 microphone, Apple dongle), custom app, reproschema-ui for pediatric; acquisition_methods: 4 entries detailing 22 adult tasks, 36 pediatric tasks, validated questionnaires, EHR linkage; data_collectors: research teams at 5 sites with compensation details; collection_timeframes: September 2022 - November 2026 with version-specific release dates; direct_collection: clinic-based with IRB consent
qualityExceptionally comprehensive collection protocol with specific hardware, software, tasks per cohort, site details, and timeline
correctnessCollection methods appropriate for clinical voice research; hardware choices (Avid AE-36 microphone) professional-grade
consistencyCollection mechanisms align with BIDS output, acoustic tasks match described features, timeframe matches funding period and version releases
direct_collection
direct_collection:
- id: voice:directcoll:1
  description: 'Data collected directly from individuals in clinic settings by trained research assistants
    using the Bridge2AI-Voice App. Participants notified and consented through IRB-approved process. Clinical
    diagnoses and EHR data obtained with participant consent from site clinicians. Data collected in USA
    and Canada.

    '
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidencecollection_mechanisms: 3 detailed entries with hardware specs (iPad 9th/10th gen, iPad Air 5th gen, Avid AE-36 microphone, Apple dongle), custom app, reproschema-ui for pediatric; acquisition_methods: 4 entries detailing 22 adult tasks, 36 pediatric tasks, validated questionnaires, EHR linkage; data_collectors: research teams at 5 sites with compensation details; collection_timeframes: September 2022 - November 2026 with version-specific release dates; direct_collection: clinic-based with IRB consent
qualityExceptionally comprehensive collection protocol with specific hardware, software, tasks per cohort, site details, and timeline
correctnessCollection methods appropriate for clinical voice research; hardware choices (Avid AE-36 microphone) professional-grade
consistencyCollection mechanisms align with BIDS output, acoustic tasks match described features, timeframe matches funding period and version releases
preprocessing_strategies
preprocessing_strategies:
- id: voice:preproc:1
  description: 'Raw audio preprocessing using b2aiprep library: conversion to monaural audio, resampling
    to 16 kHz with Butterworth anti-aliasing filter. Standardization ensures consistent format across
    all recordings for downstream feature extraction.

    '
- id: voice:preproc:2
  description: 'Spectrogram extraction using short-time Fast Fourier Transform (FFT): 25ms window size,
    10ms hop length, 512-point FFT. Output spectrograms have 513xN dimensions where N is proportional
    to audio length. Stored in Parquet format.

    '
- id: voice:preproc:3
  description: 'Mel-frequency cepstral coefficients (MFCC) extraction: 60 MFCCs extracted from spectrograms
    capturing perceptually-relevant spectral envelope characteristics. Output dimension 60xN. Stored in
    Parquet format.

    '
- id: voice:preproc:4
  description: 'OpenSMILE eGeMaps acoustic feature extraction capturing temporal dynamics and acoustic
    characteristics. One row per unique recording in static_features.tsv.

    '
- id: voice:preproc:5
  description: 'Parselmouth and Praat phonetic and prosodic feature computation providing fundamental
    frequency (F0), formants, and voice quality measures. Documented in static_features.json data dictionary.

    '
- id: voice:preproc:6
  description: 'Torchaudio-based feature extraction including pitch contour, spectrograms, mel spectrograms,
    MFCCs. SPARC-based features including electromagnetic articulography (EMA) estimates, loudness, periodicity,
    and pitch measures. Phonetic posteriorgrams (PPGs). Stored in two formats: fixed static format and
    temporal format varying by recording length.

    '
- id: voice:preproc:7
  description: 'Transcription generation using OpenAI Whisper model applied to audio recordings. Raw audio
    transcripts reviewed and any recordings containing potentially identifying information or external
    voices removed. All transcriptions, EMAs, and PPGs from open-response prompts removed from public
    feature-only dataset.

    '
- id: voice:preproc:8
  description: 'REDCap and reproschema-ui data exported and converted to BIDS v1.9.0 format using b2aiprep
    open-source library. Phenotype data organized in tab-delimited files with JSON data dictionaries.
    Pediatric data first extracted from reproschema-ui to REDCap format then converted to BIDS.

    '
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Preprocessing, Cleaning, and Labeling Strategies
evidencepreprocessing_strategies: 8 detailed strategies including b2aiprep audio preprocessing (monaural, 16 kHz resampling with Butterworth filter), spectrogram extraction (25ms window, 10ms hop, 512-point FFT), MFCC extraction (60 coefficients), OpenSMILE eGeMaps, Parselmouth/Praat features, torchaudio features, Whisper transcription, REDCap→BIDS conversion; cleaning_strategies: 4 strategies (HIPAA Safe Harbor, audio privacy protection, sensitive field removal, audit protocol); labeling_strategies: clinical diagnosis by site clinicians with gold standard validation
qualityExceptional preprocessing/cleaning/labeling documentation: 8 preprocessing strategies with technical parameters (FFT size, window/hop lengths, MFCC count, sampling rates), 4 cleaning strategies (de-identification, privacy filtering, sensitive field removal, QC audit), labeling via clinical assessment with validation methods (imaging, genomics)
semanticpreprocessing_detail: exceptional; parameter_specificity: excellent
5/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelComprehensive strategies with software versions/URLs
evidencepreprocessing_strategies: 8 detailed entries with software tools (b2aiprep, OpenSMILE eGeMaps, Parselmouth, Praat, torchaudio, SPARC, OpenAI Whisper, REDCap, reproschema-ui); cleaning_strategies: 4 entries; labeling_strategies: 1 entry with validation methods; external_resources: GitHub links to b2aiprep (Apache-2.0), bridge2ai-redcap (MIT), SenseLab; specific parameters documented (16 kHz resampling, 25ms FFT window, 60 MFCCs)
qualityExceptional software transparency with specific tools, processing parameters, open-source repository links, and licensing information
correctnessSoftware tools appropriate for acoustic feature extraction and clinical data management; processing parameters (16kHz, 60 MFCCs) standard for speech analysis
consistencyTools align with described features (OpenSMILE→eGeMaps, Whisper→transcripts, b2aiprep→BIDS conversion)
cleaning_strategies
cleaning_strategies:
- id: voice:cleaning:1
  description: 'HIPAA Safe Harbor de-identification for public release: removed direct identifiers (names,
    civic addresses, social security numbers), indirect identifiers creating significant re-identification
    risk (select geographic/demographic identifiers, household composition, cultural identity), and sensitive
    information (household income, mental health status, traumatic life experiences). Geographic data
    limited; state/province removed, country retained.

    '
- id: voice:cleaning:2
  description: 'Audio privacy protection for public release: all raw audio waveforms excluded from public
    dataset. All spectrograms, MFCCs, mel spectrograms, transcriptions, EMAs, and PPGs from open-response
    prompts removed from feature-only dataset. Free speech transcripts removed. Raw audio available only
    through controlled access with DACO approval. Raw data stored and retained separately for verified
    researchers.

    '
- id: voice:cleaning:3
  description: 'Sensitive field removal based on REDCap data dictionary: all fields encoded as sensitive
    (column "Identifier?" in REDCap data dictionary CSV) removed from dataset.

    '
- id: voice:cleaning:4
  description: 'Audit protocol applied: missingness tables generated and included with dataset; distribution
    and outlier checks; categorical responses checked against schema; audio quality control metrics including
    silence amount, duration, and speech-to-passage accuracy checks. Processing using b2aiprep and SenseLab
    toolkits.

    '
✓ 1/1 R10 5.Data Composition and Structure Data Quality Issues and Anomalies Documented
evidencesampling_strategies: Non-probability purposive sampling, not representative of general population, selection bias acknowledged; cleaning_strategies: audit protocol with missingness tables, distribution/outlier checks, categorical response validation, audio quality control (silence, duration, speech-to-passage accuracy)
qualityData quality documentation includes: sampling bias acknowledgment (clinic-based, limited geographic sites), representativeness limitations (English-only, treatment-seeking populations), audit protocol (missingness tables, outlier checks, quality metrics), audio QC metrics
semanticquality_awareness: excellent; bias_acknowledgment: yes
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Preprocessing, Cleaning, and Labeling Strategies
evidencepreprocessing_strategies: 8 detailed strategies including b2aiprep audio preprocessing (monaural, 16 kHz resampling with Butterworth filter), spectrogram extraction (25ms window, 10ms hop, 512-point FFT), MFCC extraction (60 coefficients), OpenSMILE eGeMaps, Parselmouth/Praat features, torchaudio features, Whisper transcription, REDCap→BIDS conversion; cleaning_strategies: 4 strategies (HIPAA Safe Harbor, audio privacy protection, sensitive field removal, audit protocol); labeling_strategies: clinical diagnosis by site clinicians with gold standard validation
qualityExceptional preprocessing/cleaning/labeling documentation: 8 preprocessing strategies with technical parameters (FFT size, window/hop lengths, MFCC count, sampling rates), 4 cleaning strategies (de-identification, privacy filtering, sensitive field removal, QC audit), labeling via clinical assessment with validation methods (imaging, genomics)
semanticpreprocessing_detail: exceptional; parameter_specificity: excellent
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Data Anomalies and Quality Issues Noted
evidencecleaning_strategies: audit protocol with missingness tables included with dataset, distribution and outlier checks performed, categorical response validation against schema, audio quality control metrics (silence amount, duration, speech-to-passage accuracy checks); sensitive_elements: transcriptions may contain potentially identifying information or external voices
qualityQuality issues and anomalies documented: missingness tables generated and included, outlier detection performed, categorical validation checks, audio QC metrics (silence, duration, accuracy), transcription risks identified (identifying info, external voices), audit protocol applied via b2aiprep/SenseLab
semanticquality_awareness: excellent; qc_metrics_defined: yes
5/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelComprehensive strategies with software versions/URLs
evidencepreprocessing_strategies: 8 detailed entries with software tools (b2aiprep, OpenSMILE eGeMaps, Parselmouth, Praat, torchaudio, SPARC, OpenAI Whisper, REDCap, reproschema-ui); cleaning_strategies: 4 entries; labeling_strategies: 1 entry with validation methods; external_resources: GitHub links to b2aiprep (Apache-2.0), bridge2ai-redcap (MIT), SenseLab; specific parameters documented (16 kHz resampling, 25ms FFT window, 60 MFCCs)
qualityExceptional software transparency with specific tools, processing parameters, open-source repository links, and licensing information
correctnessSoftware tools appropriate for acoustic feature extraction and clinical data management; processing parameters (16kHz, 60 MFCCs) standard for speech analysis
consistencyTools align with described features (OpenSMILE→eGeMaps, Whisper→transcripts, b2aiprep→BIDS conversion)
labeling_strategies
labeling_strategies:
- id: voice:labeling:1
  description: 'Diagnostic labels assigned by site clinicians based on clinical assessment and gold standard
    validation methods per Bridge2AI-Voice protocol Table 1 and ICD-10 codes. For voice disorders: laryngoscopy
    and stroboscopy. For neurological disorders: CT Brain, MRI Brain, genome sequencing, serum biomarkers.
    For mood/psychiatric disorders: EHR records, psychiatrist diagnosis, medication history. For respiratory
    disorders: spirometry, flow volume loops, CT scan. Single labeler (site clinician) per participant.
    Future labels should include description of exact variables used for determination.

    '
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Preprocessing, Cleaning, and Labeling Strategies
evidencepreprocessing_strategies: 8 detailed strategies including b2aiprep audio preprocessing (monaural, 16 kHz resampling with Butterworth filter), spectrogram extraction (25ms window, 10ms hop, 512-point FFT), MFCC extraction (60 coefficients), OpenSMILE eGeMaps, Parselmouth/Praat features, torchaudio features, Whisper transcription, REDCap→BIDS conversion; cleaning_strategies: 4 strategies (HIPAA Safe Harbor, audio privacy protection, sensitive field removal, audit protocol); labeling_strategies: clinical diagnosis by site clinicians with gold standard validation
qualityExceptional preprocessing/cleaning/labeling documentation: 8 preprocessing strategies with technical parameters (FFT size, window/hop lengths, MFCC count, sampling rates), 4 cleaning strategies (de-identification, privacy filtering, sensitive field removal, QC audit), labeling via clinical assessment with validation methods (imaging, genomics)
semanticpreprocessing_detail: exceptional; parameter_specificity: excellent
5/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelComprehensive strategies with software versions/URLs
evidencepreprocessing_strategies: 8 detailed entries with software tools (b2aiprep, OpenSMILE eGeMaps, Parselmouth, Praat, torchaudio, SPARC, OpenAI Whisper, REDCap, reproschema-ui); cleaning_strategies: 4 entries; labeling_strategies: 1 entry with validation methods; external_resources: GitHub links to b2aiprep (Apache-2.0), bridge2ai-redcap (MIT), SenseLab; specific parameters documented (16 kHz resampling, 25ms FFT window, 60 MFCCs)
qualityExceptional software transparency with specific tools, processing parameters, open-source repository links, and licensing information
correctnessSoftware tools appropriate for acoustic feature extraction and clinical data management; processing parameters (16kHz, 60 MFCCs) standard for speech analysis
consistencyTools align with described features (OpenSMILE→eGeMaps, Whisper→transcripts, b2aiprep→BIDS conversion)
intended_uses
intended_uses:
- id: voice:use:1
  description: 'Primary intended use: development and validation of AI/ML models for voice as a biomarker
    of health, supporting screening, diagnosis, and treatment of voice disorders, neurological disorders,
    mood disorders, respiratory disorders, and pediatric speech disorders through model pretraining, fine-tuning,
    benchmarking, and validation.

    '
- id: voice:use:2
  description: 'Research into acoustic biomarkers and development of standards for voice data collection
    and analysis. Establishing best practices for AI/ML-friendly voice datasets and contributing to the
    field''s maturation as a clinical diagnostic modality.

    '
- id: voice:use:3
  description: 'Training and education in voice AI research through workforce development initiatives,
    curriculum creation on FAIR and CARE voice AI model development, and fostering collaborations between
    medical voice researchers, acoustic engineers, and AI/ML specialists, especially from underserved
    communities.

    '
- id: voice:use:4
  description: 'Multimodal health research combining voice data with EHR information, radiomics, genomics,
    imaging, and other health biomarkers to understand complex disease relationships and improve diagnostic
    accuracy.

    '
✓ 1/1 R10 3.Data Reuse and Interoperability Use Guidance Provided
evidenceintended_uses: 4 entries (AI/ML model development, acoustic biomarker research, training/education, multimodal health research); discouraged_uses: 3 entries (hiring, surveillance, re-identification); prohibited_uses: 1 entry (unauthorized use, sale, re-identification attempts)
qualityClear use guidance across three categories: intended (4 specific use cases), discouraged (3 harmful applications), prohibited (unauthorized use, sale, re-identification) - comprehensive ethical and practical guidance
semanticguidance_completeness: excellent; ethical_considerations: well_defined
existing_uses
existing_uses:
- id: voice:existinguse:1
  description: 'A restricted version of the dataset containing raw audio has been used in the Bridge2AI
    Summer School and hackathon for education and research training purposes.

    '
no field-level feedback matched
discouraged_uses
discouraged_uses:
- id: voice:discouraged:1
  description: 'Non-clinical applications such as hiring decisions, insurance premium adjustments, or
    any form of surveillance that could lead to discrimination or harm based on health conditions or voice
    characteristics. These applications could negatively impact individuals.

    '
- id: voice:discouraged:2
  description: 'Any attempt to re-identify research participants or use data in ways that could foreseeably
    cause harm or stigmatization to research participants, their families, communities, or specific populations.
    Dataset covered under Certificate of Confidentiality.

    '
- id: voice:discouraged:3
  description: 'Development or use of intellectual property protections, database rights, or related rights
    in ways that would prevent or block access to any element of the dataset or conclusions derived from
    it. Must respect Fort Lauderdale Agreement and Open Science principles.

    '
✓ 1/1 R10 3.Data Reuse and Interoperability Use Guidance Provided
evidenceintended_uses: 4 entries (AI/ML model development, acoustic biomarker research, training/education, multimodal health research); discouraged_uses: 3 entries (hiring, surveillance, re-identification); prohibited_uses: 1 entry (unauthorized use, sale, re-identification attempts)
qualityClear use guidance across three categories: intended (4 specific use cases), discouraged (3 harmful applications), prohibited (unauthorized use, sale, re-identification) - comprehensive ethical and practical guidance
semanticguidance_completeness: excellent; ethical_considerations: well_defined
prohibited_uses
prohibited_uses:
- id: voice:prohibited:1
  description: 'Use of the dataset outside of authorized research purposes as defined in the Bridge2AI
    Voice Registered Access Agreement. The dataset is intended solely for commercial and non-commercial
    research by Authorized Researchers. Sale of all or part of the data on any media is prohibited. Re-identification
    attempts are prohibited.

    '
✓ 1/1 R10 3.Data Reuse and Interoperability Use Guidance Provided
evidenceintended_uses: 4 entries (AI/ML model development, acoustic biomarker research, training/education, multimodal health research); discouraged_uses: 3 entries (hiring, surveillance, re-identification); prohibited_uses: 1 entry (unauthorized use, sale, re-identification attempts)
qualityClear use guidance across three categories: intended (4 specific use cases), discouraged (3 harmful applications), prohibited (unauthorized use, sale, re-identification) - comprehensive ethical and practical guidance
semanticguidance_completeness: excellent; ethical_considerations: well_defined
license
license: Bridge2AI Voice Registered Access License
✓ 1/1 R10 3.Data Reuse and Interoperability License Terms Allow Reuse
evidencelicense: Bridge2AI Voice Registered Access License; Commercial and non-commercial research use permitted for Authorized Researchers; Recipient encouraged to publish results in open-access journals
qualityClear license allowing both commercial and non-commercial research use, open-access publication encouraged, registered access mechanism defined with DUA requirements
semanticreuse_permissions: clear_permissive; commercial_use: permitted
license_and_use_terms
license_and_use_terms:
  id: voice:license:1
  name: Bridge2AI Voice Registered Access License
  description: 'Public access dataset distributed through PhysioNet under Bridge2AI Voice Registered Access
    License. Only registered users who sign the specified Data Use Agreement (Bridge2AI Voice Registered
    Access Agreement) can access files. Data covered under Certificate of Confidentiality which must be
    asserted against compulsory legal demands. Raw audio available through controlled access only via
    Data Access Compliance Office (DACO) requiring distinct application and DTUA signed by institutional
    official. Recipient must adhere to PhysioNet requirements managed by MIT Laboratory for Computational
    Physiology. Recipient encouraged to publish results in open-access journals. No export controls apply.
    Commercial and non-commercial research use permitted for Authorized Researchers.

    '
✓ 1/1 R10 2.Dataset Access and Retrieval Access Policy and IP Restrictions Defined
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License with detailed DUA requirements, PhysioNet registered access for features, controlled access via DACO for raw audio requiring institutional DTUA
qualityComprehensive access policy with two-tier system (registered vs controlled access), clear DUA requirements, institutional authorization requirements (DTUA signed by official), no IP restrictions mentioned, no export controls apply
semanticpolicy_clarity: excellent; access_tiers: well_defined
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information, description: 400+ chars comprehensive, keywords: 86 keywords, license_and_use_terms: detailed with access mechanisms
qualityAll mandatory fields present with exceptional content quality
correctnessAll fields semantically appropriate and correctly formatted
consistencyFields align with dataset scope and access model
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path (platform, login, policy)
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License via PhysioNet requiring DUA signature for feature dataset; controlled access for raw audio via DACO@b2ai-voice.org requiring Data Access Request Form (DARF), institutional Data Use and Transfer Agreement (DTUA) signed by authorized official; distribution_formats specify which data in each tier; third_party_sharing restrictions detailed
qualityExceptionally clear two-tier access mechanism with specific forms, contact emails, signing authorities, and data types per tier
correctnessAccess model appropriate for sensitive health data with tiered risk-based controls
consistencyAccess tiers align with privacy risk (features=registered, raw audio=controlled) and deidentification approach
5/5 R20 Q18 (FAIRness & Accessibility) Reusability (License Clarity)
levelLicense explicitly defines reuse terms
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License explicitly permits commercial and non-commercial research use for Authorized Researchers; no geographic restriction; no restriction to specific research type; encouraged open-access publication; prohibited: sale of data, re-identification, IP restrictions blocking access; Fort Lauderdale Agreement and Open Science principles mentioned; retention and destruction requirements specified
qualityExceptional license clarity with explicit permissions (commercial use allowed), prohibitions (no sale/re-identification), and alignment with open science principles
correctnessLicense terms legally coherent and appropriate for NIH-funded data resource
consistencyReuse permissions align with intended_uses (AI/ML development, research, education) and discouraged_uses (hiring, surveillance)
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License with detailed DUA requirements, DACO controlled access for raw audio, DTUA for institutional sharing, Certificate of Confidentiality; third_party_sharing: clear restrictions on resale and redistribution; retention_limit: 2-year DTUA term with destruction requirements
qualityComprehensive governance with two-tier access model (registered for features, controlled for raw audio) and detailed compliance requirements
correctnessAccess model appropriate for sensitive biometric health data with tiered risk-based controls
consistencyLicense terms align with privacy protections, deidentification approach, and Certificate of Confidentiality
distribution_formats
distribution_formats:
- id: voice:format:1
  name: Parquet format for spectrograms and derived features
  description: 'Time-varying features (spectrograms, mel spectrograms, MFCCs, pitch contour, SPARC features,
    PPGs) stored in Parquet format compatible with Python datasets library. Each element contains participant_id,
    session_id, task_name, and feature arrays. Excluded for open-response tasks in public release.

    '
- id: voice:format:2
  name: TSV/JSON for phenotype and static features
  description: 'Phenotype data in tab-delimited format (confounders.tsv, demographics.tsv, diagnosis/
    *.tsv, enrollment/*.tsv, questionnaire/*.tsv, task/*.tsv) with JSON data dictionaries per file. Static
    acoustic features (OpenSMILE, Praat, Parselmouth, torchaudio) in static_features.tsv with static_features.json
    data dictionary. One row per participant or recording as appropriate. BIDS v1.9.0 compliant structure.

    '
- id: voice:format:3
  name: WAV audio format (controlled access only)
  description: 'Original raw audio waveforms in WAV format following BIDS structure: sub-{participant_id}/ses-{session_id}/audio/sub{id}_ses{id}_task-{task_name}.wav
    with companion JSON metadata. Available through controlled access only. Contact DACO@b2ai-voice.org
    to request access.

    '
✓ 1/1 R10 2.Dataset Access and Retrieval Distribution Formats and File Types Specified
evidencedistribution_formats: Parquet (spectrograms, derived features), TSV/JSON (phenotype, static features), WAV audio (controlled access only); BIDS v1.9.0 compliant structure; MIME types implicit from format descriptions
qualitySpecific file formats documented: Parquet for time-varying features, TSV/JSON for phenotype/static features with data dictionaries, WAV for raw audio; BIDS v1.9.0 structure compliance stated
semanticformat_specificity: excellent; standard_compliance: BIDS_v1.9.0
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path (platform, login, policy)
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License via PhysioNet requiring DUA signature for feature dataset; controlled access for raw audio via DACO@b2ai-voice.org requiring Data Access Request Form (DARF), institutional Data Use and Transfer Agreement (DTUA) signed by authorized official; distribution_formats specify which data in each tier; third_party_sharing restrictions detailed
qualityExceptionally clear two-tier access mechanism with specific forms, contact emails, signing authorities, and data types per tier
correctnessAccess model appropriate for sensitive health data with tiered risk-based controls
consistencyAccess tiers align with privacy risk (features=registered, raw audio=controlled) and deidentification approach
5/5 R20 Q4 (Structural Completeness) File Enumeration and Type Variety
level>3 file types
evidencedistribution_formats: Parquet (spectrograms, MFCCs, temporal features), TSV/JSON (phenotype, static features, BIDS metadata), WAV (raw audio, controlled access). Media types: application/parquet, text/tab-separated-values, application/json, audio/wav
qualityExcellent format diversity supporting multiple analysis approaches from raw audio to derived features
correctnessFile formats appropriate for acoustic data and clinical metadata
consistencyFormat descriptions align with preprocessing pipeline and access tiers
distribution_dates
distribution_dates:
- id: voice:distdate:1
  description: 'Dataset first published and made available late November 2024 through Health Data Nexus.
    PhysioNet releases: v1.0 January 17, 2025 (306 participants, 12,523 recordings); v1.1 January 17,
    2025 (added MFCC features); v2.0.0 April 16, 2025; v2.0.1 August 18, 2025; v3.0.0 released 2025 (833
    adult participants, ~61,937 recordings). Semi-annual releases planned. Pediatric v1.0 available separately.

    '
5/5 R20 Q13 (Technical Documentation) Version History Documentation
levelComprehensive versioning with errata, updates, and release notes
evidenceversion: 3.0.0; updates: detailed with v1.0 (Jan 17, 2025, 306 participants, 12,523 recordings), v1.1 (MFCC features added), v2.0.0 (Apr 16, 2025), v2.0.1 (Aug 18, 2025), v3.0.0 (833 adults, 61,937 recordings), pediatric v1.0 (300 participants), semi-annual release plan, target 10,000 by 2027; version_access: all versions hosted with version-specific DOIs on PhysioNet and Health Data Nexus; distribution_dates: release timeline documented
qualityExceptional version tracking with specific dates, participant/recording counts per version, feature additions, and perpetual access policy
correctnessVersion progression logical (increasing participant counts, feature additions); dates align with project timeline
consistencyVersion information consistent across updates, distribution_dates, and version_access sections
third_party_sharing
third_party_sharing:
- id: voice:thirdparty:1
  description: 'Dataset distributed broadly to individuals outside the creating entity through PhysioNet
    (registered and controlled access) and Health Data Nexus. Collaborators at other research organizations
    must apply independently and sign a DTUA. Recipient shall not disclose, release, sell, rent, or lease
    data to third parties without prior written consent of Provider. Authorized Persons at Recipient institution
    may access data under same agreement.

    '
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path (platform, login, policy)
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License via PhysioNet requiring DUA signature for feature dataset; controlled access for raw audio via DACO@b2ai-voice.org requiring Data Access Request Form (DARF), institutional Data Use and Transfer Agreement (DTUA) signed by authorized official; distribution_formats specify which data in each tier; third_party_sharing restrictions detailed
qualityExceptionally clear two-tier access mechanism with specific forms, contact emails, signing authorities, and data types per tier
correctnessAccess model appropriate for sensitive health data with tiered risk-based controls
consistencyAccess tiers align with privacy risk (features=registered, raw audio=controlled) and deidentification approach
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License with detailed DUA requirements, DACO controlled access for raw audio, DTUA for institutional sharing, Certificate of Confidentiality; third_party_sharing: clear restrictions on resale and redistribution; retention_limit: 2-year DTUA term with destruction requirements
qualityComprehensive governance with two-tier access model (registered for features, controlled for raw audio) and detailed compliance requirements
correctnessAccess model appropriate for sensitive biometric health data with tiered risk-based controls
consistencyLicense terms align with privacy protections, deidentification approach, and Certificate of Confidentiality
maintainers
maintainers:
- id: voice:maintainer:1
  description: 'Bridge2AI-Voice Consortium (University of South Florida, lead institution) supported by
    NIH Bridge2AI program. Contact: [email protected]. Responsible for dataset curation, standards development,
    ethics oversight, versioning, and updates. Data Access Compliance Office (DACO) manages controlled
    access applications. Contact: DACO@b2ai-voice.org.

    '
- id: voice:maintainer:2
  description: 'PhysioNet (MIT Laboratory for Computational Physiology, supported by NIBIB NIH grant R01EB030362)
    serves as primary distribution platform for registered access dataset. Provides technical infrastructure
    and access management.

    '
- id: voice:maintainer:3
  description: 'Health Data Nexus (Temerty Centre for Artificial Intelligence Research and Education in
    Medicine, T-CAIREM, University of Toronto) maintains alternative platform providing cloud compute
    alongside dataset. Earlier dataset versions available here. Contact: [email protected].

    '
no field-level feedback matched
updates
updates:
  id: voice:updates:1
  name: Versioned releases with ongoing data collection
  description: 'Dataset updated with versioned static releases semi-annually as data collection progresses.
    Users notified through news items on platforms and standard communication channels. v1.0 released
    January 17, 2025 (306 participants, 12,523 recordings); v1.1 added MFCC features; v2.0.0 April 16,
    2025; v2.0.1 August 18, 2025; v3.0.0 released 2025 (833 adults, ~61,937 recordings). Pediatric v1.0
    released separately. Data collection ongoing through November 30, 2026. Target: 10,000 participants
    by 2027. Future releases will expand Spanish language protocols and add additional multimodal data
    (imaging, genomics). Version-specific DOIs maintained. Older versions continue to be supported and
    hosted.

    '
✓ 1/1 R10 6.Data Provenance and Version Tracking Change Descriptions and Errata Provided
evidenceupdates: v1.0 released January 17, 2025 (306 participants, 12,523 recordings); v1.1 added MFCC features; v2.0.0 April 16, 2025; v2.0.1 August 18, 2025; v3.0.0 released 2025 (833 adults, ~61,937 recordings); Pediatric v1.0 released separately
qualityDetailed version history with changes: v1.0→v1.1 (MFCC features added), v1.1→v2.0.0 (major update), v2.0.0→v2.0.1 (patch), v2.0.1→v3.0.0 (833 participants, ~61,937 recordings); participant/recording counts per version
semanticchange_documentation: excellent; version_progression: logical
✓ 1/1 R10 6.Data Provenance and Version Tracking Update Schedule or Frequency Indicated
evidenceupdates: Semi-annual releases planned; Data collection ongoing from September 2022 through November 2026; Enrollment anticipated to reach 10,000 participants by 2027; Users notified through news items on platforms and standard communication channels
qualityClear update schedule: semi-annual releases planned, collection timeline (Sep 2022 - Nov 2026), enrollment target (10,000 by 2027), user notification mechanisms (platform news items, standard channels)
semanticschedule_clarity: excellent; ongoing_commitment: yes
5/5 R20 Q13 (Technical Documentation) Version History Documentation
levelComprehensive versioning with errata, updates, and release notes
evidenceversion: 3.0.0; updates: detailed with v1.0 (Jan 17, 2025, 306 participants, 12,523 recordings), v1.1 (MFCC features added), v2.0.0 (Apr 16, 2025), v2.0.1 (Aug 18, 2025), v3.0.0 (833 adults, 61,937 recordings), pediatric v1.0 (300 participants), semi-annual release plan, target 10,000 by 2027; version_access: all versions hosted with version-specific DOIs on PhysioNet and Health Data Nexus; distribution_dates: release timeline documented
qualityExceptional version tracking with specific dates, participant/recording counts per version, feature additions, and perpetual access policy
correctnessVersion progression logical (increasing participant counts, feature additions); dates align with project timeline
consistencyVersion information consistent across updates, distribution_dates, and version_access sections
4/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelStructured version control with timestamps
evidenceupdates: comprehensive version history with dates and participant counts per version (v1.0 Jan 17 2025, v1.1, v2.0.0 Apr 16 2025, v2.0.1 Aug 18 2025, v3.0.0); version_access: version-specific DOIs maintained; missing: explicit errata documentation or formal change log structure
qualityStrong provenance through versioned releases with dates and metrics; could benefit from formal errata/change log documentation
correctnessVersion dates follow logical temporal progression aligned with collection timeline
consistencyVersion history consistent across updates, distribution_dates, and version_access sections; participant counts increase monotonically
retention_limit
retention_limit:
  id: voice:retention:1
  name: Data retention and disposition policy
  description: 'Data Transfer and Use Agreement (DTUA) specifies two-year agreement term from start date.
    Upon termination or expiration (two years, project completion, ethics approval expiration, or provider
    termination, whichever occurs first), data shall be destroyed per provider instructions with written
    certification required within 30 days. Recipient may retain one copy to comply with records retention
    requirements under law, regulation, institutional policy, and for research integrity and verification.
    Ongoing restrictions apply to retained copies. Provider may unilaterally amend agreement if federal
    sponsor requires revision. Health Data Nexus retains dataset as long as useful for research purposes,
    possibly indefinitely. Version-specific DOIs maintained for all historical versions.

    '
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License with detailed DUA requirements, DACO controlled access for raw audio, DTUA for institutional sharing, Certificate of Confidentiality; third_party_sharing: clear restrictions on resale and redistribution; retention_limit: 2-year DTUA term with destruction requirements
qualityComprehensive governance with two-tier access model (registered for features, controlled for raw audio) and detailed compliance requirements
correctnessAccess model appropriate for sensitive biometric health data with tiered risk-based controls
consistencyLicense terms align with privacy protections, deidentification approach, and Certificate of Confidentiality
version_access
version_access:
  id: voice:versionaccess:1
  name: Version access policy
  description: 'All dataset versions available through PhysioNet at https://physionet.org/content/b2ai-voice/
    with version-specific DOIs. DOI for latest version: https://doi.org/10.13026/37yb-1t42. Earlier versions
    also available on Health Data Nexus at https://healthdatanexus.ai/content/b2ai-voice/1.0/. By default,
    older versions continue to be supported, hosted, and made available. Dataset publishers reserve right
    to remove access to older versions. Each version has unique DOI.

    '
✓ 1/1 R10 6.Data Provenance and Version Tracking Version Access Methods Documented
evidenceversion_access: All dataset versions available through PhysioNet at https://physionet.org/content/b2ai-voice/ with version-specific DOIs; earlier versions on Health Data Nexus at https://healthdatanexus.ai/content/b2ai-voice/1.0/; older versions continue to be supported, hosted, and made available; each version has unique DOI
qualityComprehensive version access documentation: all versions available on PhysioNet with version-specific DOIs, alternative platform (Health Data Nexus) for earlier versions, commitment to maintaining older versions, unique DOI per version
semanticversion_access_clarity: excellent; historical_preservation: yes
5/5 R20 Q13 (Technical Documentation) Version History Documentation
levelComprehensive versioning with errata, updates, and release notes
evidenceversion: 3.0.0; updates: detailed with v1.0 (Jan 17, 2025, 306 participants, 12,523 recordings), v1.1 (MFCC features added), v2.0.0 (Apr 16, 2025), v2.0.1 (Aug 18, 2025), v3.0.0 (833 adults, 61,937 recordings), pediatric v1.0 (300 participants), semi-annual release plan, target 10,000 by 2027; version_access: all versions hosted with version-specific DOIs on PhysioNet and Health Data Nexus; distribution_dates: release timeline documented
qualityExceptional version tracking with specific dates, participant/recording counts per version, feature additions, and perpetual access policy
correctnessVersion progression logical (increasing participant counts, feature additions); dates align with project timeline
consistencyVersion information consistent across updates, distribution_dates, and version_access sections
4/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelStructured version control with timestamps
evidenceupdates: comprehensive version history with dates and participant counts per version (v1.0 Jan 17 2025, v1.1, v2.0.0 Apr 16 2025, v2.0.1 Aug 18 2025, v3.0.0); version_access: version-specific DOIs maintained; missing: explicit errata documentation or formal change log structure
qualityStrong provenance through versioned releases with dates and metrics; could benefit from formal errata/change log documentation
correctnessVersion dates follow logical temporal progression aligned with collection timeline
consistencyVersion history consistent across updates, distribution_dates, and version_access sections; participant counts increase monotonically
extension_mechanism
extension_mechanism:
  id: voice:extension:1
  name: Dataset extension and contribution mechanisms
  description: 'Derivative datasets can be published on Health Data Nexus referencing original source
    under same access conditions. Open-source repositories (b2aiprep, SenseLab) have discussion forums,
    issue pages, and pull request mechanisms for contributing improvements to preprocessing code. REDCap
    data dictionary available at https://github.com/eipm/bridge2ai-redcap with MIT license. Bridge2AI-Voice
    documentation at https://github.com/eipm/bridge2ai-docs. Future augmentations to voice collection
    protocol coordinated through consortium.

    '
no field-level feedback matched
human_subject_research
human_subject_research:
  id: voice:hsr:1
  name: Bridge2AI-Voice Human Subjects Research
  description: 'Observational study (cross-sectional design) involving direct collection of voice recordings,
    questionnaire responses, and EHR linkage from human participants. Prospective informed consent obtained
    from all participants. IRB approved by University of South Florida Single IRB with subsite IRBs through
    Single IRB process. Not a drug or medical device study. No data monitoring committee appointed. Study
    ID: OT2OD032720.

    '
  involves_human_subjects: true
  irb_approval:
  - University of South Florida Single IRB approval with subsite IRB approvals through Single IRB process
  ethics_review_board:
  - University of South Florida Institutional Review Board (Single IRB)
  special_populations:
  - Pediatric participants (aged 2-18 years) at Hospital for Sick Children with age-appropriate protocols
  regulatory_compliance:
  - HIPAA Safe Harbor de-identification standards applied
  - Certificate of Confidentiality covering dataset against compulsory legal demands
  - 45 CFR 46 (Common Rule) compliance
  - Data covered as Personally Identifiable Information per OMB Memorandum M-07-16
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards IRB or Ethics Review Documented
evidencehuman_subject_research.irb_approval: University of South Florida Single IRB with subsite IRB approvals through Single IRB process; ethics_review_board: University of South Florida Institutional Review Board; bioethics guidance by consortium bioethicists (Belisle-Pipon, Ravitsky)
qualityComprehensive IRB documentation: University of South Florida Single IRB (lead), subsite approvals via Single IRB process, bioethics team involvement, ethics module developing consent guidelines for voice AI
semanticirb_completeness: excellent; consistency_check: pass_human_subjects_true_with_irb
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive - all human subjects protections documented
evidencehuman_subject_research: detailed with IRB approval (USF Single IRB), involves_human_subjects=true; ethical_reviews: bioethics integrated throughout; is_deidentified: HIPAA Safe Harbor detailed; participant_privacy: 8 layered protections including Certificate of Confidentiality; collection_consents: prospective informed consent; consent_revocations: withdrawal mechanism; participant_compensation: $40-120 via gift cards; at_risk_populations: pediatric protections; sensitive_elements: 5 categories documented
qualityExemplary comprehensive ethical documentation covering all protection areas with specific mechanisms and regulatory compliance
correctnessHIPAA Safe Harbor method correctly applied with appropriate removals for voice biometric data
consistencyStrong alignment: human subjects research → IRB approval → consent → deidentification → privacy protections → vulnerable populations
collection_consents
collection_consents:
- id: voice:consent:1
  description: 'All participants duly informed and provided prospective informed consent for data collection
    and use via IRB-approved consent process. Consent includes: authorization for voice data collection
    and speaking tasks, demographic and medical history questionnaires, disease- specific validated questionnaires,
    access to medical information through EHR platforms, and permission to share research data. Data that
    poses low risk of re-identification shared in registered access; heightened re-identification risk
    data shared through controlled access mechanism. No restriction on commercial vs. non-commercial use;
    no geographic restriction; no restriction to specific research type.

    '
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Informed Consent Obtained from Participants
evidencecollection_consents: All participants duly informed and provided prospective informed consent via IRB-approved process; consent includes voice data collection authorization, questionnaire permission, EHR access, data sharing permission; collection_notifications: participants notified and consented before data collection
qualityDetailed consent documentation: prospective informed consent, IRB-approved process, specific authorizations (voice tasks, questionnaires, EHR access, data sharing), notification before collection, withdrawal rights communicated
semanticconsent_completeness: excellent; consent_type: prospective_informed; consistency_check: pass_human_subjects_true_with_consent_documented
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive - all human subjects protections documented
evidencehuman_subject_research: detailed with IRB approval (USF Single IRB), involves_human_subjects=true; ethical_reviews: bioethics integrated throughout; is_deidentified: HIPAA Safe Harbor detailed; participant_privacy: 8 layered protections including Certificate of Confidentiality; collection_consents: prospective informed consent; consent_revocations: withdrawal mechanism; participant_compensation: $40-120 via gift cards; at_risk_populations: pediatric protections; sensitive_elements: 5 categories documented
qualityExemplary comprehensive ethical documentation covering all protection areas with specific mechanisms and regulatory compliance
correctnessHIPAA Safe Harbor method correctly applied with appropriate removals for voice biometric data
consistencyStrong alignment: human subjects research → IRB approval → consent → deidentification → privacy protections → vulnerable populations
collection_notifications
collection_notifications:
- id: voice:notification:1
  description: 'Participants notified and consented through IRB-approved consent process before data collection.
    Consent process described eligibility, voluntary participation, data uses, sharing mechanisms, privacy
    protections, and withdrawal rights. English language used for all communications in current releases.

    '
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Informed Consent Obtained from Participants
evidencecollection_consents: All participants duly informed and provided prospective informed consent via IRB-approved process; consent includes voice data collection authorization, questionnaire permission, EHR access, data sharing permission; collection_notifications: participants notified and consented before data collection
qualityDetailed consent documentation: prospective informed consent, IRB-approved process, specific authorizations (voice tasks, questionnaires, EHR access, data sharing), notification before collection, withdrawal rights communicated
semanticconsent_completeness: excellent; consent_type: prospective_informed; consistency_check: pass_human_subjects_true_with_consent_documented
ethical_reviews
ethical_reviews:
- id: voice:ethicalreview:1
  description: 'IRB review and approval obtained from University of South Florida Single IRB with subsite
    IRB approvals. Study reviewed and approved for human subjects research. Bioethics guidance integrated
    throughout study design and conduct by consortium bioethicists (Belisle-Pipon, Ravitsky). Ethics module
    develops new guidelines for consenting to voice data collection, voice data sharing, and utilization
    in context of voice AI technology. Project addresses ethical issues from voice data generation through
    clinical adoption and downstream health decisions.

    '
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Ethical Review Details Including Conflicts
evidenceethical_reviews: IRB review and approval from University of South Florida Single IRB with subsite approvals; bioethics guidance integrated throughout by consortium bioethicists (Belisle-Pipon, Ravitsky); ethics module developing guidelines for voice data consenting, sharing, and AI utilization; project addresses ethical issues from data generation through clinical adoption
qualityDetailed ethical review documentation: IRB approval process (Single IRB with subsites), bioethics team involvement (named ethicists: Belisle-Pipon, Ravitsky), ethics module developing consent/sharing/utilization guidelines, lifecycle ethics consideration (generation→adoption→downstream decisions); no conflicts of interest explicitly stated
semanticethics_completeness: excellent; conflicts_documented: not_explicitly_stated; issues: No explicit conflicts of interest statement - minor gap but not a scoring issue given comprehensive ethics documentation
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive - all human subjects protections documented
evidencehuman_subject_research: detailed with IRB approval (USF Single IRB), involves_human_subjects=true; ethical_reviews: bioethics integrated throughout; is_deidentified: HIPAA Safe Harbor detailed; participant_privacy: 8 layered protections including Certificate of Confidentiality; collection_consents: prospective informed consent; consent_revocations: withdrawal mechanism; participant_compensation: $40-120 via gift cards; at_risk_populations: pediatric protections; sensitive_elements: 5 categories documented
qualityExemplary comprehensive ethical documentation covering all protection areas with specific mechanisms and regulatory compliance
correctnessHIPAA Safe Harbor method correctly applied with appropriate removals for voice biometric data
consistencyStrong alignment: human subjects research → IRB approval → consent → deidentification → privacy protections → vulnerable populations
sensitive_elements
sensitive_elements:
- id: voice:sensitive:1
  description: 'Voice recordings considered biometric identifiers under HIPAA. Raw audio waveforms excluded
    from public release; available only through controlled access with DACO approval and institutional
    DTUA. Risks include voice re-identification, voice AI hacking, and illicit or unauthorized use of
    voice data.

    '
- id: voice:sensitive:2
  description: 'Electronic health record (EHR) data accessed with participant consent for gold standard
    validation of diagnoses and symptoms. Sensitive health information including diagnoses, symptoms,
    disease-specific clinical data, and multimodal health biomarkers.

    '
- id: voice:sensitive:3
  description: 'Dataset contains sensitive demographic information: racial and ethnic origins, sexual
    orientation, financial and socioeconomic status, and health data. All direct identifiers removed;
    indirect identifiers removed where creating significant re-identification risk.

    '
- id: voice:sensitive:4
  description: 'Free speech task transcriptions may contain potentially identifying information or external
    voices. All open-response audio features (spectrograms, MFCCs, mel spectrograms, transcriptions, EMAs,
    PPGs) removed from public feature-only dataset.

    '
- id: voice:sensitive:5
  description: 'Dataset covered under Certificate of Confidentiality protecting against compulsory legal
    demands such as court orders and subpoenas for identifying information or characteristics of research
    participants. Provides additional legal protections beyond standard de-identification.

    '
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Data Anomalies and Quality Issues Noted
evidencecleaning_strategies: audit protocol with missingness tables included with dataset, distribution and outlier checks performed, categorical response validation against schema, audio quality control metrics (silence amount, duration, speech-to-passage accuracy checks); sensitive_elements: transcriptions may contain potentially identifying information or external voices
qualityQuality issues and anomalies documented: missingness tables generated and included, outlier detection performed, categorical validation checks, audio QC metrics (silence, duration, accuracy), transcription risks identified (identifying info, external voices), audit protocol applied via b2aiprep/SenseLab
semanticquality_awareness: excellent; qc_metrics_defined: yes
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Sensitive Content and Warnings Provided
evidencesensitive_elements: 5 entries covering voice as biometric identifier (re-identification risks, voice AI hacking), EHR data sensitivity (diagnoses, clinical data), demographic sensitivity (racial/ethnic origins, sexual orientation, financial status, health data), free speech transcription risks (identifying info, external voices), Certificate of Confidentiality coverage
qualityComprehensive sensitive content documentation: 5 categories of sensitive elements (biometric voice data with re-identification/hacking risks, sensitive EHR/health info, demographic PII, transcription risks, legal protections via Certificate of Confidentiality), specific risk types identified (re-identification, unauthorized use, voice AI hacking)
semanticsensitivity_awareness: excellent; risk_types_identified: 5
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive - all human subjects protections documented
evidencehuman_subject_research: detailed with IRB approval (USF Single IRB), involves_human_subjects=true; ethical_reviews: bioethics integrated throughout; is_deidentified: HIPAA Safe Harbor detailed; participant_privacy: 8 layered protections including Certificate of Confidentiality; collection_consents: prospective informed consent; consent_revocations: withdrawal mechanism; participant_compensation: $40-120 via gift cards; at_risk_populations: pediatric protections; sensitive_elements: 5 categories documented
qualityExemplary comprehensive ethical documentation covering all protection areas with specific mechanisms and regulatory compliance
correctnessHIPAA Safe Harbor method correctly applied with appropriate removals for voice biometric data
consistencyStrong alignment: human subjects research → IRB approval → consent → deidentification → privacy protections → vulnerable populations
is_deidentified
is_deidentified:
  id: voice:deident:1
  name: HIPAA Safe Harbor de-identification
  description: 'HIPAA Safe Harbor de-identification applied. All direct identifiers removed (names, civic
    addresses, social security numbers). Indirect identifiers removed where creating significant re-identification
    risk (geographic/demographic identifiers, household composition, cultural identity). Non-identifying
    sensitive information removed (household income, mental health status, traumatic life experiences).
    All raw voice data removed from public release. Sensitive fields per REDCap data dictionary removed.
    Direct identifiers removed: Yes. HIPAA de-identification rules applied: Yes. Dates rebased: Yes. Geographic
    information removed/generalized: Yes. Narrative text fields removed: Yes. K-anonymization: No.

    '
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Deidentification Method Described
evidenceis_deidentified: HIPAA Safe Harbor de-identification applied; all direct identifiers removed (names, civic addresses, SSN); indirect identifiers removed where creating significant re-identification risk; dates rebased; geographic info removed/generalized; narrative text fields removed
qualitySpecific deidentification method (HIPAA Safe Harbor) with detailed implementation: direct identifiers removed, indirect identifiers assessed and removed where risky, dates rebased, geographic generalization, narrative text removal - comprehensive approach appropriate for health data
semanticmethod_specificity: excellent; appropriateness: HIPAA_Safe_Harbor_suitable_for_voice_biomarker_health_data; consistency_check: pass_deidentified_with_method_specified
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive - all human subjects protections documented
evidencehuman_subject_research: detailed with IRB approval (USF Single IRB), involves_human_subjects=true; ethical_reviews: bioethics integrated throughout; is_deidentified: HIPAA Safe Harbor detailed; participant_privacy: 8 layered protections including Certificate of Confidentiality; collection_consents: prospective informed consent; consent_revocations: withdrawal mechanism; participant_compensation: $40-120 via gift cards; at_risk_populations: pediatric protections; sensitive_elements: 5 categories documented
qualityExemplary comprehensive ethical documentation covering all protection areas with specific mechanisms and regulatory compliance
correctnessHIPAA Safe Harbor method correctly applied with appropriate removals for voice biometric data
consistencyStrong alignment: human subjects research → IRB approval → consent → deidentification → privacy protections → vulnerable populations
participant_privacy
participant_privacy:
- id: voice:privacy:1
  description: 'Multiple layered privacy protections: (1) HIPAA Safe Harbor de-identification; (2) removal
    of all raw audio from public dataset; (3) removal of open-response features from public dataset; (4)
    Certificate of Confidentiality; (5) registered access requiring DUA signature; (6) controlled access
    for raw audio requiring institutional DTUA; (7) federated learning technology planned for multi-institutional
    analysis minimizing data sharing; (8) secure storage requirements in DTUA (administrative, physical,
    and technical safeguards). No analysis of potential re-identification impact conducted.

    '
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Privacy Protections Beyond Deidentification
evidenceparticipant_privacy: 8 layered protections including HIPAA Safe Harbor, raw audio removal from public dataset, open-response feature removal, Certificate of Confidentiality, registered access with DUA, controlled access for raw audio with institutional DTUA, federated learning planned, secure storage requirements
qualityExceptional multi-layered privacy framework: de-identification + data access controls (registered/controlled) + legal protections (Certificate of Confidentiality) + technical safeguards (federated learning, secure storage) + content filtering (raw audio/open-response removal)
semanticprivacy_depth: exceptional_8_layers; legal_protections: Certificate_of_Confidentiality
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive - all human subjects protections documented
evidencehuman_subject_research: detailed with IRB approval (USF Single IRB), involves_human_subjects=true; ethical_reviews: bioethics integrated throughout; is_deidentified: HIPAA Safe Harbor detailed; participant_privacy: 8 layered protections including Certificate of Confidentiality; collection_consents: prospective informed consent; consent_revocations: withdrawal mechanism; participant_compensation: $40-120 via gift cards; at_risk_populations: pediatric protections; sensitive_elements: 5 categories documented
qualityExemplary comprehensive ethical documentation covering all protection areas with specific mechanisms and regulatory compliance
correctnessHIPAA Safe Harbor method correctly applied with appropriate removals for voice biometric data
consistencyStrong alignment: human subjects research → IRB approval → consent → deidentification → privacy protections → vulnerable populations
at_risk_populations
at_risk_populations:
  id: voice:atrisk:1
  name: Pediatric participants protection
  description: 'Pediatric participants (aged 2-18) require additional protections. Pediatric data collected
    exclusively at Hospital for Sick Children (SickKids) with age-appropriate protocols. Distinct IRB
    considerations for pediatric cohort. Pediatric questionnaires and acoustic tasks specifically designed
    for different age groups (2-4, 4-6, 6-10, 10+ years). Minimum age for adult dataset is 18 years. Pediatric
    dataset released separately with additional privacy precautions.

    '
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Vulnerable Populations and Compensation Documented
evidencespecial_populations: Pediatric participants (aged 2-18) with age-appropriate protocols; at_risk_populations: Pediatric protections including exclusive SickKids collection, distinct IRB, age-grouped protocols, separate dataset release; participant_compensation: $40 for <90min sessions, $80 for >90min, max $120 total
qualityComprehensive vulnerable population protections: pediatric cohort identified as special population, age-appropriate protocols (2-4, 4-6, 6-10, 10+ years), dedicated site (SickKids), distinct IRB considerations, separate release with additional privacy precautions; compensation clearly documented
semanticvulnerable_pop_protections: excellent; compensation_clarity: clear
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive - all human subjects protections documented
evidencehuman_subject_research: detailed with IRB approval (USF Single IRB), involves_human_subjects=true; ethical_reviews: bioethics integrated throughout; is_deidentified: HIPAA Safe Harbor detailed; participant_privacy: 8 layered protections including Certificate of Confidentiality; collection_consents: prospective informed consent; consent_revocations: withdrawal mechanism; participant_compensation: $40-120 via gift cards; at_risk_populations: pediatric protections; sensitive_elements: 5 categories documented
qualityExemplary comprehensive ethical documentation covering all protection areas with specific mechanisms and regulatory compliance
correctnessHIPAA Safe Harbor method correctly applied with appropriate removals for voice biometric data
consistencyStrong alignment: human subjects research → IRB approval → consent → deidentification → privacy protections → vulnerable populations
participant_compensation
participant_compensation:
- id: voice:compensation:1
  description: 'Adult participants compensated via electronic gift cards: $40 for data collection sessions
    under 90 minutes; $80 for sessions over 90 minutes. Maximum 3 sessions and maximum total compensation
    of $120 per participant. Research staff (clinicians, co-investigators) compensated through consortium-level
    publication authorship.

    '
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Vulnerable Populations and Compensation Documented
evidencespecial_populations: Pediatric participants (aged 2-18) with age-appropriate protocols; at_risk_populations: Pediatric protections including exclusive SickKids collection, distinct IRB, age-grouped protocols, separate dataset release; participant_compensation: $40 for <90min sessions, $80 for >90min, max $120 total
qualityComprehensive vulnerable population protections: pediatric cohort identified as special population, age-appropriate protocols (2-4, 4-6, 6-10, 10+ years), dedicated site (SickKids), distinct IRB considerations, separate release with additional privacy precautions; compensation clearly documented
semanticvulnerable_pop_protections: excellent; compensation_clarity: clear
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Creators and Acknowledgements Documented
evidencecreators: Bridge2AI-Voice Consortium with 50+ investigators from 12+ institutions; Contact PI: Dr. Yael Bensoussan (USF); Co-PI: Dr. Olivier Elemento (Weill Cornell); 14 key co-investigators named with institutional affiliations; participant_compensation documented; research staff authorship acknowledgment
qualityComprehensive creator documentation: consortium structure (50+ investigators, 12+ institutions), leadership (Contact PI, Co-PI), 14 key co-investigators with names and affiliations, participant compensation ($40-120), research staff acknowledgment (authorship)
semanticcreator_completeness: excellent; institutional_attribution: yes
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive - all human subjects protections documented
evidencehuman_subject_research: detailed with IRB approval (USF Single IRB), involves_human_subjects=true; ethical_reviews: bioethics integrated throughout; is_deidentified: HIPAA Safe Harbor detailed; participant_privacy: 8 layered protections including Certificate of Confidentiality; collection_consents: prospective informed consent; consent_revocations: withdrawal mechanism; participant_compensation: $40-120 via gift cards; at_risk_populations: pediatric protections; sensitive_elements: 5 categories documented
qualityExemplary comprehensive ethical documentation covering all protection areas with specific mechanisms and regulatory compliance
correctnessHIPAA Safe Harbor method correctly applied with appropriate removals for voice biometric data
consistencyStrong alignment: human subjects research → IRB approval → consent → deidentification → privacy protections → vulnerable populations
external_resources
external_resources:
- id: voice:resource:1
  name: PhysioNet Dataset Landing Page
  description: Primary registered access distribution platform for adult and pediatric datasets
  external_resources:
  - https://physionet.org/content/b2ai-voice/
- id: voice:resource:2
  name: Bridge2AI-Voice Project Documentation
  description: Comprehensive project documentation, collection methods, governance, healthsheet
  external_resources:
  - https://docs.b2ai-voice.org
- id: voice:resource:3
  name: Bridge2AI-Voice GitHub Documentation Repository
  description: Source code for docs and dashboard at docs.b2ai-voice.org (MIT license)
  external_resources:
  - https://github.com/eipm/bridge2ai-docs
- id: voice:resource:4
  name: b2aiprep Software Library
  description: 'Open source library (Apache-2.0 license) for preprocessing raw audio waveforms into parquet
    files and merging source data into phenotype files

    '
  external_resources:
  - https://github.com/sensein/b2aiprep
- id: voice:resource:5
  name: Bridge2AI REDCap Data Dictionary
  description: REDCap data dictionary and metadata (MIT license) at github.com/eipm/bridge2ai-redcap
  external_resources:
  - https://github.com/eipm/bridge2ai-redcap
- id: voice:resource:6
  name: NIH RePORTER Project Details
  description: Federal grant information for project 3OT2OD032720-01S3
  external_resources:
  - https://reporter.nih.gov/project-details/11376382
- id: voice:resource:7
  name: Health Data Nexus
  description: Alternative platform providing cloud compute alongside earlier dataset versions
  external_resources:
  - https://healthdatanexus.ai/content/b2ai-voice/1.0/
- id: voice:resource:8
  name: Zenodo Archive (REDCap data dictionary)
  description: 'Bensoussan Y. et al. (2024). eipm/bridge2ai-redcap. Zenodo. https://zenodo.org/doi/10.5281/zenodo.12760724

    '
  external_resources:
  - https://doi.org/10.5281/zenodo.13834653
- id: voice:resource:9
  name: Interspeech 2024 Protocol Publication
  description: 'Bensoussan et al. "Developing Multi-Disorder Voice Protocols: A team science approach
    involving clinical expertise, bioethics, standards, and DEI." Proc. Interspeech 2024.

    '
  external_resources:
  - https://doi.org/10.21437/Interspeech.2024-1926
- id: voice:resource:10
  name: Data Access Compliance Office (DACO)
  description: Contact for controlled access to raw audio data and institutional DTUA
  external_resources:
  - mailto:DACO@b2ai-voice.org
- id: voice:resource:11
  name: Bridge2AI Program
  description: Parent NIH Common Fund program supporting AI-ready biomedical datasets
  external_resources:
  - https://bridge2ai.org
- id: voice:resource:12
  name: Bridge2AI Voice Scholars Training
  description: Training opportunities for using the dataset
  external_resources:
  - https://www.b2aivoicescholars.org/
- id: voice:resource:13
  name: SenseLab Software
  description: Software toolkit for processing audio files for research tasks
  external_resources:
  - https://github.com/sensein/senselab
- id: voice:resource:14
  name: AI-Readiness Recommendations
  description: AI-readiness for Biomedical Data - Bridge2AI Recommendations document
  external_resources:
  - https://docs.b2ai-voice.org
⚠ low R20 · consistency
issueCitation provided but no additional external publications referenced despite multi-year project
fieldscitation, external_resources
fixConsider adding additional project publications if available
✓ 1/1 R10 1.Dataset Discovery and Identification Landing Page and Resources
evidencepage: https://docs.b2ai-voice.org; external_resources includes PhysioNet landing page, GitHub repos, documentation sites
qualityMultiple accessible landing pages documented: project docs (docs.b2ai-voice.org), PhysioNet distribution (physionet.org/content/b2ai-voice/), Health Data Nexus platform, plus 14 external resource entries
semanticaccessibility: multiple_platforms; url_validity: valid; issues: page field points to docs site rather than primary DOI landing page (PhysioNet) - minor inconsistency but both are valid
✓ 1/1 R10 10.Cross-Platform and Community Integration Outreach Materials and Documentation Links
evidenceexternal_resources: Project documentation site (docs.b2ai-voice.org), GitHub documentation repository, training program (B2AI Voice Scholars at b2aivoicescholars.org), software libraries (b2aiprep, SenseLab on GitHub), REDCap data dictionary (GitHub + Zenodo), PhysioNet landing page, publications (Interspeech 2024), NIH RePORTER project details
qualityComprehensive outreach and documentation: dedicated documentation site, training program (B2AI Voice Scholars), GitHub repositories (4 repos for software and documentation), published protocols (Interspeech 2024), data dictionary (public on GitHub and Zenodo), platform landing pages (PhysioNet, Health Data Nexus), funding transparency (NIH RePORTER)
semanticoutreach_completeness: excellent; educational_resources: yes; documentation_accessibility: public
✓ 1/1 R10 2.Dataset Access and Retrieval Download URL or Platform Link Available
evidenceexternal_resources: PhysioNet (https://physionet.org/content/b2ai-voice/), Health Data Nexus (https://healthdatanexus.ai/content/b2ai-voice/1.0/), controlled access via DACO@b2ai-voice.org
qualityMultiple download platforms documented with specific URLs: PhysioNet (primary registered access), Health Data Nexus (alternative with cloud compute), controlled access contact (DACO email) for raw audio
semanticplatform_availability: multiple_platforms; url_validity: valid
✓ 1/1 R10 2.Dataset Access and Retrieval Related Datasets and External Resources Linked
evidenceexternal_resources: 14 entries including PhysioNet, GitHub repos (b2aiprep, bridge2ai-redcap, bridge2ai-docs, SenseLab), Zenodo archive, NIH RePORTER, publications (Interspeech 2024), Health Data Nexus
qualityComprehensive external resources (14 distinct entries) with software repositories (GitHub), publications (DOI links), platform links, funding details (NIH RePORTER), and alternative distribution platforms
semanticresource_count: 14; link_types: diverse_comprehensive
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) External Standards and Resources Referenced
evidenceexternal_resources: 14 entries including publications (Interspeech 2024 DOI), standards (BIDS v1.9.0, FAIR, CARE, ICD-10, HIPAA Safe Harbor), platforms (PhysioNet, Health Data Nexus), repositories (GitHub, Zenodo), documentation (docs.b2ai-voice.org), funding (NIH RePORTER)
qualityExtensive external documentation: published protocols (Interspeech 2024 with DOI), data standards (BIDS v1.9.0, FAIR, CARE), coding systems (ICD-10), regulatory standards (HIPAA), platform documentation (PhysioNet, Health Data Nexus), software repos (4 GitHub links), funding details (NIH RePORTER), archive (Zenodo DOI)
semanticexternal_resource_count: 14; standard_compliance: multiple
5/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelComprehensive strategies with software versions/URLs
evidencepreprocessing_strategies: 8 detailed entries with software tools (b2aiprep, OpenSMILE eGeMaps, Parselmouth, Praat, torchaudio, SPARC, OpenAI Whisper, REDCap, reproschema-ui); cleaning_strategies: 4 entries; labeling_strategies: 1 entry with validation methods; external_resources: GitHub links to b2aiprep (Apache-2.0), bridge2ai-redcap (MIT), SenseLab; specific parameters documented (16 kHz resampling, 25ms FFT window, 60 MFCCs)
qualityExceptional software transparency with specific tools, processing parameters, open-source repository links, and licensing information
correctnessSoftware tools appropriate for acoustic feature extraction and clinical data management; processing parameters (16kHz, 60 MFCCs) standard for speech analysis
consistencyTools align with described features (OpenSMILE→eGeMaps, Whisper→transcripts, b2aiprep→BIDS conversion)
4/5 R20 Q14 (Technical Documentation) Associated Publications
levelMultiple references and dataset citation
evidencecitation: Bensoussan et al. Interspeech 2024 (https://doi.org/10.21437/Interspeech.2024-1926); external_resources: Zenodo archive (10.5281/zenodo.13834653), NIH RePORTER project link; missing: additional methodological papers or dataset descriptor publications
qualityGood publication documentation with formal citation and DOIs; could benefit from additional methodological papers given project scope
correctnessDOIs valid format for Interspeech conference and Zenodo archive
consistencyCitation describes protocol development aligning with collection methods; publication date (2024) consistent with data collection timeline
1/1 R20 Q16 (FAIRness & Accessibility) Findability (Persistent Links)
levelPass - Multiple persistent URLs present
evidencepage: https://docs.b2ai-voice.org; external_resources: 14 URLs including https://physionet.org/content/b2ai-voice/, https://healthdatanexus.ai/content/b2ai-voice/1.0/, GitHub repos, DOIs (10.5281/zenodo.13834653, 10.21437/Interspeech.2024-1926), NIH RePORTER
qualityExcellent findability with multiple persistent access points across platforms
correctnessAll URLs follow standard patterns and include appropriate domains (.org, .ai, github.com, doi.org)
consistencyURLs align with described platforms (PhysioNet, Health Data Nexus, GitHub, Zenodo)
1/1 R20 Q20 (FAIRness & Accessibility) Interlinking Across Platforms
levelPass - Cross-platform links verified
evidenceexternal_resources: 14 resources linking PhysioNet (primary distribution), Health Data Nexus (alternative platform with cloud compute), GitHub (bridge2ai-docs, b2aiprep, bridge2ai-redcap, senselab), Zenodo (archive), NIH RePORTER (grant info), doi.org (publications), Bridge2AI program site, training site (b2aivoicescholars.org)
qualityExceptional cross-platform interlinking connecting data repositories, code, documentation, publications, and training resources
correctnessPlatform links appropriate for roles: PhysioNet/Nexus for data, GitHub for code, Zenodo for archival, NIH RePORTER for funding
consistencyCross-platform references align with described distribution (PhysioNet primary), software (GitHub repos), and infrastructure (Health Data Nexus for compute)
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass - DOI and persistent URLs present
evidencedoi: 10.13026/37yb-1t42 (PhysioNet), page: https://docs.b2ai-voice.org, external_resources: 14 persistent URLs including PhysioNet, GitHub repos, Zenodo
qualityExcellent identifier coverage with PhysioNet DOI and multiple persistent access points
correctnessDOI prefix 10.13026 correctly identifies PhysioNet registrar
consistencyDOI matches page URL domain and distribution platform

Unmatched feedback

Feedback that referenced multiple fields or no specific field.
✓ 1/1 R10 2.Dataset Access and Retrieval Regulatory Restrictions and Confidentiality Level Specified
evidenceregulatory_compliance: HIPAA Safe Harbor de-identification, Certificate of Confidentiality, 45 CFR 46 (Common Rule), PII per OMB M-07-16; no export controls; Certificate of Confidentiality protects against legal demands
qualityExplicit regulatory framework documented: HIPAA compliance, Common Rule (45 CFR 46), Certificate of Confidentiality coverage, PII classification, no export controls - comprehensive regulatory compliance statements
semanticregulatory_completeness: excellent; confidentiality_level: clearly_defined
✓ 1/1 R10 3.Data Reuse and Interoperability Data Formats Are Standardized
evidenceformat: Parquet (Python datasets library compatible), TSV/JSON (tab-delimited with JSON data dictionaries), WAV audio; encoding: BIDS v1.9.0 compliant structure; monaural 16 kHz audio after preprocessing
qualityStandard formats throughout: Parquet (Apache standard), TSV/JSON (universal), WAV audio (standard audio format), BIDS v1.9.0 compliance (community neuroimaging standard), 16 kHz sampling rate specified
semanticformat_standardization: excellent; encoding_specified: yes
✓ 1/1 R10 3.Data Reuse and Interoperability Variable Metadata with Identifiers Defined
evidenceStatic features: static_features.tsv with static_features.json data dictionary; Phenotype files: confounders.tsv, demographics.tsv, diagnosis/*.tsv, questionnaire/*.tsv with JSON data dictionaries per file; REDCap data dictionary available on GitHub
qualityComprehensive variable-level metadata: JSON data dictionaries for all TSV files, OpenSMILE/Praat/Parselmouth/torchaudio features documented, REDCap data dictionary published on GitHub with MIT license
semanticmetadata_completeness: excellent; data_dictionary_availability: public_github
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Software and Tools Documented
evidenceSoftware tools: b2aiprep (preprocessing library, Apache-2.0), SenseLab (audio processing), OpenSMILE (eGeMaps features), Parselmouth/Praat (phonetic features), torchaudio (pitch/spectrograms), OpenAI Whisper (transcription), REDCap (data collection), reproschema-ui (pediatric protocol), Bridge2AI-Voice App (iOS data collection); GitHub repos linked for b2aiprep, SenseLab, bridge2ai-redcap, bridge2ai-docs
qualityComprehensive software documentation: 9+ tools named with purposes (b2aiprep, SenseLab, OpenSMILE, Parselmouth, Praat, torchaudio, Whisper, REDCap, reproschema-ui, Bridge2AI-Voice App), licensing (Apache-2.0, MIT), GitHub repository links (4 repos), version specified for BIDS (v1.9.0)
semantictool_completeness: excellent; repository_links: yes
✓ 1/1 R10 10.Cross-Platform and Community Integration Dataset Published on a Recognized Platform
evidencePhysioNet (primary distribution, MIT Laboratory for Computational Physiology, NIH NIBIB supported via R01EB030362); Health Data Nexus (T-CAIREM, University of Toronto); datasets distributed broadly through registered and controlled access mechanisms
qualityMultiple recognized platforms: PhysioNet (primary, NIH-supported repository for physiological data), Health Data Nexus (T-CAIREM/University of Toronto cloud compute platform), both well-established biomedical data repositories
semanticplatform_recognition: excellent_multiple_recognized_platforms; platform_type: domain_specific_biomedical
5/5 R20 Q10 (Metadata Quality & Content) Interoperability and Standardization
levelStandard formats + schema/ontology compliance
evidenceformat: Parquet, TSV, JSON, WAV; conforms_to: BIDS v1.9.0 structure explicitly specified; standards mentioned: ICD-10 codes for diagnoses, HIPAA Safe Harbor, FAIR principles, CARE principles; preprocessing uses standardized features (OpenSMILE eGeMaps, Praat, MFCCs)
qualityExcellent interoperability with explicit BIDS conformance, standard acoustic features, clinical coding (ICD-10), and FAIR/CARE alignment
correctnessBIDS v1.9.0 appropriate for neuroimaging/acoustic data; standard feature extraction methods correctly cited
consistencyFormat choices (Parquet for arrays, TSV/JSON for tabular) align with BIDS conventions and data types

Recommendations

  1. R10 · Consider aligning 'page' field with DOI landing page (PhysioNet) for consistency, or clarify that docs.b2ai-voice.org is intentionally the primary project page
  2. R10 · Add explicit conflicts of interest statement to ethical_reviews section to complement comprehensive ethics documentation
  3. R10 · Document any reidentification risk analysis or privacy impact assessment if conducted (mentioned as not performed in participant_privacy)
  4. R10 · Consider adding more granular missingness statistics from audit protocol as part of data quality documentation
  5. R10 · Document future Spanish language protocol timeline and implementation details when available
  6. R10 · Add cross-references to multimodal data (imaging, genomics) planned for future releases with expected timelines
  7. R20 · Add dataset descriptor publication to citation field and external_resources (e.g., Data in Brief or Scientific Data article describing v3.0.0)
  8. R20 · Include RRIDs for key software resources: b2aiprep, OpenSMILE, Praat, reproschema-ui to improve research resource tracking and citation
  9. R20 · Create structured errata/change log documentation in updates section beyond version release notes
  10. R20 · Consider adding additional methodological papers covering bioethics framework, BIDS adaptation for voice, or federated learning implementation when available
  11. R20 · Expand version_access to explicitly mention errata mechanisms if issues discovered in specific versions
  12. R20 · Add clarifying note that grant number 3OT2OD032720-01S3 represents supplemental funding (prefix '3', suffix 'S3') to NIH base award