Interleaved Semantic Evaluation

Project: VOICE · Method: claudecode_agent_core
YAML: data/d4d_concatenated/claudecode_agent_core/VOICE_d4d_core.yaml
R10 JSON: data/evaluation_llm/rubric10_semantic/concatenated/VOICE_claudecode_agent_core_evaluation.json
R20 JSON: data/evaluation_llm/rubric20_semantic/concatenated/VOICE_claudecode_agent_core_evaluation.json
Model: claude-sonnet-4-5-20250929
Rubric10 (semantic)
42/50 (84.0%)
Rubric20 (semantic)
68.0/84 (81.0%)
Consistency checks (R10/R20)
34 pass · 0 fail · 4 warn
Mapped feedback / fields
165 across 51 fields
R10 sub-element R20 question Semantic issue

Strengths

  • Exceptional structural completeness with all mandatory fields populated at high quality (73 keywords, 1148-char description, comprehensive purposes/tasks/gaps)
  • Outstanding ethical documentation covering IRB approval, HIPAA Safe Harbor deidentification, informed consent, Certificate of Confidentiality, and pediatric protections
  • Exemplary version history with version-specific DOIs, timestamped releases, participant count tracking, and clear update roadmap (v1.0→v3.0 progression documented)
  • Comprehensive collection protocol detail with specific hardware (iPad + Avid AE-36 mic), 22 adult + 36 pediatric acoustic tasks, validated questionnaires, and multi-site coordination
  • Strong FAIR compliance with PhysioNet DOI, persistent URLs, BIDS v1.9.0 standard, multi-platform distribution (PhysioNet, Health Data Nexus, GitHub), and detailed access mechanisms
  • Rich funding documentation with NIH grant number, funding amounts, project dates, and detailed multi-institutional consortium structure (50+ investigators, 12+ institutions)
  • Detailed subpopulation characterization with 6 disease cohorts, clinical validation methods, and age-appropriate protocols for pediatric population
  • Excellent technical documentation of preprocessing (8 strategies with specific parameters: 16kHz, 25ms window, 60 MFCCs) and cleaning (HIPAA Safe Harbor, audio privacy tiers)
  • Well-defined two-tier access governance: registered access (PhysioNet DUA) for de-identified features, controlled access (DACO + institutional DTUA) for raw audio biometrics
  • Comprehensive external resources with 12 links including 2 DOI-linked publications, 3 GitHub repos with licenses, Zenodo archive, and multi-platform documentation

Weaknesses

  • D4D-core schema limitation: lacks dedicated conforms_to_schema field present in full D4D schema (BIDS compliance mentioned in descriptions but not explicit structured field)
  • Software tool documentation incomplete in core schema: tool names mentioned (b2aiprep, OpenSMILE, Whisper) but versions/release dates not in dedicated structured fields (though GitHub links in external_resources partially compensate)
  • Minor grant number format question: '3OT2OD032720-01S3' follows NIH supplemental pattern but leading '3' unusual (likely correct as supplement indicator)
  • No RRID identifiers present for software tools (OpenSMILE, Praat, Whisper could have RRIDs for better tool citation)
  • Citation field absent in D4D-core schema: dataset citation not in dedicated structured field (dataset DOI present but no BibTeX/RIS export mentioned)

Field-by-field

id
id: https://doi.org/10.13026/37yb-1t42
✓ 1/1 R10 1.Dataset Discovery and Identification Persistent Identifier (DOI, RRID, or URI)
evidencedoi: 10.13026/37yb-1t42, id: https://doi.org/10.13026/37yb-1t42
qualityDOI present with correct PhysioNet prefix (10.13026), properly formatted, and semantically valid
semanticDOI prefix 10.13026 matches PhysioNet registrar; full DOI resolves to valid dataset landing page
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice, description: 1148 chars comprehensive, keywords: 73 keywords, license_and_use_terms: detailed registered access license
qualityAll mandatory fields present with exceptional completeness and detail
correctnessDOI format valid, PhysioNet prefix 10.13026 verified
consistencyAll mandatory fields semantically coherent and appropriate for voice biomarker dataset
name
name: Bridge2AI-Voice
no field-level feedback matched
title
title: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information
✓ 1/1 R10 1.Dataset Discovery and Identification Dataset Title and Description Completeness
evidencetitle: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information; description: 1,157 characters providing specific details on 833 participants, 61,937 recordings, disease categories, collection methods
qualityComprehensive description exceeds 200 characters, provides quantitative specifics, explains scientific rationale, describes data types and access model
semanticDescription semantically rich with specific participant counts, data modalities, ethical considerations, and research objectives
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice, description: 1148 chars comprehensive, keywords: 73 keywords, license_and_use_terms: detailed registered access license
qualityAll mandatory fields present with exceptional completeness and detail
correctnessDOI format valid, PhysioNet prefix 10.13026 verified
consistencyAll mandatory fields semantically coherent and appropriate for voice biomarker dataset
description
description: 'The Bridge2AI-Voice project seeks to create an ethically sourced flagship dataset to enable
  future research in artificial intelligence and support critical insights into the use of voice as a
  biomarker of health. The human voice contains complex acoustic markers which have been linked to important
  health conditions including dementia, mood disorders, and cancer. When viewed as a biomarker, voice
  is a promising characteristic to measure as it is simple to collect, cost-effective, and has broad clinical
  utility. This comprehensive collection provides voice recordings with corresponding clinical information
  from participants selected based on known conditions which manifest within the voice waveform including
  voice disorders, neurological disorders, mood disorders, and respiratory disorders. The dataset is designed
  to fuel voice AI research, establish data standards, and promote ethical and trustworthy AI/ML development
  for voice biomarkers of health. Data collection occurs through a multi-institutional collaborative effort
  using standardized protocols, custom smartphone applications, and rigorous ethical oversight. Version
  3.0 provides approximately 61,937 voice-derived recordings from 833 adult participants collected across
  multiple sites in North America, with derived features such as spectrograms, MFCCs, acoustic features,
  and clinical phenotype data. The pediatric dataset v1.0 is also available with data from 300 participants.
  Raw audio data is available through controlled access to protect participant privacy.

  '
✓ 1/1 R10 1.Dataset Discovery and Identification Dataset Title and Description Completeness
evidencetitle: Bridge2AI-Voice - An ethically-sourced, diverse voice dataset linked to health information; description: 1,157 characters providing specific details on 833 participants, 61,937 recordings, disease categories, collection methods
qualityComprehensive description exceeds 200 characters, provides quantitative specifics, explains scientific rationale, describes data types and access model
semanticDescription semantically rich with specific participant counts, data modalities, ethical considerations, and research objectives
✓ 1/1 R10 5.Data Composition and Structure Number of Instances or Samples Reported
evidenceinstances.counts: 833 adult participants; description specifies ~61,937 voice-derived recordings; pediatric v1.0 adds 300 participants; enrollment target 10,000 by 2027
qualitySpecific instance counts: 833 adult participants, ~61,937 recordings, 300 pediatric participants; progression documented (v1.0: 306 participants, 12,523 recordings → v3.0: 833, ~61,937)
semanticInstance counts semantically detailed and versioned - current state (v3.0: 833), historical (v1.0: 306), future target (10,000), recordings-per-participant ratio provided
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice, description: 1148 chars comprehensive, keywords: 73 keywords, license_and_use_terms: detailed registered access license
qualityAll mandatory fields present with exceptional completeness and detail
correctnessDOI format valid, PhysioNet prefix 10.13026 verified
consistencyAll mandatory fields semantically coherent and appropriate for voice biomarker dataset
4/5 R20 Q10 (Metadata Quality & Content) Interoperability and Standardization
levelStandard formats + partial schema compliance
evidenceformat: TSV, JSON, GZ, Parquet (standard formats); encoding: implicit UTF-8 for text; description mentions 'BIDS v1.9.0 compliant structure' for directory organization and TSV/JSON phenotype files; keywords include BIDS, FAIR principles, CARE principles
qualityStrong standardization with BIDS v1.9.0 compliance documented in description/distribution fields, though D4D-core schema lacks explicit conforms_to_schema field present in full schema
correctnessBIDS v1.9.0 is appropriate standard for neuroimaging/voice biomarker data; Parquet and TSV formats semantically correct for time-series and tabular phenotype data
consistencyBIDS compliance claim consistent across multiple distribution format descriptions and phenotype data organization
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1148 chars with specific participant counts and version details, purposes: 4 entries averaging 350+ chars each with detailed research objectives
qualityExceptional narrative depth with concrete metrics and comprehensive research goals
correctnessParticipant numbers (833 adults, 61,937 recordings) semantically consistent with v3.0 release scope
consistencyPurposes align with NIH Bridge2AI program mission and funding goals
doi
doi: 10.13026/37yb-1t42
✓ 1/1 R10 1.Dataset Discovery and Identification Persistent Identifier (DOI, RRID, or URI)
evidencedoi: 10.13026/37yb-1t42, id: https://doi.org/10.13026/37yb-1t42
qualityDOI present with correct PhysioNet prefix (10.13026), properly formatted, and semantically valid
semanticDOI prefix 10.13026 matches PhysioNet registrar; full DOI resolves to valid dataset landing page
✓ 1/1 R10 10.Cross-Platform and Community Integration Citation and DOI for Cross-referencing
evidencedoi: 10.13026/37yb-1t42; version_access documents version-specific DOIs for all releases; external_resources reference Interspeech 2024 publication (doi.org/10.21437/Interspeech.2024-1926), Zenodo REDCap archive (doi.org/10.5281/zenodo.13834653)
qualityDOI provided for dataset (10.13026/37yb-1t42), version-specific DOIs documented, related publications have DOIs (Interspeech 2024, Zenodo archive); no formal citation string but publisher and version enable construction
semanticCitation infrastructure semantically complete - dataset DOI enables permanent reference, version-specific DOIs enable precise replication, related resource DOIs support comprehensive citation
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) External Standards and Resources Referenced
evidenceexternal_resources: 12 documented including BIDS v1.9.0 specification (referenced 6+ times), Interspeech 2024 protocol publication (doi.org/10.21437/Interspeech.2024-1926), NIH RePORTER grant details, Zenodo REDCap archive (doi.org/10.5281/zenodo.13834653), Bridge2AI program site
qualityExternal standards explicitly cited: BIDS v1.9.0 (data structure), ICD-10 (diagnoses), Bridge2AI protocols; publications provide methodological validation (Interspeech 2024); code/documentation archived (Zenodo, GitHub)
semanticExternal resources semantically comprehensive - standards (BIDS), publications (Interspeech), repositories (GitHub, Zenodo), program context (Bridge2AI), grant information (NIH RePORTER)
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice, description: 1148 chars comprehensive, keywords: 73 keywords, license_and_use_terms: detailed registered access license
qualityAll mandatory fields present with exceptional completeness and detail
correctnessDOI format valid, PhysioNet prefix 10.13026 verified
consistencyAll mandatory fields semantically coherent and appropriate for voice biomarker dataset
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass
evidencedoi: 10.13026/37yb-1t42, page: https://docs.b2ai-voice.org, download_url: https://physionet.org/content/b2ai-voice/
qualityMultiple persistent identifiers present with PhysioNet DOI and project documentation URLs
correctnessDOI prefix 10.13026 matches PhysioNet registrar, DOI format valid
consistencyDOI, page URL, and download URL all resolve to consistent Bridge2AI-Voice project resources
page
page: https://docs.b2ai-voice.org
✓ 1/1 R10 1.Dataset Discovery and Identification Landing Page and Resources (page, hierarchical resources)
evidencepage: https://docs.b2ai-voice.org; download_url: https://physionet.org/content/b2ai-voice/; external_resources: 12 documented resources including PhysioNet, GitHub repos, documentation
qualityMultiple accessible landing pages (project docs and PhysioNet), hierarchical resources structure with 12 external resources providing documentation, code, and access points
semanticURLs semantically appropriate - docs.b2ai-voice.org for documentation, physionet.org for data access, github.com for code repositories
1/1 R20 Q16 (FAIRness & Accessibility) Findability (Persistent Links)
levelPass
evidencepage: https://docs.b2ai-voice.org, download_url: https://physionet.org/content/b2ai-voice/, external_resources: 12 entries with URLs to PhysioNet, GitHub (bridge2ai-docs, b2aiprep, bridge2ai-redcap), NIH RePORTER, Health Data Nexus, Zenodo, Interspeech DOI, Bridge2AI program, training site
qualityExcellent persistent link coverage with documentation, download, and external resource URLs across multiple platforms
correctnessAll URLs follow valid format with appropriate domains (physionet.org, github.com, doi.org, nih.gov, zenodo.org)
consistencyURLs consistent with described distribution platforms (PhysioNet primary, Health Data Nexus alternative)
1/1 R20 Q20 (FAIRness & Accessibility) Interlinking Across Platforms
levelPass
evidenceexternal_resources: cross-platform links verified - PhysioNet (primary distribution), Health Data Nexus (alternative with cloud compute), GitHub (3 repos: bridge2ai-docs, b2aiprep, bridge2ai-redcap), Zenodo (archive), NIH RePORTER (grant tracking), Bridge2AI program site, training portal; page: https://docs.b2ai-voice.org; download_url: https://physionet.org/content/b2ai-voice/
qualityExcellent cross-platform interlinking with 7+ distinct platforms (PhysioNet, Health Data Nexus, GitHub, Zenodo, NIH, Bridge2AI, training site) providing complementary access points
correctnessPlatform links semantically appropriate: PhysioNet for biomedical data distribution, GitHub for code/docs, Zenodo for archival, NIH RePORTER for grant transparency
consistencyCross-platform links consistent with described distribution strategy (PhysioNet primary, Health Data Nexus alternative, GitHub for open-source tools)
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass
evidencedoi: 10.13026/37yb-1t42, page: https://docs.b2ai-voice.org, download_url: https://physionet.org/content/b2ai-voice/
qualityMultiple persistent identifiers present with PhysioNet DOI and project documentation URLs
correctnessDOI prefix 10.13026 matches PhysioNet registrar, DOI format valid
consistencyDOI, page URL, and download URL all resolve to consistent Bridge2AI-Voice project resources
language
language: en
✓ 1/1 R10 3.Data Reuse and Interoperability Data Formats Are Standardized (encoding, format)
evidencedistribution_formats specify standard formats: Parquet (Python datasets library compatible), TSV (BIDS v1.9.0 compliant), JSON data dictionaries, WAV audio; language: en; BIDS v1.9.0 structure throughout
qualityStandards-compliant formats: BIDS v1.9.0 for directory structure and metadata, Parquet for ML-friendly time-series, TSV/JSON pairing for interoperability
semanticFormat choices semantically AI/ML-friendly - Parquet for efficient feature loading, BIDS for neuroimaging community compatibility
version
version: 3.0.0
✓ 1/1 R10 10.Cross-Platform and Community Integration Citation and DOI for Cross-referencing
evidencedoi: 10.13026/37yb-1t42; version_access documents version-specific DOIs for all releases; external_resources reference Interspeech 2024 publication (doi.org/10.21437/Interspeech.2024-1926), Zenodo REDCap archive (doi.org/10.5281/zenodo.13834653)
qualityDOI provided for dataset (10.13026/37yb-1t42), version-specific DOIs documented, related publications have DOIs (Interspeech 2024, Zenodo archive); no formal citation string but publisher and version enable construction
semanticCitation infrastructure semantically complete - dataset DOI enables permanent reference, version-specific DOIs enable precise replication, related resource DOIs support comprehensive citation
✓ 1/1 R10 6.Data Provenance and Version Tracking Dataset Version Number Provided
evidenceversion: 3.0.0; distribution_dates and updates document version progression: v1.0 (Jan 17, 2025, 306 participants), v1.1 (Jan 17, 2025, added MFCCs), v2.0.0 (Apr 16, 2025), v2.0.1 (Aug 18, 2025), v3.0.0 (2025, 833 participants)
qualityCurrent version (3.0.0) clearly stated, comprehensive version history with release dates and participant counts, semantic versioning pattern followed
semanticVersion numbering semantically meaningful - major version increments (1.0→2.0→3.0) correspond to significant participant count increases
✓ 1/1 R10 6.Data Provenance and Version Tracking Version Access Methods Documented
evidenceversion_access: all versions available at https://physionet.org/content/b2ai-voice/ with version-specific DOIs; earlier versions also at Health Data Nexus; DOI for latest version: 10.13026/37yb-1t42; older versions continue to be supported and hosted
qualityExplicit version access policy: PhysioNet hosts all versions with unique DOIs, Health Data Nexus provides alternative access, commitment to ongoing support of older versions
semanticVersion access semantically robust - unique DOIs enable permanent citation of specific versions, multiple platforms reduce single-point-of-failure risk
✓ 1/1 R10 6.Data Provenance and Version Tracking Change Descriptions and Errata Provided
evidenceupdates and distribution_dates document version changes: v1.0 (306 participants, 12,523 recordings), v1.1 (added MFCC features), v2.0.0, v2.0.1, v3.0.0 (833 adults, ~61,937 recordings); users notified through platform news items
qualityVersion change descriptions provided: participant count increases, feature additions (MFCCs in v1.1), notification mechanism (platform news items); pediatric v1.0 as separate release
semanticChange documentation semantically informative - quantitative details (participant counts, recording counts, feature types) enable impact assessment
5/5 R20 Q13 (Technical Documentation) Version History Documentation
levelComprehensive versioning with errata, updates, and release notes
evidenceversion: 3.0.0; version_access: all versions available via PhysioNet with version-specific DOIs, older versions on Health Data Nexus; updates: semi-annual releases, v1.0 Jan 17 2025 (306 participants, 12,523 recordings), v1.1 added MFCC features, v2.0.0 Apr 16 2025, v2.0.1 Aug 18 2025, v3.0.0 2025 (833 adults, ~61,937 recordings), pediatric v1.0 separate, target 10,000 by 2027; distribution_dates: detailed release timeline with participant counts per version
qualityExemplary version tracking with specific release dates, participant/recording counts per version, version-specific DOIs, and clear update roadmap
correctnessVersion progression plausible (306→833 participants over 1 year with semi-annual releases); version numbering follows semantic versioning (1.0→1.1→2.0.0→2.0.1→3.0.0)
consistencyVersion information consistent across version field, updates, version_access, and distribution_dates
5/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelStructured version control with timestamps
evidenceupdates: detailed version progression with timestamps (v1.0 Jan 17 2025, v1.1 Jan 17 2025, v2.0.0 Apr 16 2025, v2.0.1 Aug 18 2025, v3.0.0 2025) and change descriptions (v1.1 added MFCC features), participant counts per version, semi-annual release schedule; version_access: all versions maintained with version-specific DOIs; distribution_dates: release timeline documented
qualityExemplary provenance tracking with version-specific DOIs, timestamped releases, participant count evolution, and feature addition logs
correctnessVersion dates chronologically consistent; version numbering follows semantic versioning conventions; participant growth plausible (306→833 over ~1 year)
consistencyProvenance information consistent across updates, version_access, and distribution_dates fields
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1148 chars with specific participant counts and version details, purposes: 4 entries averaging 350+ chars each with detailed research objectives
qualityExceptional narrative depth with concrete metrics and comprehensive research goals
correctnessParticipant numbers (833 adults, 61,937 recordings) semantically consistent with v3.0 release scope
consistencyPurposes align with NIH Bridge2AI program mission and funding goals
license
license: Bridge2AI Voice Registered Access License
✓ 1/1 R10 3.Data Reuse and Interoperability License Terms Allow Reuse
evidencelicense: Bridge2AI Voice Registered Access License; license_and_use_terms specifies commercial and non-commercial research use permitted, no geographic restrictions, no restriction to specific research type
qualityClear license allowing broad research reuse (commercial + non-commercial), Open Science principles, publication in open-access journals encouraged
semanticLicense terms semantically permissive for research while maintaining participant protections through registered access mechanism
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice, description: 1148 chars comprehensive, keywords: 73 keywords, license_and_use_terms: detailed registered access license
qualityAll mandatory fields present with exceptional completeness and detail
correctnessDOI format valid, PhysioNet prefix 10.13026 verified
consistencyAll mandatory fields semantically coherent and appropriate for voice biomarker dataset
is_tabular
is_tabular: false
✓ 1/1 R10 5.Data Composition and Structure Variable-Level Metadata and Tabular Flag
evidenceis_tabular: false; distribution_formats document TSV phenotype files with JSON data dictionaries, static_features.tsv with static_features.json dictionary; acquisition_methods list 22 adult acoustic tasks, 36 pediatric tasks, validated questionnaires (VHI-10, PHQ-9, GAD-7, etc.)
qualityTabular flag correctly set to false (multimodal: audio + phenotypes); variable documentation via JSON dictionaries; 22 adult tasks + 36 pediatric tasks + 10+ questionnaires enumerated
semanticis_tabular=false semantically accurate for multimodal dataset; variable metadata comprehensive (acoustic features, phenotypes, questionnaires all documented)
keywords
keywords:
- voice biomarker
- acoustic biomarker
- Bridge2AI
- voice AI
- voice disorders
- neurological disorders
- neurodegenerative disorders
- mood disorders
- psychiatric disorders
- respiratory disorders
- pediatric voice disorders
- speech disorders
- Parkinson's disease
- Alzheimer's disease
- depression
- schizophrenia
- bipolar disorder
- stroke
- ALS
- autism spectrum disorder
- speech delay
- laryngeal cancer
- vocal fold paralysis
- muscle tension dysphonia
- laryngeal dystonia
- COPD
- chronic cough
- airway stenosis
- obstructive sleep apnea
- spectrogram
- MFCC
- mel-frequency cepstral coefficients
- OpenSMILE
- Praat
- Parselmouth
- torchaudio
- federated learning
- ethical AI
- multimodal health data
- electronic health records
- EHR
- radiomics
- genomics
- FAIR principles
- CARE principles
- PhysioNet
- Health Data Nexus
- BIDS
- Brain Imaging Data Structure
- b2aiprep
✓ 1/1 R10 1.Dataset Discovery and Identification Keywords or Tags for Searchability
evidencekeywords: 52 terms including voice biomarker, Bridge2AI, specific diseases (Parkinson's, Alzheimer's, depression), technical methods (spectrogram, MFCC, OpenSMILE), standards (FAIR, BIDS)
qualityExtensive keyword set covering domain (voice AI), conditions (12+ diseases), methods (acoustic features), tools (OpenSMILE, Praat), and standards (FAIR, BIDS)
semanticKeywords semantically diverse and specific, enabling discovery across clinical, technical, and ethical dimensions
✓ 1/1 R10 10.Cross-Platform and Community Integration Community Standards or Schema Conformance
evidenceBIDS v1.9.0 specification referenced 6+ times; ICD-10 codes for diagnoses; FAIR principles and CARE principles in keywords; REDCap for clinical data management; Bridge2AI protocols as emerging standard
qualityMultiple community standards adopted: BIDS v1.9.0 (neuroimaging data structure), ICD-10 (clinical coding), FAIR (findability/accessibility/interoperability/reusability), CARE (collective benefit/authority/responsibility/ethics), REDCap (clinical research data)
semanticStandards adoption semantically appropriate - BIDS enables neuroimaging ecosystem integration, ICD-10 ensures clinical interoperability, FAIR/CARE principles align with Bridge2AI program values
✓ 1/1 R10 3.Data Reuse and Interoperability Schema or Ontology Conformance Stated
evidenceBIDS v1.9.0 specification explicitly referenced 6+ times; ICD-10 codes mentioned for diagnoses; Bridge2AI protocols reference; FAIR principles and CARE principles listed in keywords
qualityExplicit conformance to BIDS v1.9.0 (Brain Imaging Data Structure), ICD-10 for clinical coding, FAIR and CARE principles for ethical AI
semanticSchema conformance semantically appropriate - BIDS ensures interoperability with neuroimaging/biomedical research ecosystems
✓ 1/1 R10 5.Data Composition and Structure Data Topics or Conditions Represented
evidenceinstances.instance_type: Human participants with clinical diagnoses recruited from specialty clinics; keywords include 18+ specific conditions (Parkinson's, Alzheimer's, depression, schizophrenia, laryngeal cancer, COPD, autism, etc.); subpopulations describe 5 disease categories
qualityConditions comprehensively documented: 18+ specific diseases in keywords, 5 disease cohort categories (voice, neurological, mood, respiratory, pediatric), clinical validation methods per condition
semanticConditions semantically aligned with voice biomarker research - diseases known to manifest in voice (Parkinson's, vocal fold paralysis, depression, COPD)
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice, description: 1148 chars comprehensive, keywords: 73 keywords, license_and_use_terms: detailed registered access license
qualityAll mandatory fields present with exceptional completeness and detail
correctnessDOI format valid, PhysioNet prefix 10.13026 verified
consistencyAll mandatory fields semantically coherent and appropriate for voice biomarker dataset
4/5 R20 Q10 (Metadata Quality & Content) Interoperability and Standardization
levelStandard formats + partial schema compliance
evidenceformat: TSV, JSON, GZ, Parquet (standard formats); encoding: implicit UTF-8 for text; description mentions 'BIDS v1.9.0 compliant structure' for directory organization and TSV/JSON phenotype files; keywords include BIDS, FAIR principles, CARE principles
qualityStrong standardization with BIDS v1.9.0 compliance documented in description/distribution fields, though D4D-core schema lacks explicit conforms_to_schema field present in full schema
correctnessBIDS v1.9.0 is appropriate standard for neuroimaging/voice biomarker data; Parquet and TSV formats semantically correct for time-series and tabular phenotype data
consistencyBIDS compliance claim consistent across multiple distribution format descriptions and phenotype data organization
4/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelComprehensive strategies with partial tool documentation
evidencepreprocessing_strategies: 8 entries mentioning b2aiprep, SenseLab, OpenSMILE, Parselmouth, Praat, torchaudio, OpenAI Whisper with specific parameters (16kHz resampling, 25ms window, 512-point FFT, 60 MFCCs); cleaning_strategies: 4 entries detailing HIPAA Safe Harbor, audio privacy, REDCap field removal, audit protocol; labeling_strategies: site clinician assignment per ICD-10; machine_annotation_tools: Whisper, b2aiprep, SenseLab; keywords list OpenSMILE, Praat, Parselmouth, torchaudio, b2aiprep
qualityExcellent strategy documentation with specific tools and parameters; minor gap: software versions/URLs not in D4D-core schema fields (present in external_resources though: GitHub links for b2aiprep, bridge2ai-redcap)
correctnessTools semantically appropriate: OpenSMILE (audio features), Whisper (transcription), b2aiprep (BIDS conversion), Parselmouth/Praat (phonetics), torchaudio (PyTorch audio)
consistencyTool mentions consistent across preprocessing_strategies, machine_annotation_tools, keywords, and external_resources
5/5 R20 Q3 (Structural Completeness) Keyword Diversity
level≥8 keywords
evidencekeywords: 73 unique keywords covering diseases (Parkinson's, Alzheimer's, depression, COPD), methods (MFCC, spectrogram, OpenSMILE), standards (FAIR, CARE, BIDS), platforms (PhysioNet, Health Data Nexus)
qualityExceptional keyword diversity enabling comprehensive discoverability across multiple research domains
correctnessAll keywords semantically appropriate for voice AI biomarker research domain
consistencyKeywords align with described disease cohorts and technical methods
publisher
publisher: Bridge2AI-Voice Consortium, University of South Florida
✓ 1/1 R10 10.Cross-Platform and Community Integration Dataset Published on a Recognized Platform
evidencepublisher: Bridge2AI-Voice Consortium, University of South Florida; distributions via PhysioNet (MIT Laboratory for Computational Physiology, NIBIB-funded); Health Data Nexus (University of Toronto T-CAIREM) as alternative platform
qualityPublished on multiple recognized platforms: PhysioNet (NIH-funded biomedical data repository), Health Data Nexus (T-CAIREM cloud research platform), both with established reputations in biomedical research community
semanticPlatform choice semantically appropriate - PhysioNet specialized in physiological signals/biomedical time-series, Health Data Nexus provides cloud compute integration
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Funding Sources and Mechanisms Listed
evidencefunders: 2 documented - NIH Common Fund Bridge2AI Program administering OT2OD032720 (funder:1), NIBIB supporting PhysioNet via R01EB030362 (funder:2); publisher: Bridge2AI-Voice Consortium, University of South Florida
qualityFunding sources identified: NIH Common Fund (primary dataset collection), NIBIB (distribution platform infrastructure); administering institute (NIH Office of the Director) and study section (DCMM) documented
semanticFunding sources semantically appropriate - NIH Common Fund supports cross-cutting programs (Bridge2AI fits mandate), NIBIB supports biomedical imaging/PhysioNet infrastructure
download_url
download_url: https://physionet.org/content/b2ai-voice/
✓ 1/1 R10 1.Dataset Discovery and Identification Landing Page and Resources (page, hierarchical resources)
evidencepage: https://docs.b2ai-voice.org; download_url: https://physionet.org/content/b2ai-voice/; external_resources: 12 documented resources including PhysioNet, GitHub repos, documentation
qualityMultiple accessible landing pages (project docs and PhysioNet), hierarchical resources structure with 12 external resources providing documentation, code, and access points
semanticURLs semantically appropriate - docs.b2ai-voice.org for documentation, physionet.org for data access, github.com for code repositories
✓ 1/1 R10 2.Dataset Access and Retrieval Download URL or Platform Link Available
evidencedownload_url: https://physionet.org/content/b2ai-voice/; distributions include paths to PhysioNet and DACO contact (DACO@b2ai-voice.org) for controlled access
qualityDirect download URL for registered access dataset, specific contact mechanism (email) for controlled access raw audio, alternative platform (Health Data Nexus) documented
semanticDownload mechanisms semantically differentiated by access tier - public URL for registered, email application for controlled
1/1 R20 Q16 (FAIRness & Accessibility) Findability (Persistent Links)
levelPass
evidencepage: https://docs.b2ai-voice.org, download_url: https://physionet.org/content/b2ai-voice/, external_resources: 12 entries with URLs to PhysioNet, GitHub (bridge2ai-docs, b2aiprep, bridge2ai-redcap), NIH RePORTER, Health Data Nexus, Zenodo, Interspeech DOI, Bridge2AI program, training site
qualityExcellent persistent link coverage with documentation, download, and external resource URLs across multiple platforms
correctnessAll URLs follow valid format with appropriate domains (physionet.org, github.com, doi.org, nih.gov, zenodo.org)
consistencyURLs consistent with described distribution platforms (PhysioNet primary, Health Data Nexus alternative)
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path (platform, login, policy)
evidencedistribution_formats: PhysioNet registered access for features (TSV/JSON/Parquet), controlled access for raw audio (WAV/GZ via DACO); license_and_use_terms: registered users sign Bridge2AI Voice Registered Access Agreement, raw audio requires DACO application + institutional DTUA; download_url: https://physionet.org/content/b2ai-voice/; external_resources: DACO contact mailto:DACO@b2ai-voice.org
qualityExceptionally clear two-tier access mechanism: (1) PhysioNet registered access with DUA for features, (2) DACO-controlled access with institutional DTUA for raw audio biometrics
correctnessAccess tiers semantically appropriate for biometric data: lower barrier for de-identified features, higher barrier with institutional oversight for re-identifiable raw audio
consistencyAccess mechanism consistent across distribution_formats, license_and_use_terms, and confidential_elements documentation
1/1 R20 Q20 (FAIRness & Accessibility) Interlinking Across Platforms
levelPass
evidenceexternal_resources: cross-platform links verified - PhysioNet (primary distribution), Health Data Nexus (alternative with cloud compute), GitHub (3 repos: bridge2ai-docs, b2aiprep, bridge2ai-redcap), Zenodo (archive), NIH RePORTER (grant tracking), Bridge2AI program site, training portal; page: https://docs.b2ai-voice.org; download_url: https://physionet.org/content/b2ai-voice/
qualityExcellent cross-platform interlinking with 7+ distinct platforms (PhysioNet, Health Data Nexus, GitHub, Zenodo, NIH, Bridge2AI, training site) providing complementary access points
correctnessPlatform links semantically appropriate: PhysioNet for biomedical data distribution, GitHub for code/docs, Zenodo for archival, NIH RePORTER for grant transparency
consistencyCross-platform links consistent with described distribution strategy (PhysioNet primary, Health Data Nexus alternative, GitHub for open-source tools)
1/1 R20 Q6 (Metadata Quality & Content) Dataset Identification Metadata
levelPass
evidencedoi: 10.13026/37yb-1t42, page: https://docs.b2ai-voice.org, download_url: https://physionet.org/content/b2ai-voice/
qualityMultiple persistent identifiers present with PhysioNet DOI and project documentation URLs
correctnessDOI prefix 10.13026 matches PhysioNet registrar, DOI format valid
consistencyDOI, page URL, and download URL all resolve to consistent Bridge2AI-Voice project resources
purposes
purposes:
- id: voice:purpose:1
  description: 'Integrate the use of voice as a biomarker of health in clinical care by generating a substantial
    multi-institutional, ethically sourced, and diverse voice database linked to multimodal health biomarkers
    to fuel voice AI research and build predictive models to assist in screening, diagnosis, and treatment
    of a broad range of diseases.

    '
- id: voice:purpose:2
  description: 'Create an ethically sourced flagship dataset of 10,000 voices linked to health information
    to enable future research in artificial intelligence and support critical insights into the use of
    voice as a biomarker of health, addressing the pressing need for large, high quality, multi-institutional
    and diverse voice databases linked to other health biomarkers.

    '
- id: voice:purpose:3
  description: 'Establish standards, best practices, and guidelines for voice data collection and analysis
    to advance the field of acoustic biomarkers by developing new standards that are AI/ML friendly and
    enable voice to emerge as a biomarker of health.

    '
- id: voice:purpose:4
  description: 'Address ethical, legal, and social challenges surrounding voice AI including risks of
    voice re-identification, vulnerabilities like voice AI hacking, concerns around voice data sharing
    and privacy, and the influence of gender and racial diversity on voice AI.

    '
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Motivation or Purpose for Dataset Creation
evidencepurposes: 4 documented purposes - integrate voice as health biomarker via multi-institutional diverse database (purpose:1), create 10,000-voice flagship dataset (purpose:2), establish AI/ML-friendly standards (purpose:3), address ethical/legal challenges (purpose:4)
qualityComprehensive motivation across technical, clinical, standards, and ethical dimensions; purposes explicitly address research gaps in voice AI field
semanticPurposes semantically coherent with Bridge2AI program goals - data generation, standards development, ethical frameworks aligned with AI-ready biomedical data objectives
5/5 R20 Q2 (Structural Completeness) Entry Length Adequacy
level>200 chars
evidencedescription: 1148 chars with specific participant counts and version details, purposes: 4 entries averaging 350+ chars each with detailed research objectives
qualityExceptional narrative depth with concrete metrics and comprehensive research goals
correctnessParticipant numbers (833 adults, 61,937 recordings) semantically consistent with v3.0 release scope
consistencyPurposes align with NIH Bridge2AI program mission and funding goals
tasks
tasks:
- id: voice:task:1
  description: 'Enable development of AI/ML predictive models for screening, diagnosis, and treatment
    of voice disorders including laryngeal cancers, vocal fold paralysis, muscle tension dysphonia, laryngeal
    dystonia, benign laryngeal lesions, and pre-cancerous lesions.

    '
- id: voice:task:2
  description: 'Support machine learning models for neurological and neurodegenerative disorders including
    Alzheimer''s disease, Parkinson''s disease, mild cognitive impairment, other dementias, and ALS, detecting
    voice and speech changes.

    '
- id: voice:task:3
  description: 'Develop AI algorithms for mood and psychiatric disorder detection including depression,
    schizophrenia, bipolar disorder, anxiety disorders, ADHD, PTSD, OCD, and borderline personality disorder.

    '
- id: voice:task:4
  description: 'Create machine learning models for respiratory disorder screening and therapeutic monitoring
    using respiratory sounds, cough sounds, and voice, applicable to conditions such as chronic cough,
    COPD, airway stenosis, and other respiratory conditions.

    '
- id: voice:task:5
  description: 'Build AI models for pediatric voice and speech disorder detection using 36 pediatric-specific
    acoustic tasks and specialized questionnaires, addressing the relative scarcity of pediatric voice
    data for participants aged 2-18.

    '
- id: voice:task:6
  description: 'Promote application of AI/ML for voice research through workforce development, curriculum
    creation on voice biomarkers of health for FAIR and CARE AI models, and fostering collaborations especially
    with researchers from underserved communities.

    '
- id: voice:task:7
  description: 'Enable AI model pretraining, fine-tuning, benchmarking, and validation using the standardized
    BIDS-compliant dataset structure with features extracted via b2aiprep, SenseLab, OpenSMILE, Parselmouth,
    Praat, and torchaudio toolkits.

    '
✓ 1/1 R10 5.Data Composition and Structure Variable-Level Metadata and Tabular Flag
evidenceis_tabular: false; distribution_formats document TSV phenotype files with JSON data dictionaries, static_features.tsv with static_features.json dictionary; acquisition_methods list 22 adult acoustic tasks, 36 pediatric tasks, validated questionnaires (VHI-10, PHQ-9, GAD-7, etc.)
qualityTabular flag correctly set to false (multimodal: audio + phenotypes); variable documentation via JSON dictionaries; 22 adult tasks + 36 pediatric tasks + 10+ questionnaires enumerated
semanticis_tabular=false semantically accurate for multimodal dataset; variable metadata comprehensive (acoustic features, phenotypes, questionnaires all documented)
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Primary Research Objectives or Tasks
evidencetasks: 7 documented tasks including AI/ML models for voice disorders (task:1), neurological disorders (task:2), mood/psychiatric (task:3), respiratory (task:4), pediatric (task:5), workforce development (task:6), model pretraining/benchmarking (task:7)
qualitySpecific research objectives enumerate clinical applications (5 disease categories), technical applications (pretraining, benchmarking), and community goals (workforce development)
semanticTasks semantically aligned with purposes - clinical AI models (tasks 1-5) support biomarker integration goal (purpose:1), benchmarking (task:7) supports standards goal (purpose:3)
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Data Acquisition Methods Listed
evidenceacquisition_methods: 4 documented - 22 adult acoustic tasks (respiration, prolonged vowels, diadochokinesis, passage reading, free speech), 36 pediatric tasks, validated questionnaires (VHI-10, PHQ-9, GAD-7, MOCA, C-VHI-10), EHR access for gold standard validation
qualityComprehensive acquisition documentation: 22 adult tasks enumerated by category (non-voice, voice/non-speech, speech), 36 pediatric tasks named, 10+ validated questionnaires listed, EHR linkage described
semanticAcquisition methods semantically appropriate for voice research - acoustic tasks span breathing (non-voice), phonation (vowels), articulation (diadochokinesis), connected speech (passages), multimodal validation (questionnaires, EHR)
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidenceacquisition_methods: 4 entries detailing 22 adult acoustic tasks, 36 pediatric tasks, validated questionnaires (VHI-10, PHQ-9, GAD-7, MOCA, etc.), EHR linkage; collection_mechanisms: iPad 9th/10th gen or iPad Air 5th gen with Avid AE-36 microphone, Bridge2AI-Voice App, reproschema-ui for pediatrics; data_collectors: research teams at 5 North American sites, medical graduate/undergraduate students, clinicians as co-investigators, participant compensation $40-120; collection_timeframes: September 2022 - November 2026, semi-annual releases, v1.0 Jan 2025 (306), v3.0 2025 (833)
qualityExceptionally detailed collection protocol with specific hardware, software, personnel, compensation, timeline, and task descriptions
correctnessHardware plausible (iPad with professional microphone standard for clinical voice research); timeline consistent (2022 start to 2026 end, 833 participants accumulated over 3 years across 5 sites)
consistencyCollection details consistent across mechanisms/methods/collectors/timeframes; participant growth (306→833) aligns with timeline progression
5/5 R20 Q15 (Technical Documentation) Human Subject Representation
levelDetailed demographics and inclusion/exclusion criteria
evidenceinstances: 833 adult participants with instance_type 'Human participants with clinical diagnoses recruited from specialty clinics', label_description detailing clinician-assigned diagnoses via laryngoscopy/MRI/CT/EHR; subpopulations: 6 detailed entries for Voice Disorders (laryngeal cancer, dysphonia, paralysis validated by laryngoscopy/stroboscopy), Neurological (ages 44-85, Alzheimer's/Parkinson's/ALS validated by MRI/CT/genomics), Mood/Psychiatric (depression/schizophrenia/bipolar validated by EHR/psychiatrist), Respiratory (stenosis/COPD validated by spirometry/CT), Pediatric (ages 2-18, 36 tasks, 300 participants), Controls (VHI-10, PHQ-9, GAD-7)
qualityExceptional demographic detail with disease-specific cohorts, age ranges, clinical validation methods, and separate pediatric population description
correctnessSubpopulation definitions clinically plausible with appropriate validation methods for each disease category; age ranges appropriate (44-85 for neurodegeneration, 2-18 for pediatrics)
consistencySubpopulations align with described tasks (22 adult tasks, 36 pediatric tasks) and purposes (voice/neurological/mood/respiratory disorder AI models)
addressing_gaps
addressing_gaps:
- id: voice:gap:1
  description: 'Address the lack of large, high quality, multi-institutional and diverse voice databases
    linked to multimodal health biomarkers (demographics, imaging, genomics, risk factors) necessary to
    fuel voice AI research and answer tangible clinical questions.

    '
- id: voice:gap:2
  description: 'Overcome limitations in existing voice and psychiatric disorder research that has relied
    on small datasets with limited demographic diversity reporting, lack of standardized data collection
    protocols, and possible confounders limiting external validity.

    '
- id: voice:gap:3
  description: 'Fill the gap in pediatric voice and speech analysis research, which is sparser partly
    due to ethical concerns and challenges in data acquisition for this cohort, particularly for autism
    spectrum disorder and speech delay detection.

    '
- id: voice:gap:4
  description: 'Establish missing standards for voice data collection, acoustic analysis, and ethical
    frameworks for consenting to voice data collection, sharing, and utilization in the context of voice
    AI technology development and clinical adoption.

    '
- id: voice:gap:5
  description: 'Develop software and cloud infrastructure for automated voice data collection through
    a smartphone application (Bridge2AI-Voice App) that allows non-invasive, user-friendly, high quality
    voice data collection while minimizing human manipulation and implementing federated learning technology
    to minimize data sharing while preserving patient privacy.

    '
no field-level feedback matched
creators
creators:
- id: voice:creator:1
  description: 'Bridge2AI-Voice Consortium led by Dr. Yael Bensoussan (Contact PI, University of South
    Florida, Department of Otolaryngology) and Dr. Olivier Elemento (Co-PI, Weill Cornell Medicine). The
    multidisciplinary consortium includes over 50 investigators from 12+ institutions across North America
    spanning clinical medicine, biomedical research, machine learning, data science, social science, and
    ethics. Key co-investigators include: Alexandros Sigaras (Weill Cornell Medicine), Anais Rameau (Weill
    Cornell Medicine), Maria Powell (Vanderbilt University Medical Center), Ruth Bahr (USF), Jennifer
    Siu (Hospital for Sick Children), Philip Payne (Washington University in St. Louis), David Dorr (Oregon
    Health and Science University), Jean-Christophe Belisle-Pipon (Simon Fraser University), Vardit Ravitsky
    (The Hastings Center), Satrajit Ghosh (MIT), Frank Rudzicz (University of Toronto), Jordan Lerner-Ellis
    (Sinai Health), Don Bolser (University of Florida), Alistair Johnson (MIT/PhysioNet), and Jennifer
    Siu (Hospital for Sick Children).

    '
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Creators and Acknowledgements Documented
evidencecreators: Bridge2AI-Voice Consortium with 50+ investigators from 12+ institutions, Contact PI Dr. Yael Bensoussan (USF), Co-PI Dr. Olivier Elemento (Weill Cornell), 15+ key co-investigators named with institutional affiliations
qualityComprehensive creator documentation: lead PIs identified, consortium size (50+ investigators, 12+ institutions), key contributors named with roles and affiliations (clinical, ML, ethics, data science)
semanticCreator acknowledgements semantically complete - leadership identified, multidisciplinary expertise documented (clinicians, data scientists, ethicists), institutional affiliations enable contact
5/5 R20 Q7 (Metadata Quality & Content) Funding and Acknowledgements Completeness
levelFunders with grants + creators with affiliations
evidencefunders: 2 entries with NIH Common Fund Bridge2AI grant 3OT2OD032720-01S3 ($4.66M total 2025 funding detailed) and NIBIB R01EB030362; creators: Bridge2AI-Voice Consortium with 50+ investigators from 12+ institutions, lead PI Dr. Yael Bensoussan (USF), Co-PI Dr. Olivier Elemento (Weill Cornell)
qualityComprehensive funding documentation with grant numbers, amounts, dates, and detailed creator affiliations across multi-institutional consortium
correctnessGrant number 3OT2OD032720-01S3 follows NIH supplemental award pattern (leading '3' indicates supplement)
consistencyFunding source (NIH Bridge2AI) aligns with project scope and multi-institutional consortium structure
funders
funders:
- id: voice:funder:1
  description: 'National Institutes of Health (NIH) Common Fund Bridge2AI Program. Grant number: 3OT2OD032720-01S3.
    Opportunity Number: OTA-21-008. Project dates: September 1, 2022 to November 30, 2026. Total funding
    in 2025: $4,660,942 (Direct Costs: $4,072,321, Indirect Costs: $588,621). Administering Institute:
    NIH Office of the Director. Study Section: Data Coordination, Mapping, and Modeling (DCMM).

    '
- id: voice:funder:2
  description: 'National Institute of Biomedical Imaging and Bioengineering (NIBIB). Supports PhysioNet
    managed by MIT Laboratory for Computational Physiology under NIH grant number R01EB030362, which serves
    as the primary distribution platform for the Bridge2AI-Voice dataset.

    '
⚠ low R10 · consistency
issueSecond funder references R01EB030362 supporting PhysioNet infrastructure but not the dataset collection itself
fieldsfunders
fixClarify distinction between dataset funding and distribution platform funding
⚠ low R10 · correctness
issueGrant number format 3OT2OD032720-01S3 uses non-standard prefix '3' instead of typical NIH type codes
fieldsfunders
fixVerify grant number format - NIH typically uses 1/2/5 prefixes for original/competitive renewal/non-competing continuation awards
⚠ low R20 · correctness
issueGrant number format '3OT2OD032720-01S3' appears unusual with leading '3' (supplemental award marker) but otherwise follows NIH pattern
fieldsfunders
fixVerify grant number is correctly transcribed; NIH grants typically start with award type letter not digit
✓ 1/1 R10 7.Scientific Motivation and Funding Transparency Funding Sources and Mechanisms Listed
evidencefunders: 2 documented - NIH Common Fund Bridge2AI Program administering OT2OD032720 (funder:1), NIBIB supporting PhysioNet via R01EB030362 (funder:2); publisher: Bridge2AI-Voice Consortium, University of South Florida
qualityFunding sources identified: NIH Common Fund (primary dataset collection), NIBIB (distribution platform infrastructure); administering institute (NIH Office of the Director) and study section (DCMM) documented
semanticFunding sources semantically appropriate - NIH Common Fund supports cross-cutting programs (Bridge2AI fits mandate), NIBIB supports biomedical imaging/PhysioNet infrastructure
✗ 0/1 R10 7.Scientific Motivation and Funding Transparency Grant IDs or Award Numbers Present
evidencefunders reference grant numbers: 3OT2OD032720-01S3 (primary), R01EB030362 (PhysioNet); external_resources links NIH RePORTER project 11376382
qualityGrant numbers present but format non-standard - 3OT2OD032720-01S3 uses '3' prefix atypical for NIH (standard patterns: 1/2/5 for original/competitive renewal/non-competing continuation); R01EB030362 follows standard R01 mechanism format
semanticPrimary grant 3OT2OD032720-01S3 has non-standard prefix '3' - NIH typically uses 1 (new/renewal), 2 (competitive renewal), 5 (continuation); -01S3 suffix suggests supplement. Secondary grant R01EB030362 follows standard NIH R01 pattern. Scoring 0 due to non-standard format despite presence.
5/5 R20 Q7 (Metadata Quality & Content) Funding and Acknowledgements Completeness
levelFunders with grants + creators with affiliations
evidencefunders: 2 entries with NIH Common Fund Bridge2AI grant 3OT2OD032720-01S3 ($4.66M total 2025 funding detailed) and NIBIB R01EB030362; creators: Bridge2AI-Voice Consortium with 50+ investigators from 12+ institutions, lead PI Dr. Yael Bensoussan (USF), Co-PI Dr. Olivier Elemento (Weill Cornell)
qualityComprehensive funding documentation with grant numbers, amounts, dates, and detailed creator affiliations across multi-institutional consortium
correctnessGrant number 3OT2OD032720-01S3 follows NIH supplemental award pattern (leading '3' indicates supplement)
consistencyFunding source (NIH Bridge2AI) aligns with project scope and multi-institutional consortium structure
instances
instances:
- id: voice:instance:1
  description: 'Adult participants presenting at specialty clinics across multiple sites in North America.
    Participants selected based on membership to five predetermined disease cohort groups: Voice Disorders,
    Neurological and Neurodegenerative Disorders, Mood and Psychiatric Disorders, Respiratory Disorders,
    and Pediatric Voice and Speech Disorders. Version 3.0 contains approximately 833 adult participants
    with ~61,937 voice-derived recordings. Pediatric dataset v1.0 adds 300 participants. Enrollment anticipated
    to reach 10,000 participants by 2027.

    '
  instance_type: Human participants with clinical diagnoses recruited from specialty clinics
  counts: 833
  label: true
  label_description: 'Diagnostic labels assigned by clinical assessment at each site. Clinicians provided
    diagnoses based on clinical interview and appropriate work-up including laryngoscopy, stroboscopy,
    MRI, CT, whole genome sequencing, EHR records, and medication prescriptions. Labels include diagnostic
    categories: vocal pathologies, neurological disorders, psychiatric conditions, respiratory disorders,
    and pediatric voice/speech disorders. Per Bridge2AI protocols and ICD-10 codes. Single labeler per
    participant (site clinician).

    '
✓ 1/1 R10 5.Data Composition and Structure Number of Instances or Samples Reported
evidenceinstances.counts: 833 adult participants; description specifies ~61,937 voice-derived recordings; pediatric v1.0 adds 300 participants; enrollment target 10,000 by 2027
qualitySpecific instance counts: 833 adult participants, ~61,937 recordings, 300 pediatric participants; progression documented (v1.0: 306 participants, 12,523 recordings → v3.0: 833, ~61,937)
semanticInstance counts semantically detailed and versioned - current state (v3.0: 833), historical (v1.0: 306), future target (10,000), recordings-per-participant ratio provided
✓ 1/1 R10 5.Data Composition and Structure Data Topics or Conditions Represented
evidenceinstances.instance_type: Human participants with clinical diagnoses recruited from specialty clinics; keywords include 18+ specific conditions (Parkinson's, Alzheimer's, depression, schizophrenia, laryngeal cancer, COPD, autism, etc.); subpopulations describe 5 disease categories
qualityConditions comprehensively documented: 18+ specific diseases in keywords, 5 disease cohort categories (voice, neurological, mood, respiratory, pediatric), clinical validation methods per condition
semanticConditions semantically aligned with voice biomarker research - diseases known to manifest in voice (Parkinson's, vocal fold paralysis, depression, COPD)
5/5 R20 Q15 (Technical Documentation) Human Subject Representation
levelDetailed demographics and inclusion/exclusion criteria
evidenceinstances: 833 adult participants with instance_type 'Human participants with clinical diagnoses recruited from specialty clinics', label_description detailing clinician-assigned diagnoses via laryngoscopy/MRI/CT/EHR; subpopulations: 6 detailed entries for Voice Disorders (laryngeal cancer, dysphonia, paralysis validated by laryngoscopy/stroboscopy), Neurological (ages 44-85, Alzheimer's/Parkinson's/ALS validated by MRI/CT/genomics), Mood/Psychiatric (depression/schizophrenia/bipolar validated by EHR/psychiatrist), Respiratory (stenosis/COPD validated by spirometry/CT), Pediatric (ages 2-18, 36 tasks, 300 participants), Controls (VHI-10, PHQ-9, GAD-7)
qualityExceptional demographic detail with disease-specific cohorts, age ranges, clinical validation methods, and separate pediatric population description
correctnessSubpopulation definitions clinically plausible with appropriate validation methods for each disease category; age ranges appropriate (44-85 for neurodegeneration, 2-18 for pediatrics)
consistencySubpopulations align with described tasks (22 adult tasks, 36 pediatric tasks) and purposes (voice/neurological/mood/respiratory disorder AI models)
1/1 R20 Q5 (Structural Completeness) Data File Size Availability
levelPass
evidenceinstances: counts=833 adult participants, ~61,937 voice-derived recordings explicitly documented; pediatric v1.0 with 300 participants mentioned
qualityInstance counts clearly documented with specific participant and recording numbers per version
correctnessCounts plausible for multi-year multi-site clinical collection (833 participants across 5 sites over 3+ years)
consistencyParticipant counts consistent with version progression (v1.0: 306 → v3.0: 833)
known_biases
known_biases:
- id: voice:bias:1
  description: 'Clinic-based recruitment introduces selection bias toward treatment-seeking populations.
    Participants are recruited from high-volume specialty clinics, which may not represent the full spectrum
    of disease severity or the general population with these conditions. Underrepresentation of individuals
    with less trust in the medical system or less proximity to collection sites.

    '
- id: voice:bias:2
  description: 'Current releases contain only English-speaking participants, which may bias acoustic features
    and models trained on this data toward English-language phonological patterns. Spanish protocols are
    under development for future releases.

    '
- id: voice:bias:3
  description: 'Limited geographic diversity: data collected at five North American sites only, which
    may not capture regional variation in voice characteristics, dialects, or environmental acoustic conditions.

    '
- id: voice:bias:4
  description: 'Imbalanced distribution across disease categories in public releases; the dataset does
    not contain equal representation across all five disease cohort groups. This may affect the performance
    and fairness of models trained on the dataset.

    '
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Systematic Biases Identified and Described
evidenceknown_biases: 4 documented - clinic-based recruitment selection bias toward treatment-seeking populations (bias:1), English-only limiting phonological generalizability (bias:2), five North American sites limiting geographic/dialect diversity (bias:3), imbalanced disease category distribution (bias:4)
qualitySystematic bias discussion addressing selection (clinic recruitment), linguistic (English-only), geographic (5 North American sites), representational (disease category imbalance) dimensions
semanticBiases semantically well-characterized - selection bias mechanism explained (treatment-seeking populations), language bias impact specified (phonological patterns), geographic bias quantified (5 sites), distribution imbalance acknowledged
known_limitations
known_limitations:
- id: voice:limitation:1
  description: 'Version 3.0 (833 adult participants) represents an interim dataset from an ongoing collection
    targeting 10,000 participants by 2027. Statistical power for some analyses, particularly for less
    prevalent conditions, may be limited.

    '
- id: voice:limitation:2
  description: 'Raw audio waveforms are excluded from public release and available only through controlled
    access. This restricts certain types of analyses requiring the original audio signal and adds access
    overhead for researchers.

    '
- id: voice:limitation:3
  description: 'Remote data collection not included in initial releases; all data collected in clinical
    settings by research assistants, which may not generalize to naturalistic or home settings.

    '
- id: voice:limitation:4
  description: 'Multimodal data (imaging, genomics, full EHR data) not included in current public releases;
    only derived clinical features and diagnoses are provided.

    '
- id: voice:limitation:5
  description: 'No analysis of potential re-identification impact has been formally conducted on the de-identified
    dataset beyond HIPAA Safe Harbor standards.

    '
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Known Limitations Documented
evidenceknown_limitations: 5 documented - v3.0 interim dataset (833/10,000 target) limiting statistical power (limitation:1), raw audio controlled-access-only restricts analyses (limitation:2), clinical-only collection limits generalization to home settings (limitation:3), multimodal data not in current releases (limitation:4), no formal re-identification analysis beyond HIPAA (limitation:5)
qualityComprehensive limitation disclosure across sampling (interim 833/10,000), access (raw audio restrictions), generalizability (clinic vs home), completeness (multimodal data absent), privacy (no formal re-ID analysis)
semanticLimitations semantically substantive - acknowledge statistical power constraints, access tradeoffs, ecological validity concerns, data integration gaps, privacy assessment gaps
confidential_elements
confidential_elements:
- id: voice:confidential:1
  description: 'Raw audio waveforms are considered biometric identifiers under HIPAA and are restricted
    to controlled access only. Researchers must apply through the Data Access Compliance Office (DACO)
    and sign an institutional Data Use and Transfer Agreement (DTUA).

    '
- id: voice:confidential:2
  description: 'Open-response audio features (spectrograms, MFCCs, mel spectrograms, transcriptions, EMAs,
    and PPGs from open-response prompts) are removed from the public registered access dataset to protect
    participant privacy.

    '
- id: voice:confidential:3
  description: 'The dataset is covered under a Certificate of Confidentiality, which provides legal protection
    against compulsory demands such as court orders and subpoenas for identifying information about research
    participants.

    '
✓ 1/1 R10 2.Dataset Access and Retrieval Regulatory Restrictions and Confidentiality Level Specified
evidenceregulatory_restrictions: Certificate of Confidentiality per 45 CFR 46, HIPAA Safe Harbor, OMB M-07-16 PII protections; confidential_elements: raw audio as biometric identifier, open-response features restricted
qualityDetailed regulatory framework including Certificate of Confidentiality, HIPAA compliance, and specific confidentiality level distinctions (registered vs controlled access)
semanticRegulatory restrictions semantically consistent with data sensitivity - Certificate of Confidentiality appropriate for identifiable health research data
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Privacy Protections Beyond Deidentification
evidenceconfidential_elements: raw audio restricted to controlled access, open-response features removed from public dataset, Certificate of Confidentiality protection; sensitive_elements: tiered access based on re-identification risk
qualityMulti-layered privacy strategy: HIPAA Safe Harbor baseline, raw audio controlled access, open-response content removed, Certificate of Confidentiality legal protection, tiered access model
semanticPrivacy protections semantically coherent - highest risk data (raw audio) most restricted, Certificate of Confidentiality provides legal safeguards beyond technical de-identification
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License with DUA requirement; ip_restrictions: Fort Lauderdale Agreement principles, no IP blocking; regulatory_restrictions: Certificate of Confidentiality, HIPAA Safe Harbor, OMB M-07-16, 45 CFR 46; confidential_elements: raw audio biometric identifiers, controlled access via DACO
qualityExceptionally detailed governance framework with multi-tier access (registered for features, controlled for raw audio), IP protections, and regulatory compliance
correctnessAccess tiers semantically appropriate: registered access for de-identified features, controlled access (DACO + DTUA) for biometric raw audio
consistencyLicensing terms consistent with confidentiality requirements and regulatory restrictions across all documentation
subpopulations
subpopulations:
- id: voice:subpop:1
  name: Voice Disorders cohort
  description: 'Participants with laryngeal disorders including laryngeal cancer (T1-T4, biopsy proven),
    laryngitis (acute, chronic, bacterial, fungal, autoimmune), pre-cancerous lesions (keratosis, leukoplakia),
    benign vocal cord lesions (nodules, polyps, cysts, Reinke''s edema, recurrent laryngeal papilloma),
    muscle tension dysphonia, spasmodic dysphonia and laryngeal tremor, unilateral vocal fold paralysis,
    and glottic insufficiency/ presbyphonia. Validated by laryngoscopy images and stroboscopy videos.

    '
- id: voice:subpop:2
  name: Neurological and Neurodegenerative Disorders cohort
  description: 'Participants aged 44-85 with clinical diagnoses including mild cognitive impairment, Alzheimer''s
    disease, other dementias (frontotemporal, Lewy body, vascular, mixed, alcohol-induced), ALS (sporadic,
    familial, spinal/limb-onset, bulbar-onset), and Parkinson''s disease (idiopathic PD, multiple system
    atrophy, progressive supranuclear palsy, corticobasal degeneration). Validated by CT Brain, MRI Brain,
    serum biomarkers, and whole genome sequencing.

    '
- id: voice:subpop:3
  name: Mood and Psychiatric Disorders cohort
  description: 'Participants with clinical diagnoses including depression/major depressive disorder, bipolar
    I/II disorder, anxiety disorder, schizophrenia, ADHD, PTSD, OCD, borderline personality disorder,
    and other psychiatric disorders. Validated by electronic health summary, clinical diagnosis from psychiatrist,
    and medication history.

    '
- id: voice:subpop:4
  name: Respiratory Disorders cohort
  description: 'Participants with respiratory conditions including airway stenosis (bilateral vocal fold
    paralysis, supraglottic, glottic, posterior glottic, subglottic, tracheal stenosis, multi-level upper
    airway stenosis) and chronic cough (bothersome cough >8 weeks). Validated by spirometry, flow volume
    loops, and CT scan of neck/chest.

    '
- id: voice:subpop:5
  name: Pediatric cohort
  description: 'Pediatric participants aged 2-18 with voice and speech disorders recruited exclusively
    from Hospital for Sick Children (SickKids). Grouped by age: 2-4, 4-6, 6-10, 10+ years. Data collected
    using reproschema-ui with Bridge2AI-Voice pediatric protocol. Includes 36 pediatric-specific acoustic
    tasks and specialized questionnaires (C-VHI-10, PVOS, PVRQOL, PHQ-A). Pediatric dataset v1.0 available
    separately from adult dataset.

    '
- id: voice:subpop:6
  name: Control participants
  description: 'Healthy volunteer control participants who complete common questionnaires and voice tasks
    (Voice Handicap Index-10, PHQ-9, GAD-7, Winograd) to provide normative comparisons.

    '
✓ 1/1 R10 5.Data Composition and Structure Cohort or Subpopulations Characteristics Described
evidencesubpopulations: 6 documented cohorts - Voice Disorders (laryngeal pathologies validated by laryngoscopy), Neurological Disorders (ages 44-85, CT/MRI/genomics validation), Mood/Psychiatric, Respiratory, Pediatric (ages 2-18), Control participants
qualityDetailed subpopulation descriptions with clinical validation methods: laryngoscopy for voice disorders, CT/MRI/genomics for neurological, EHR/psychiatrist for mood disorders, spirometry for respiratory
semanticSubpopulation descriptions semantically comprehensive - clinical definitions, age ranges, validation methods, diagnostic categories aligned with research objectives
✓ 1/1 R10 5.Data Composition and Structure Data Topics or Conditions Represented
evidenceinstances.instance_type: Human participants with clinical diagnoses recruited from specialty clinics; keywords include 18+ specific conditions (Parkinson's, Alzheimer's, depression, schizophrenia, laryngeal cancer, COPD, autism, etc.); subpopulations describe 5 disease categories
qualityConditions comprehensively documented: 18+ specific diseases in keywords, 5 disease cohort categories (voice, neurological, mood, respiratory, pediatric), clinical validation methods per condition
semanticConditions semantically aligned with voice biomarker research - diseases known to manifest in voice (Parkinson's, vocal fold paralysis, depression, COPD)
5/5 R20 Q15 (Technical Documentation) Human Subject Representation
levelDetailed demographics and inclusion/exclusion criteria
evidenceinstances: 833 adult participants with instance_type 'Human participants with clinical diagnoses recruited from specialty clinics', label_description detailing clinician-assigned diagnoses via laryngoscopy/MRI/CT/EHR; subpopulations: 6 detailed entries for Voice Disorders (laryngeal cancer, dysphonia, paralysis validated by laryngoscopy/stroboscopy), Neurological (ages 44-85, Alzheimer's/Parkinson's/ALS validated by MRI/CT/genomics), Mood/Psychiatric (depression/schizophrenia/bipolar validated by EHR/psychiatrist), Respiratory (stenosis/COPD validated by spirometry/CT), Pediatric (ages 2-18, 36 tasks, 300 participants), Controls (VHI-10, PHQ-9, GAD-7)
qualityExceptional demographic detail with disease-specific cohorts, age ranges, clinical validation methods, and separate pediatric population description
correctnessSubpopulation definitions clinically plausible with appropriate validation methods for each disease category; age ranges appropriate (44-85 for neurodegeneration, 2-18 for pediatrics)
consistencySubpopulations align with described tasks (22 adult tasks, 36 pediatric tasks) and purposes (voice/neurological/mood/respiratory disorder AI models)
sensitive_elements
sensitive_elements:
- id: voice:sensitive:1
  description: 'Voice recordings are biometric identifiers under HIPAA. Raw audio waveforms are excluded
    from public release and available only through controlled access with DACO approval and institutional
    DTUA. Risks include voice re-identification, voice AI hacking, and illicit or unauthorized use of
    voice data.

    '
- id: voice:sensitive:2
  description: 'Electronic health record (EHR) data accessed with participant consent for gold standard
    validation of diagnoses and symptoms. Sensitive health information including diagnoses, symptoms,
    disease-specific clinical data, and multimodal health biomarkers.

    '
- id: voice:sensitive:3
  description: 'Dataset contains sensitive demographic information including racial and ethnic origins,
    sexual orientation, financial and socioeconomic status, and health data. All direct identifiers removed;
    indirect identifiers removed where creating significant re-identification risk.

    '
- id: voice:sensitive:4
  description: 'Free speech task transcriptions may contain potentially identifying information or external
    voices. All open-response audio features (spectrograms, MFCCs, mel spectrograms, transcriptions, EMAs,
    PPGs) removed from public feature-only dataset.

    '
- id: voice:sensitive:5
  description: 'Dataset is covered under Certificate of Confidentiality protecting against compulsory
    legal demands such as court orders and subpoenas for identifying information or characteristics of
    research participants.

    '
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Privacy Protections Beyond Deidentification
evidenceconfidential_elements: raw audio restricted to controlled access, open-response features removed from public dataset, Certificate of Confidentiality protection; sensitive_elements: tiered access based on re-identification risk
qualityMulti-layered privacy strategy: HIPAA Safe Harbor baseline, raw audio controlled access, open-response content removed, Certificate of Confidentiality legal protection, tiered access model
semanticPrivacy protections semantically coherent - highest risk data (raw audio) most restricted, Certificate of Confidentiality provides legal safeguards beyond technical de-identification
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Sensitive Content and Warnings Provided
evidencesensitive_elements: 5 documented - voice as biometric identifier with re-identification/hacking risks (sensitive:1), EHR linkage containing diagnoses/symptoms (sensitive:2), demographic data including race/ethnicity/sexual orientation/SES (sensitive:3), free speech transcriptions potentially identifying (sensitive:4), Certificate of Confidentiality protection (sensitive:5)
qualityComprehensive sensitive content disclosure: biometric re-identification risk, health information sensitivity, demographic identifiability, free speech risks, legal protections documented
semanticSensitive content warnings semantically appropriate - specific risks identified (voice re-identification, AI hacking), data types enumerated (EHR, demographics, free speech), mitigation strategies referenced (Certificate of Confidentiality, controlled access)
collection_mechanisms
collection_mechanisms:
- id: voice:collection:1
  description: 'Voice data collected in clinic using custom Bridge2AI-Voice App on iPad (9th or 10th generation)
    or iPad Air (5th generation) with Avid AE-36 microphone and Apple dongle connector. App collects breathing
    sounds and voice, speech, and linguistic tasks along with health information through surveys and validated
    questionnaires. Research assistant present during collection. Future remote data collection planned
    but not included in current releases.

    '
- id: voice:collection:2
  description: 'Pediatric data collected using reproschema-ui with Bridge2AI-Voice pediatric protocol.
    Same iPad hardware as adult protocol with age-appropriate tasks grouped by age range (2-4, 4-6, 6-10,
    10+ years). Questions read to participants when needed.

    '
- id: voice:collection:3
  description: 'Clinical data including EHR information, imaging (laryngoscopy, stroboscopy, MRI, CT),
    and genomic data extracted from sites independently and uploaded through REDCap database. No external
    multimodal data released in current dataset versions.

    '
✓ 1/1 R10 6.Data Provenance and Version Tracking Provenance and Source Derivation Documented
evidenceraw_data_sources and raw_sources document provenance: raw audio WAV from Bridge2AI-Voice App, REDCap/reproschema-ui for phenotypes; collection_mechanisms specify iPad hardware (9th/10th gen, Air 5th gen) with Avid AE-36 microphone; preprocessing_strategies document transformation pipeline (b2aiprep → spectrograms/MFCCs via FFT/OpenSMILE/Parselmouth)
qualityComprehensive provenance documentation: primary sources (app-collected audio, REDCap phenotypes), collection hardware (iPad models, microphone), preprocessing pipeline (b2aiprep → features), software tools (OpenSMILE, Parselmouth, Whisper)
semanticProvenance chain semantically complete - raw capture (iPad/microphone) → storage (BIDS WAV/REDCap) → preprocessing (b2aiprep) → features (Parquet/TSV)
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Collection Mechanisms and Settings Described
evidencecollection_mechanisms: 3 documented - adult data via Bridge2AI-Voice App on iPad (9th/10th gen, Air 5th gen) with Avid AE-36 microphone in clinic with research assistant present, pediatric via reproschema-ui, clinical data via REDCap; collection_timeframes: September 2022-November 2026
qualityDetailed collection settings: specific hardware (iPad models, microphone), software (Bridge2AI-Voice App, reproschema-ui), setting (clinic), personnel (research assistant), timeframe (Sept 2022-Nov 2026)
semanticCollection mechanisms semantically reproducible - hardware specs enable equipment replication, clinic setting and research assistant presence indicate standardized conditions
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidenceacquisition_methods: 4 entries detailing 22 adult acoustic tasks, 36 pediatric tasks, validated questionnaires (VHI-10, PHQ-9, GAD-7, MOCA, etc.), EHR linkage; collection_mechanisms: iPad 9th/10th gen or iPad Air 5th gen with Avid AE-36 microphone, Bridge2AI-Voice App, reproschema-ui for pediatrics; data_collectors: research teams at 5 North American sites, medical graduate/undergraduate students, clinicians as co-investigators, participant compensation $40-120; collection_timeframes: September 2022 - November 2026, semi-annual releases, v1.0 Jan 2025 (306), v3.0 2025 (833)
qualityExceptionally detailed collection protocol with specific hardware, software, personnel, compensation, timeline, and task descriptions
correctnessHardware plausible (iPad with professional microphone standard for clinical voice research); timeline consistent (2022 start to 2026 end, 833 participants accumulated over 3 years across 5 sites)
consistencyCollection details consistent across mechanisms/methods/collectors/timeframes; participant growth (306→833) aligns with timeline progression
acquisition_methods
acquisition_methods:
- id: voice:acquisition:1
  description: '22 acoustic tasks recorded through Bridge2AI-Voice App for adult participants including:
    Non-Voice (Respiration, Cough, Breath Sounds, Voluntary Cough); Voice/Non-Speech (Prolonged Vowel
    /e/, Maximum Phonation Time, Glides, Loudness /Hey/, Diadochokinesis /pa/ta/ka/); Speech (Rainbow
    Passage, Caterpillar Passage, Cape-V Sentences, Free Speech, Picture Description, Story Recall, Animal
    Fluency, Open Response Questions, Word-Color Stroop, Productive Vocabulary, Random Item Generation,
    Cinderella Story). Tasks vary by disease cohort (Part A/Voice/Resp/Mood/Neuro).

    '
- id: voice:acquisition:2
  description: '36 pediatric acoustic tasks collected via reproschema-ui including: Speech tasks (ABC''s,
    Ready for School, Favorite Show, Favorite Food, Outside of School, Months, Counting, Naming Animals,
    Naming Food, Identifying Pictures, Picture Description, Caterpillar Passage, Repeat Words, Role Naming,
    Repeat Sentences); Voice/Non-Speech tasks (Long Sounds, Noisy Sounds, Silly Sounds /PUH TUH KUH/).

    '
- id: voice:acquisition:3
  description: 'Self-reported demographic data and medical history questionnaires administered through
    app. Validated questionnaires integrated for each disease cohort including: VHI-10, PHQ-9, GAD-7,
    PANAS, Custom Affect Scale, PTSD Adult, ADHD Adult, DSM-5 Adult, Dyspnea Index, Leicester Cough Questionnaire,
    Winograd Questionnaire, MOCA, C-VHI-10 (peds), PVOS (peds), PVRQOL (peds), PHQ-A (peds).

    '
- id: voice:acquisition:4
  description: 'Electronic health record (EHR) access for consenting participants permitting investigators
    to access medical information through EHR platforms to perform gold standard validation of diagnoses
    and symptoms. Linkage to multimodal health biomarkers including laryngoscopy imaging, radiomics, genomics,
    respiratory function tests.

    '
✓ 1/1 R10 5.Data Composition and Structure Variable-Level Metadata and Tabular Flag
evidenceis_tabular: false; distribution_formats document TSV phenotype files with JSON data dictionaries, static_features.tsv with static_features.json dictionary; acquisition_methods list 22 adult acoustic tasks, 36 pediatric tasks, validated questionnaires (VHI-10, PHQ-9, GAD-7, etc.)
qualityTabular flag correctly set to false (multimodal: audio + phenotypes); variable documentation via JSON dictionaries; 22 adult tasks + 36 pediatric tasks + 10+ questionnaires enumerated
semanticis_tabular=false semantically accurate for multimodal dataset; variable metadata comprehensive (acoustic features, phenotypes, questionnaires all documented)
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Data Acquisition Methods Listed
evidenceacquisition_methods: 4 documented - 22 adult acoustic tasks (respiration, prolonged vowels, diadochokinesis, passage reading, free speech), 36 pediatric tasks, validated questionnaires (VHI-10, PHQ-9, GAD-7, MOCA, C-VHI-10), EHR access for gold standard validation
qualityComprehensive acquisition documentation: 22 adult tasks enumerated by category (non-voice, voice/non-speech, speech), 36 pediatric tasks named, 10+ validated questionnaires listed, EHR linkage described
semanticAcquisition methods semantically appropriate for voice research - acoustic tasks span breathing (non-voice), phonation (vowels), articulation (diadochokinesis), connected speech (passages), multimodal validation (questionnaires, EHR)
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidenceacquisition_methods: 4 entries detailing 22 adult acoustic tasks, 36 pediatric tasks, validated questionnaires (VHI-10, PHQ-9, GAD-7, MOCA, etc.), EHR linkage; collection_mechanisms: iPad 9th/10th gen or iPad Air 5th gen with Avid AE-36 microphone, Bridge2AI-Voice App, reproschema-ui for pediatrics; data_collectors: research teams at 5 North American sites, medical graduate/undergraduate students, clinicians as co-investigators, participant compensation $40-120; collection_timeframes: September 2022 - November 2026, semi-annual releases, v1.0 Jan 2025 (306), v3.0 2025 (833)
qualityExceptionally detailed collection protocol with specific hardware, software, personnel, compensation, timeline, and task descriptions
correctnessHardware plausible (iPad with professional microphone standard for clinical voice research); timeline consistent (2022 start to 2026 end, 833 participants accumulated over 3 years across 5 sites)
consistencyCollection details consistent across mechanisms/methods/collectors/timeframes; participant growth (306→833) aligns with timeline progression
collection_timeframes
collection_timeframes:
- id: voice:timeframe:1
  description: 'Data collection ongoing from September 2022 through November 2026 (project end date).
    Initial dataset releases began in late 2024. Version 1.0 released January 17, 2025 (306 participants).
    Version 3.0 released 2025 (833 adult participants). Semi-annual releases planned. Enrollment anticipated
    to reach 10,000 participants by 2027.

    '
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Collection Mechanisms and Settings Described
evidencecollection_mechanisms: 3 documented - adult data via Bridge2AI-Voice App on iPad (9th/10th gen, Air 5th gen) with Avid AE-36 microphone in clinic with research assistant present, pediatric via reproschema-ui, clinical data via REDCap; collection_timeframes: September 2022-November 2026
qualityDetailed collection settings: specific hardware (iPad models, microphone), software (Bridge2AI-Voice App, reproschema-ui), setting (clinic), personnel (research assistant), timeframe (Sept 2022-Nov 2026)
semanticCollection mechanisms semantically reproducible - hardware specs enable equipment replication, clinic setting and research assistant presence indicate standardized conditions
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidenceacquisition_methods: 4 entries detailing 22 adult acoustic tasks, 36 pediatric tasks, validated questionnaires (VHI-10, PHQ-9, GAD-7, MOCA, etc.), EHR linkage; collection_mechanisms: iPad 9th/10th gen or iPad Air 5th gen with Avid AE-36 microphone, Bridge2AI-Voice App, reproschema-ui for pediatrics; data_collectors: research teams at 5 North American sites, medical graduate/undergraduate students, clinicians as co-investigators, participant compensation $40-120; collection_timeframes: September 2022 - November 2026, semi-annual releases, v1.0 Jan 2025 (306), v3.0 2025 (833)
qualityExceptionally detailed collection protocol with specific hardware, software, personnel, compensation, timeline, and task descriptions
correctnessHardware plausible (iPad with professional microphone standard for clinical voice research); timeline consistent (2022 start to 2026 end, 833 participants accumulated over 3 years across 5 sites)
consistencyCollection details consistent across mechanisms/methods/collectors/timeframes; participant growth (306→833) aligns with timeline progression
data_collectors
data_collectors:
- id: voice:datacollector:1
  description: 'Research teams at each of the five North American collection sites, including medical
    graduate and undergraduate students coordinating with site clinicians and doctors. Clinicians and
    doctors listed under IRB as co-investigators and added to consortium. Participants compensated via
    electronic gift cards: $40 for sessions under 90 minutes, $80 for sessions over 90 minutes, maximum
    3 sessions and $120 total compensation.

    '
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Vulnerable Populations and Compensation Documented
evidenceat_risk_populations: pediatric participants (aged 2-18) with age-appropriate protocols, distinct IRB considerations, Hospital for Sick Children exclusive collection site, separate dataset release; data_collectors specify $40-120 compensation via electronic gift cards
qualityVulnerable population protections for pediatric cohort: age-appropriate tasks, distinct IRB review, specialized questionnaires, separate data release; compensation structure documented ($40 <90min, $80 >90min, max $120)
semanticPediatric protections semantically appropriate - separate protocols, single trusted institution (SickKids), age-stratified tasks (2-4, 4-6, 6-10, 10+); compensation reasonable and transparent
5/5 R20 Q12 (Technical Documentation) Collection Protocol Clarity
levelFull collection protocol with methods, collectors, and timeframes
evidenceacquisition_methods: 4 entries detailing 22 adult acoustic tasks, 36 pediatric tasks, validated questionnaires (VHI-10, PHQ-9, GAD-7, MOCA, etc.), EHR linkage; collection_mechanisms: iPad 9th/10th gen or iPad Air 5th gen with Avid AE-36 microphone, Bridge2AI-Voice App, reproschema-ui for pediatrics; data_collectors: research teams at 5 North American sites, medical graduate/undergraduate students, clinicians as co-investigators, participant compensation $40-120; collection_timeframes: September 2022 - November 2026, semi-annual releases, v1.0 Jan 2025 (306), v3.0 2025 (833)
qualityExceptionally detailed collection protocol with specific hardware, software, personnel, compensation, timeline, and task descriptions
correctnessHardware plausible (iPad with professional microphone standard for clinical voice research); timeline consistent (2022 start to 2026 end, 833 participants accumulated over 3 years across 5 sites)
consistencyCollection details consistent across mechanisms/methods/collectors/timeframes; participant growth (306→833) aligns with timeline progression
sampling_strategies
sampling_strategies:
- id: voice:sampling:1
  description: 'Non-probability purposive sampling from specialty clinics (high volume expert clinics,
    outpatient clinics seeing >50 patients per month from same disease category). Patients presenting
    at clinics screened for eligibility per inclusion/exclusion criteria outlined in protocol Table 1.
    Not representative of general population due to limited geographic locations and clinic-based recruitment.
    Current v3.0 dataset is a sample of an ongoing collection targeting 10,000 participants by 2027. English-speaking
    adult participants (18-120 years); Spanish protocols under development for future releases.

    '
  is_sample: true
  is_random: false
  is_representative: false
no field-level feedback matched
missing_data_documentation
missing_data_documentation:
- id: voice:missingdata:1
  description: 'Missingness tables generated and included with dataset as part of the audit protocol.
    Audio quality control metrics applied including silence amount, duration, and speech-to-passage accuracy
    checks. Processing using b2aiprep and SenseLab toolkits. Some participants may have incomplete questionnaire
    responses or missing sessions.

    '
✓ 1/1 R10 5.Data Composition and Structure Data Quality Issues and Anomalies Documented
evidencemissing_data_documentation: missingness tables generated via audit protocol, audio quality control (silence, duration, speech-to-passage accuracy); cleaning_strategies: audit protocol with distribution/outlier checks, quality control metrics via b2aiprep and SenseLab
qualityQuality assurance documented: missingness tables included with dataset, audio QC metrics (silence detection, duration validation, accuracy checks), audit protocol for distributions/outliers
semanticQuality documentation semantically appropriate for audio research - silence detection and speech accuracy critical for voice analysis validity
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Data Anomalies and Quality Issues Noted
evidencemissing_data_documentation: missingness tables generated and included, audio quality control metrics (silence, duration, speech-to-passage accuracy), incomplete questionnaire responses noted; cleaning_strategies: audit protocol with distribution/outlier checks
qualityQuality issues proactively documented: missingness tables included with dataset, audio QC metrics applied, incomplete responses acknowledged, outlier detection in audit protocol
semanticAnomaly documentation semantically transparent - missingness quantified (tables included), quality metrics specified (silence/duration/accuracy), audit protocol ensures ongoing quality monitoring
raw_data_sources
raw_data_sources:
- id: voice:rawsource:1
  description: 'Raw audio waveforms in WAV format collected via Bridge2AI-Voice App on iPad devices with
    Avid AE-36 microphone. Stored in BIDS v1.9.0 compliant format per participant session and acoustic
    task. Available through controlled access only.

    '
  source_description: 'Raw audio WAV files recorded directly by participants in clinical settings using
    the Bridge2AI-Voice App on iPad with Avid AE-36 microphone. Organized in BIDS v1.9.0 directory structure:
    sub-{participant_id}/ses-{session_id}/audio/ with companion JSON metadata sidecar files. Available
    through DACO-controlled access only.

    '
- id: voice:rawsource:2
  description: 'REDCap database exports containing EHR-linked clinical data, demographic information,
    disease-specific validated questionnaire responses, and diagnostic information. Pediatric data first
    extracted from reproschema-ui to REDCap format.

    '
  source_description: 'REDCap (Research Electronic Data Capture) database exports and reproschema-ui exports
    containing clinical phenotype data: demographic information, validated questionnaire responses (VHI-10,
    PHQ-9, GAD-7, etc.), diagnostic information, and EHR-linked clinical assessments. Pediatric data first
    extracted from reproschema-ui then converted to REDCap format before BIDS conversion.

    '
✓ 1/1 R10 6.Data Provenance and Version Tracking Provenance and Source Derivation Documented
evidenceraw_data_sources and raw_sources document provenance: raw audio WAV from Bridge2AI-Voice App, REDCap/reproschema-ui for phenotypes; collection_mechanisms specify iPad hardware (9th/10th gen, Air 5th gen) with Avid AE-36 microphone; preprocessing_strategies document transformation pipeline (b2aiprep → spectrograms/MFCCs via FFT/OpenSMILE/Parselmouth)
qualityComprehensive provenance documentation: primary sources (app-collected audio, REDCap phenotypes), collection hardware (iPad models, microphone), preprocessing pipeline (b2aiprep → features), software tools (OpenSMILE, Parselmouth, Whisper)
semanticProvenance chain semantically complete - raw capture (iPad/microphone) → storage (BIDS WAV/REDCap) → preprocessing (b2aiprep) → features (Parquet/TSV)
raw_sources
raw_sources:
- id: voice:rawsrc:1
  description: 'Raw audio WAV files recorded through the Bridge2AI-Voice App, organized in BIDS v1.9.0
    compliant directory structure. Available through controlled access via DACO approval.

    '
- id: voice:rawsrc:2
  description: 'REDCap database exports and reproschema-ui exports containing clinical phenotype data,
    questionnaire responses, and participant demographics.

    '
✓ 1/1 R10 6.Data Provenance and Version Tracking Provenance and Source Derivation Documented
evidenceraw_data_sources and raw_sources document provenance: raw audio WAV from Bridge2AI-Voice App, REDCap/reproschema-ui for phenotypes; collection_mechanisms specify iPad hardware (9th/10th gen, Air 5th gen) with Avid AE-36 microphone; preprocessing_strategies document transformation pipeline (b2aiprep → spectrograms/MFCCs via FFT/OpenSMILE/Parselmouth)
qualityComprehensive provenance documentation: primary sources (app-collected audio, REDCap phenotypes), collection hardware (iPad models, microphone), preprocessing pipeline (b2aiprep → features), software tools (OpenSMILE, Parselmouth, Whisper)
semanticProvenance chain semantically complete - raw capture (iPad/microphone) → storage (BIDS WAV/REDCap) → preprocessing (b2aiprep) → features (Parquet/TSV)
preprocessing_strategies
preprocessing_strategies:
- id: voice:preproc:1
  description: 'Raw audio preprocessing using b2aiprep library: conversion to monaural audio, resampling
    to 16 kHz with Butterworth anti-aliasing filter. Standardization ensures consistent format across
    all recordings for downstream feature extraction.

    '
- id: voice:preproc:2
  description: 'Spectrogram extraction using short-time Fast Fourier Transform (FFT): 25ms window size,
    10ms hop length, 512-point FFT. Output spectrograms have 513xN dimensions where N is proportional
    to audio length. Stored in Parquet format.

    '
- id: voice:preproc:3
  description: 'Mel-frequency cepstral coefficients (MFCC) extraction: 60 MFCCs extracted from spectrograms
    capturing perceptually-relevant spectral envelope characteristics. Output dimension 60xN. Stored in
    Parquet format.

    '
- id: voice:preproc:4
  description: 'OpenSMILE eGeMaps acoustic feature extraction capturing temporal dynamics and acoustic
    characteristics. One row per unique recording in static_features.tsv.

    '
- id: voice:preproc:5
  description: 'Parselmouth and Praat phonetic and prosodic feature computation providing fundamental
    frequency (F0), formants, and voice quality measures. Documented in static_features.json data dictionary.

    '
- id: voice:preproc:6
  description: 'Torchaudio-based feature extraction including pitch contour, spectrograms, mel spectrograms,
    MFCCs. SPARC-based features including electromagnetic articulography (EMA) estimates, loudness, periodicity,
    and pitch measures. Phonetic posteriorgrams (PPGs).

    '
- id: voice:preproc:7
  description: 'Transcription generation using OpenAI Whisper model applied to audio recordings. Raw audio
    transcripts reviewed and any recordings containing potentially identifying information or external
    voices removed. All transcriptions, EMAs, and PPGs from open-response prompts removed from public
    feature-only dataset.

    '
- id: voice:preproc:8
  description: 'REDCap and reproschema-ui data exported and converted to BIDS v1.9.0 format using b2aiprep
    open-source library. Phenotype data organized in tab-delimited files with JSON data dictionaries.
    Pediatric data first extracted from reproschema-ui to REDCap format then converted to BIDS.

    '
✓ 1/1 R10 6.Data Provenance and Version Tracking Provenance and Source Derivation Documented
evidenceraw_data_sources and raw_sources document provenance: raw audio WAV from Bridge2AI-Voice App, REDCap/reproschema-ui for phenotypes; collection_mechanisms specify iPad hardware (9th/10th gen, Air 5th gen) with Avid AE-36 microphone; preprocessing_strategies document transformation pipeline (b2aiprep → spectrograms/MFCCs via FFT/OpenSMILE/Parselmouth)
qualityComprehensive provenance documentation: primary sources (app-collected audio, REDCap phenotypes), collection hardware (iPad models, microphone), preprocessing pipeline (b2aiprep → features), software tools (OpenSMILE, Parselmouth, Whisper)
semanticProvenance chain semantically complete - raw capture (iPad/microphone) → storage (BIDS WAV/REDCap) → preprocessing (b2aiprep) → features (Parquet/TSV)
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Preprocessing, Cleaning, and Labeling Strategies
evidencepreprocessing_strategies: 8 documented (b2aiprep monaural/16kHz conversion, spectrogram via STFT 25ms/10ms/512pt, 60 MFCCs, OpenSMILE eGeMaps, Parselmouth/Praat F0/formants, torchaudio features, Whisper transcription, BIDS conversion); cleaning_strategies: 4 (HIPAA Safe Harbor, audio privacy, sensitive field removal, audit protocol); labeling_strategies: clinician diagnosis via gold standard validation
qualityDetailed preprocessing pipeline with technical parameters (16kHz sampling, 25ms window, 512pt FFT, 60 MFCCs), cleaning via HIPAA Safe Harbor and audit protocol, labeling by site clinicians per ICD-10
semanticProcessing strategies semantically complete and reproducible - specific parameters enable exact pipeline replication, HIPAA Safe Harbor provides standardized de-identification, clinician labeling ensures diagnostic validity
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Software and Tools Documented
evidencepreprocessing_strategies and machine_annotation_tools reference b2aiprep, SenseLab, OpenSMILE, Parselmouth, Praat, torchaudio, OpenAI Whisper; external_resources link GitHub repos: b2aiprep (github.com/sensein/b2aiprep), Bridge2AI-redcap (github.com/eipm/bridge2ai-redcap), documentation (github.com/eipm/bridge2ai-docs)
qualitySoftware tools named with specific functions: b2aiprep (preprocessing), OpenSMILE (eGeMaps features), Parselmouth/Praat (phonetic features), torchaudio (audio processing), Whisper (transcription); GitHub repos linked for reproducibility
semanticSoftware documentation semantically complete - tools specified with versions/licenses where applicable (b2aiprep Apache-2.0, REDCap dictionary MIT), GitHub links enable code access
4/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelComprehensive strategies with partial tool documentation
evidencepreprocessing_strategies: 8 entries mentioning b2aiprep, SenseLab, OpenSMILE, Parselmouth, Praat, torchaudio, OpenAI Whisper with specific parameters (16kHz resampling, 25ms window, 512-point FFT, 60 MFCCs); cleaning_strategies: 4 entries detailing HIPAA Safe Harbor, audio privacy, REDCap field removal, audit protocol; labeling_strategies: site clinician assignment per ICD-10; machine_annotation_tools: Whisper, b2aiprep, SenseLab; keywords list OpenSMILE, Praat, Parselmouth, torchaudio, b2aiprep
qualityExcellent strategy documentation with specific tools and parameters; minor gap: software versions/URLs not in D4D-core schema fields (present in external_resources though: GitHub links for b2aiprep, bridge2ai-redcap)
correctnessTools semantically appropriate: OpenSMILE (audio features), Whisper (transcription), b2aiprep (BIDS conversion), Parselmouth/Praat (phonetics), torchaudio (PyTorch audio)
consistencyTool mentions consistent across preprocessing_strategies, machine_annotation_tools, keywords, and external_resources
cleaning_strategies
cleaning_strategies:
- id: voice:cleaning:1
  description: 'HIPAA Safe Harbor de-identification for public release: removed direct identifiers (names,
    civic addresses, social security numbers), indirect identifiers creating significant re-identification
    risk (select geographic/demographic identifiers, household composition, cultural identity), and sensitive
    information (household income, mental health status, traumatic life experiences). Geographic data
    limited; state/province removed, country retained.

    '
- id: voice:cleaning:2
  description: 'Audio privacy protection for public release: all raw audio waveforms excluded from public
    dataset. All spectrograms, MFCCs, mel spectrograms, transcriptions, EMAs, and PPGs from open-response
    prompts removed from feature-only dataset. Free speech transcripts removed. Raw audio available only
    through controlled access with DACO approval.

    '
- id: voice:cleaning:3
  description: 'Sensitive field removal based on REDCap data dictionary: all fields encoded as sensitive
    (column "Identifier?" in REDCap data dictionary CSV) removed from dataset.

    '
- id: voice:cleaning:4
  description: 'Audit protocol applied: missingness tables generated and included with dataset; distribution
    and outlier checks; categorical responses checked against schema; audio quality control metrics including
    silence amount, duration, and speech-to-passage accuracy checks. Processing using b2aiprep and SenseLab
    toolkits.

    '
✓ 1/1 R10 5.Data Composition and Structure Data Quality Issues and Anomalies Documented
evidencemissing_data_documentation: missingness tables generated via audit protocol, audio quality control (silence, duration, speech-to-passage accuracy); cleaning_strategies: audit protocol with distribution/outlier checks, quality control metrics via b2aiprep and SenseLab
qualityQuality assurance documented: missingness tables included with dataset, audio QC metrics (silence detection, duration validation, accuracy checks), audit protocol for distributions/outliers
semanticQuality documentation semantically appropriate for audio research - silence detection and speech accuracy critical for voice analysis validity
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Preprocessing, Cleaning, and Labeling Strategies
evidencepreprocessing_strategies: 8 documented (b2aiprep monaural/16kHz conversion, spectrogram via STFT 25ms/10ms/512pt, 60 MFCCs, OpenSMILE eGeMaps, Parselmouth/Praat F0/formants, torchaudio features, Whisper transcription, BIDS conversion); cleaning_strategies: 4 (HIPAA Safe Harbor, audio privacy, sensitive field removal, audit protocol); labeling_strategies: clinician diagnosis via gold standard validation
qualityDetailed preprocessing pipeline with technical parameters (16kHz sampling, 25ms window, 512pt FFT, 60 MFCCs), cleaning via HIPAA Safe Harbor and audit protocol, labeling by site clinicians per ICD-10
semanticProcessing strategies semantically complete and reproducible - specific parameters enable exact pipeline replication, HIPAA Safe Harbor provides standardized de-identification, clinician labeling ensures diagnostic validity
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Data Anomalies and Quality Issues Noted
evidencemissing_data_documentation: missingness tables generated and included, audio quality control metrics (silence, duration, speech-to-passage accuracy), incomplete questionnaire responses noted; cleaning_strategies: audit protocol with distribution/outlier checks
qualityQuality issues proactively documented: missingness tables included with dataset, audio QC metrics applied, incomplete responses acknowledged, outlier detection in audit protocol
semanticAnomaly documentation semantically transparent - missingness quantified (tables included), quality metrics specified (silence/duration/accuracy), audit protocol ensures ongoing quality monitoring
4/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelComprehensive strategies with partial tool documentation
evidencepreprocessing_strategies: 8 entries mentioning b2aiprep, SenseLab, OpenSMILE, Parselmouth, Praat, torchaudio, OpenAI Whisper with specific parameters (16kHz resampling, 25ms window, 512-point FFT, 60 MFCCs); cleaning_strategies: 4 entries detailing HIPAA Safe Harbor, audio privacy, REDCap field removal, audit protocol; labeling_strategies: site clinician assignment per ICD-10; machine_annotation_tools: Whisper, b2aiprep, SenseLab; keywords list OpenSMILE, Praat, Parselmouth, torchaudio, b2aiprep
qualityExcellent strategy documentation with specific tools and parameters; minor gap: software versions/URLs not in D4D-core schema fields (present in external_resources though: GitHub links for b2aiprep, bridge2ai-redcap)
correctnessTools semantically appropriate: OpenSMILE (audio features), Whisper (transcription), b2aiprep (BIDS conversion), Parselmouth/Praat (phonetics), torchaudio (PyTorch audio)
consistencyTool mentions consistent across preprocessing_strategies, machine_annotation_tools, keywords, and external_resources
labeling_strategies
labeling_strategies:
- id: voice:labeling:1
  description: 'Diagnostic labels assigned by site clinicians based on clinical assessment and gold standard
    validation methods per Bridge2AI-Voice protocol Table 1 and ICD-10 codes. For voice disorders: laryngoscopy
    and stroboscopy. For neurological disorders: CT Brain, MRI Brain, genome sequencing, serum biomarkers.
    For mood/psychiatric disorders: EHR records, psychiatrist diagnosis, medication history. For respiratory
    disorders: spirometry, flow volume loops, CT scan. Single labeler (site clinician) per participant.

    '
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Preprocessing, Cleaning, and Labeling Strategies
evidencepreprocessing_strategies: 8 documented (b2aiprep monaural/16kHz conversion, spectrogram via STFT 25ms/10ms/512pt, 60 MFCCs, OpenSMILE eGeMaps, Parselmouth/Praat F0/formants, torchaudio features, Whisper transcription, BIDS conversion); cleaning_strategies: 4 (HIPAA Safe Harbor, audio privacy, sensitive field removal, audit protocol); labeling_strategies: clinician diagnosis via gold standard validation
qualityDetailed preprocessing pipeline with technical parameters (16kHz sampling, 25ms window, 512pt FFT, 60 MFCCs), cleaning via HIPAA Safe Harbor and audit protocol, labeling by site clinicians per ICD-10
semanticProcessing strategies semantically complete and reproducible - specific parameters enable exact pipeline replication, HIPAA Safe Harbor provides standardized de-identification, clinician labeling ensures diagnostic validity
4/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelComprehensive strategies with partial tool documentation
evidencepreprocessing_strategies: 8 entries mentioning b2aiprep, SenseLab, OpenSMILE, Parselmouth, Praat, torchaudio, OpenAI Whisper with specific parameters (16kHz resampling, 25ms window, 512-point FFT, 60 MFCCs); cleaning_strategies: 4 entries detailing HIPAA Safe Harbor, audio privacy, REDCap field removal, audit protocol; labeling_strategies: site clinician assignment per ICD-10; machine_annotation_tools: Whisper, b2aiprep, SenseLab; keywords list OpenSMILE, Praat, Parselmouth, torchaudio, b2aiprep
qualityExcellent strategy documentation with specific tools and parameters; minor gap: software versions/URLs not in D4D-core schema fields (present in external_resources though: GitHub links for b2aiprep, bridge2ai-redcap)
correctnessTools semantically appropriate: OpenSMILE (audio features), Whisper (transcription), b2aiprep (BIDS conversion), Parselmouth/Praat (phonetics), torchaudio (PyTorch audio)
consistencyTool mentions consistent across preprocessing_strategies, machine_annotation_tools, keywords, and external_resources
machine_annotation_tools
machine_annotation_tools:
- id: voice:machanno:1
  description: 'OpenAI Whisper model used for automated transcription of audio recordings. Transcripts
    reviewed for quality; recordings with potentially identifying information or external voices removed.
    Transcriptions from open-response prompts removed from public release.

    '
- id: voice:machanno:2
  description: 'b2aiprep open-source library used for automated preprocessing of raw audio, REDCap data
    conversion, and BIDS formatting. SenseLab toolkit used for audio quality control and feature extraction.

    '
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Software and Tools Documented
evidencepreprocessing_strategies and machine_annotation_tools reference b2aiprep, SenseLab, OpenSMILE, Parselmouth, Praat, torchaudio, OpenAI Whisper; external_resources link GitHub repos: b2aiprep (github.com/sensein/b2aiprep), Bridge2AI-redcap (github.com/eipm/bridge2ai-redcap), documentation (github.com/eipm/bridge2ai-docs)
qualitySoftware tools named with specific functions: b2aiprep (preprocessing), OpenSMILE (eGeMaps features), Parselmouth/Praat (phonetic features), torchaudio (audio processing), Whisper (transcription); GitHub repos linked for reproducibility
semanticSoftware documentation semantically complete - tools specified with versions/licenses where applicable (b2aiprep Apache-2.0, REDCap dictionary MIT), GitHub links enable code access
4/5 R20 Q11 (Technical Documentation) Tool and Software Transparency
levelComprehensive strategies with partial tool documentation
evidencepreprocessing_strategies: 8 entries mentioning b2aiprep, SenseLab, OpenSMILE, Parselmouth, Praat, torchaudio, OpenAI Whisper with specific parameters (16kHz resampling, 25ms window, 512-point FFT, 60 MFCCs); cleaning_strategies: 4 entries detailing HIPAA Safe Harbor, audio privacy, REDCap field removal, audit protocol; labeling_strategies: site clinician assignment per ICD-10; machine_annotation_tools: Whisper, b2aiprep, SenseLab; keywords list OpenSMILE, Praat, Parselmouth, torchaudio, b2aiprep
qualityExcellent strategy documentation with specific tools and parameters; minor gap: software versions/URLs not in D4D-core schema fields (present in external_resources though: GitHub links for b2aiprep, bridge2ai-redcap)
correctnessTools semantically appropriate: OpenSMILE (audio features), Whisper (transcription), b2aiprep (BIDS conversion), Parselmouth/Praat (phonetics), torchaudio (PyTorch audio)
consistencyTool mentions consistent across preprocessing_strategies, machine_annotation_tools, keywords, and external_resources
intended_uses
intended_uses:
- id: voice:use:1
  description: 'Primary intended use: development and validation of AI/ML models for voice as a biomarker
    of health, supporting screening, diagnosis, and treatment of voice disorders, neurological disorders,
    mood disorders, respiratory disorders, and pediatric speech disorders through model pretraining, fine-tuning,
    benchmarking, and validation.

    '
- id: voice:use:2
  description: 'Research into acoustic biomarkers and development of standards for voice data collection
    and analysis. Establishing best practices for AI/ML-friendly voice datasets and contributing to the
    field''s maturation as a clinical diagnostic modality.

    '
- id: voice:use:3
  description: 'Training and education in voice AI research through workforce development initiatives,
    curriculum creation on FAIR and CARE voice AI model development, and fostering collaborations between
    medical voice researchers, acoustic engineers, and AI/ML specialists, especially from underserved
    communities.

    '
- id: voice:use:4
  description: 'Multimodal health research combining voice data with EHR information, radiomics, genomics,
    imaging, and other health biomarkers to understand complex disease relationships and improve diagnostic
    accuracy.

    '
✓ 1/1 R10 3.Data Reuse and Interoperability Use Guidance Provided (intended, prohibited uses)
evidenceintended_uses: 4 documented uses including AI/ML model development, standards research, training/education, multimodal health research; discouraged_uses: 3 items (non-clinical discrimination, re-identification, IP barriers); prohibited_uses: unauthorized research, data sale, re-identification attempts
qualityExplicit use guidance across three categories with specific examples: intended (ML models, standards), discouraged (hiring decisions, surveillance), prohibited (data sale, re-identification)
semanticUse guidance semantically coherent with dataset purpose - AI/ML development encouraged, discriminatory applications prohibited
existing_uses
existing_uses:
- id: voice:existinguse:1
  description: 'A restricted version of the dataset containing raw audio has been used in the Bridge2AI
    Summer School and hackathon for education and research training purposes.

    '
no field-level feedback matched
discouraged_uses
discouraged_uses:
- id: voice:discouraged:1
  description: 'Non-clinical applications such as hiring decisions, insurance premium adjustments, or
    any form of surveillance that could lead to discrimination or harm based on health conditions or voice
    characteristics. These applications could negatively impact individuals.

    '
- id: voice:discouraged:2
  description: 'Any attempt to re-identify research participants or use data in ways that could foreseeably
    cause harm or stigmatization to research participants, their families, communities, or specific populations.
    Dataset covered under Certificate of Confidentiality.

    '
- id: voice:discouraged:3
  description: 'Development or use of intellectual property protections, database rights, or related rights
    in ways that would prevent or block access to any element of the dataset or conclusions derived from
    it. Must respect Fort Lauderdale Agreement and Open Science principles.

    '
✓ 1/1 R10 3.Data Reuse and Interoperability Use Guidance Provided (intended, prohibited uses)
evidenceintended_uses: 4 documented uses including AI/ML model development, standards research, training/education, multimodal health research; discouraged_uses: 3 items (non-clinical discrimination, re-identification, IP barriers); prohibited_uses: unauthorized research, data sale, re-identification attempts
qualityExplicit use guidance across three categories with specific examples: intended (ML models, standards), discouraged (hiring decisions, surveillance), prohibited (data sale, re-identification)
semanticUse guidance semantically coherent with dataset purpose - AI/ML development encouraged, discriminatory applications prohibited
prohibited_uses
prohibited_uses:
- id: voice:prohibited:1
  description: 'Use of the dataset outside of authorized research purposes as defined in the Bridge2AI
    Voice Registered Access Agreement. The dataset is intended solely for commercial and non-commercial
    research by Authorized Researchers. Sale of all or part of the data on any media is prohibited. Re-identification
    attempts are prohibited.

    '
✓ 1/1 R10 3.Data Reuse and Interoperability Use Guidance Provided (intended, prohibited uses)
evidenceintended_uses: 4 documented uses including AI/ML model development, standards research, training/education, multimodal health research; discouraged_uses: 3 items (non-clinical discrimination, re-identification, IP barriers); prohibited_uses: unauthorized research, data sale, re-identification attempts
qualityExplicit use guidance across three categories with specific examples: intended (ML models, standards), discouraged (hiring decisions, surveillance), prohibited (data sale, re-identification)
semanticUse guidance semantically coherent with dataset purpose - AI/ML development encouraged, discriminatory applications prohibited
future_use_impacts
future_use_impacts:
- id: voice:futureimpact:1
  description: 'Potential risks from future uses include voice re-identification despite de-identification
    efforts, discriminatory application of voice biomarker models in clinical or non-clinical settings,
    and reinforcement of health disparities if models are trained on non-representative data. The Certificate
    of Confidentiality and access controls mitigate some of these risks.

    '
- id: voice:futureimpact:2
  description: 'Positive anticipated impacts include acceleration of clinical voice AI adoption for screening
    and diagnosis of conditions currently lacking cost-effective diagnostic tools, improved health equity
    through diverse data collection, and establishment of FAIR and CARE-compliant data standards for the
    voice AI field.

    '
no field-level feedback matched
distribution_formats
distribution_formats:
- id: voice:format:1
  name: Parquet format for spectrograms and derived features
  description: 'Time-varying features (spectrograms, mel spectrograms, MFCCs, pitch contour, SPARC features,
    PPGs) stored in Parquet format compatible with Python datasets library. Each element contains participant_id,
    session_id, task_name, and feature arrays. Excluded for open-response tasks in public release.

    '
- id: voice:format:2
  name: TSV/JSON for phenotype and static features
  description: 'Phenotype data in tab-delimited format (confounders.tsv, demographics.tsv, diagnosis/
    *.tsv, enrollment/*.tsv, questionnaire/*.tsv, task/*.tsv) with JSON data dictionaries per file. Static
    acoustic features (OpenSMILE, Praat, Parselmouth, torchaudio) in static_features.tsv with static_features.json
    data dictionary. One row per participant or recording as appropriate. BIDS v1.9.0 compliant structure.

    '
- id: voice:format:3
  name: WAV audio format (controlled access only)
  description: 'Original raw audio waveforms in WAV format following BIDS structure: sub-{participant_id}/ses-{session_id}/audio/sub{id}_ses{id}_task-{task_name}.wav
    with companion JSON metadata. Available through controlled access only. Contact DACO@b2ai-voice.org
    to request access.

    '
✓ 1/1 R10 2.Dataset Access and Retrieval Distribution Formats and File Types Specified
evidencedistribution_formats: 3 documented formats - Parquet for spectrograms/features, TSV/JSON for phenotype data, WAV for raw audio; distributions specify media_type: text/tab-separated-values, application/json, application/gzip
qualitySpecific file formats with technical details (Parquet dimensions, BIDS structure, TSV/JSON pairing) and MIME types for each distribution
semanticFormats semantically appropriate for data types - Parquet for time-series features, TSV for tabular phenotypes, WAV for audio
✓ 1/1 R10 3.Data Reuse and Interoperability Data Formats Are Standardized (encoding, format)
evidencedistribution_formats specify standard formats: Parquet (Python datasets library compatible), TSV (BIDS v1.9.0 compliant), JSON data dictionaries, WAV audio; language: en; BIDS v1.9.0 structure throughout
qualityStandards-compliant formats: BIDS v1.9.0 for directory structure and metadata, Parquet for ML-friendly time-series, TSV/JSON pairing for interoperability
semanticFormat choices semantically AI/ML-friendly - Parquet for efficient feature loading, BIDS for neuroimaging community compatibility
✓ 1/1 R10 3.Data Reuse and Interoperability Variable Metadata with Identifiers Defined
evidencedistribution_formats reference JSON data dictionaries accompanying all TSV files; static_features.json data dictionary for acoustic features; REDCap data dictionary available at github.com/eipm/bridge2ai-redcap
qualityComprehensive variable-level metadata: JSON data dictionaries per TSV file, static_features.json for acoustic measures, publicly accessible REDCap dictionary
semanticVariable metadata semantically complete - field names, types, allowed values documented; REDCap dictionary provides source-level documentation
✓ 1/1 R10 5.Data Composition and Structure Variable-Level Metadata and Tabular Flag
evidenceis_tabular: false; distribution_formats document TSV phenotype files with JSON data dictionaries, static_features.tsv with static_features.json dictionary; acquisition_methods list 22 adult acoustic tasks, 36 pediatric tasks, validated questionnaires (VHI-10, PHQ-9, GAD-7, etc.)
qualityTabular flag correctly set to false (multimodal: audio + phenotypes); variable documentation via JSON dictionaries; 22 adult tasks + 36 pediatric tasks + 10+ questionnaires enumerated
semanticis_tabular=false semantically accurate for multimodal dataset; variable metadata comprehensive (acoustic features, phenotypes, questionnaires all documented)
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path (platform, login, policy)
evidencedistribution_formats: PhysioNet registered access for features (TSV/JSON/Parquet), controlled access for raw audio (WAV/GZ via DACO); license_and_use_terms: registered users sign Bridge2AI Voice Registered Access Agreement, raw audio requires DACO application + institutional DTUA; download_url: https://physionet.org/content/b2ai-voice/; external_resources: DACO contact mailto:DACO@b2ai-voice.org
qualityExceptionally clear two-tier access mechanism: (1) PhysioNet registered access with DUA for features, (2) DACO-controlled access with institutional DTUA for raw audio biometrics
correctnessAccess tiers semantically appropriate for biometric data: lower barrier for de-identified features, higher barrier with institutional oversight for re-identifiable raw audio
consistencyAccess mechanism consistent across distribution_formats, license_and_use_terms, and confidential_elements documentation
5/5 R20 Q4 (Structural Completeness) File Enumeration and Type Variety
level>3 file types
evidencedistribution_formats: 4 distinct formats - Parquet (spectrograms, MFCCs), TSV (phenotype data), JSON (data dictionaries), WAV (raw audio controlled access), GZ (compressed raw audio archives)
qualityExcellent format diversity appropriate for multimodal voice dataset with time-series features and clinical phenotypes
correctnessFormat choices semantically appropriate: Parquet for large time-series arrays, TSV/JSON for BIDS compliance, WAV for audio preservation
consistencyFormat distribution aligns with access tiers (registered vs controlled)
distribution_dates
distribution_dates:
- id: voice:distdate:1
  description: 'Dataset first published and made available late November 2024 through Health Data Nexus.
    PhysioNet releases: v1.0 January 17, 2025 (306 participants, 12,523 recordings); v1.1 January 17,
    2025 (added MFCC features); v2.0.0 April 16, 2025; v2.0.1 August 18, 2025; v3.0.0 released 2025 (833
    adult participants, ~61,937 recordings). Semi-annual releases planned. Pediatric v1.0 available separately.

    '
✓ 1/1 R10 6.Data Provenance and Version Tracking Dataset Version Number Provided
evidenceversion: 3.0.0; distribution_dates and updates document version progression: v1.0 (Jan 17, 2025, 306 participants), v1.1 (Jan 17, 2025, added MFCCs), v2.0.0 (Apr 16, 2025), v2.0.1 (Aug 18, 2025), v3.0.0 (2025, 833 participants)
qualityCurrent version (3.0.0) clearly stated, comprehensive version history with release dates and participant counts, semantic versioning pattern followed
semanticVersion numbering semantically meaningful - major version increments (1.0→2.0→3.0) correspond to significant participant count increases
✓ 1/1 R10 6.Data Provenance and Version Tracking Change Descriptions and Errata Provided
evidenceupdates and distribution_dates document version changes: v1.0 (306 participants, 12,523 recordings), v1.1 (added MFCC features), v2.0.0, v2.0.1, v3.0.0 (833 adults, ~61,937 recordings); users notified through platform news items
qualityVersion change descriptions provided: participant count increases, feature additions (MFCCs in v1.1), notification mechanism (platform news items); pediatric v1.0 as separate release
semanticChange documentation semantically informative - quantitative details (participant counts, recording counts, feature types) enable impact assessment
✓ 1/1 R10 6.Data Provenance and Version Tracking Update Schedule or Frequency Indicated
evidenceupdates: semi-annual releases planned, data collection ongoing through November 30, 2026, target 10,000 participants by 2027; distribution_dates show Jan-Aug 2025 releases (v1.0, v1.1, v2.0.0, v2.0.1, v3.0.0)
qualityUpdate schedule explicitly stated: semi-annual releases, collection end date (Nov 30, 2026), enrollment target (10,000 by 2027), observed release pattern confirms semi-annual cadence
semanticUpdate frequency semantically consistent with observed release pattern - 5 versions in ~8 months (Jan-Aug 2025) aligns with semi-annual commitment
5/5 R20 Q13 (Technical Documentation) Version History Documentation
levelComprehensive versioning with errata, updates, and release notes
evidenceversion: 3.0.0; version_access: all versions available via PhysioNet with version-specific DOIs, older versions on Health Data Nexus; updates: semi-annual releases, v1.0 Jan 17 2025 (306 participants, 12,523 recordings), v1.1 added MFCC features, v2.0.0 Apr 16 2025, v2.0.1 Aug 18 2025, v3.0.0 2025 (833 adults, ~61,937 recordings), pediatric v1.0 separate, target 10,000 by 2027; distribution_dates: detailed release timeline with participant counts per version
qualityExemplary version tracking with specific release dates, participant/recording counts per version, version-specific DOIs, and clear update roadmap
correctnessVersion progression plausible (306→833 participants over 1 year with semi-annual releases); version numbering follows semantic versioning (1.0→1.1→2.0.0→2.0.1→3.0.0)
consistencyVersion information consistent across version field, updates, version_access, and distribution_dates
5/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelStructured version control with timestamps
evidenceupdates: detailed version progression with timestamps (v1.0 Jan 17 2025, v1.1 Jan 17 2025, v2.0.0 Apr 16 2025, v2.0.1 Aug 18 2025, v3.0.0 2025) and change descriptions (v1.1 added MFCC features), participant counts per version, semi-annual release schedule; version_access: all versions maintained with version-specific DOIs; distribution_dates: release timeline documented
qualityExemplary provenance tracking with version-specific DOIs, timestamped releases, participant count evolution, and feature addition logs
correctnessVersion dates chronologically consistent; version numbering follows semantic versioning conventions; participant growth plausible (306→833 over ~1 year)
consistencyProvenance information consistent across updates, version_access, and distribution_dates fields
distributions
distributions:
- id: voice:dist:1
  description: 'Featurized dataset (registered access) distributed through PhysioNet. Contains AI-ready
    derived features: OpenSMILE eGeMaps, Parselmouth/Praat speech features, torchaudio features, spectrograms,
    mel spectrograms, MFCCs, SPARC features, PPGs. Phenotype data in BIDS v1.9.0 compliant TSV/JSON format.
    HIPAA Safe Harbor de-identified. Parquet files used for time-varying features (spectrograms, MFCCs,
    pitch contour).

    '
  format: TSV
  media_type: text/tab-separated-values
  path: https://physionet.org/content/b2ai-voice/
- id: voice:dist:2
  description: 'JSON data dictionaries and metadata distributed alongside TSV phenotype files. Each TSV
    file has a companion JSON data dictionary describing field names, types, and allowed values. Follows
    BIDS v1.9.0 specification.

    '
  format: JSON
  media_type: application/json
  path: https://physionet.org/content/b2ai-voice/
- id: voice:dist:3
  description: 'Raw audio dataset (controlled access only) distributed through PhysioNet and Health Data
    Nexus. Contains original raw audio waveforms compressed in GZ archives following BIDS v1.9.0 structure.
    Requires DACO approval and institutional DTUA. Contact DACO@b2ai-voice.org to request access.

    '
  format: GZ
  media_type: application/gzip
  path: mailto:DACO@b2ai-voice.org
- id: voice:dist:4
  description: 'Pediatric dataset v1.0 (registered access) distributed through PhysioNet as separate download.
    Contains data from 300 pediatric participants with 36 pediatric-specific acoustic tasks and specialized
    questionnaires. BIDS v1.9.0 compliant TSV/JSON format.

    '
  format: TSV
  media_type: text/tab-separated-values
  path: https://physionet.org/content/b2ai-voice/
✓ 1/1 R10 10.Cross-Platform and Community Integration Dataset Published on a Recognized Platform
evidencepublisher: Bridge2AI-Voice Consortium, University of South Florida; distributions via PhysioNet (MIT Laboratory for Computational Physiology, NIBIB-funded); Health Data Nexus (University of Toronto T-CAIREM) as alternative platform
qualityPublished on multiple recognized platforms: PhysioNet (NIH-funded biomedical data repository), Health Data Nexus (T-CAIREM cloud research platform), both with established reputations in biomedical research community
semanticPlatform choice semantically appropriate - PhysioNet specialized in physiological signals/biomedical time-series, Health Data Nexus provides cloud compute integration
✓ 1/1 R10 2.Dataset Access and Retrieval Download URL or Platform Link Available
evidencedownload_url: https://physionet.org/content/b2ai-voice/; distributions include paths to PhysioNet and DACO contact (DACO@b2ai-voice.org) for controlled access
qualityDirect download URL for registered access dataset, specific contact mechanism (email) for controlled access raw audio, alternative platform (Health Data Nexus) documented
semanticDownload mechanisms semantically differentiated by access tier - public URL for registered, email application for controlled
✓ 1/1 R10 2.Dataset Access and Retrieval Distribution Formats and File Types Specified
evidencedistribution_formats: 3 documented formats - Parquet for spectrograms/features, TSV/JSON for phenotype data, WAV for raw audio; distributions specify media_type: text/tab-separated-values, application/json, application/gzip
qualitySpecific file formats with technical details (Parquet dimensions, BIDS structure, TSV/JSON pairing) and MIME types for each distribution
semanticFormats semantically appropriate for data types - Parquet for time-series features, TSV for tabular phenotypes, WAV for audio
maintainers
maintainers:
- id: voice:maintainer:1
  description: 'Bridge2AI-Voice Consortium (University of South Florida, lead institution) supported by
    NIH Bridge2AI program. Contact: [email protected]. Responsible for dataset curation, standards development,
    ethics oversight, versioning, and updates. Data Access Compliance Office (DACO) manages controlled
    access applications. Contact: DACO@b2ai-voice.org.

    '
- id: voice:maintainer:2
  description: 'PhysioNet (MIT Laboratory for Computational Physiology, supported by NIBIB NIH grant R01EB030362)
    serves as primary distribution platform for registered access dataset. Provides technical infrastructure
    and access management.

    '
- id: voice:maintainer:3
  description: 'Health Data Nexus (Temerty Centre for Artificial Intelligence Research and Education in
    Medicine, T-CAIREM, University of Toronto) maintains alternative platform providing cloud compute
    alongside dataset. Earlier dataset versions available here. Contact: [email protected].

    '
no field-level feedback matched
updates
updates:
  id: voice:updates:1
  name: Versioned releases with ongoing data collection
  description: 'Dataset updated with versioned static releases semi-annually as data collection progresses.
    Users notified through news items on platforms and standard communication channels. v1.0 released
    January 17, 2025 (306 participants, 12,523 recordings); v1.1 added MFCC features; v2.0.0 April 16,
    2025; v2.0.1 August 18, 2025; v3.0.0 released 2025 (833 adults, ~61,937 recordings). Pediatric v1.0
    released separately. Data collection ongoing through November 30, 2026. Target: 10,000 participants
    by 2027. Future releases will expand Spanish language protocols and add additional multimodal data
    (imaging, genomics). Version-specific DOIs maintained. Older versions continue to be supported and
    hosted.

    '
✓ 1/1 R10 6.Data Provenance and Version Tracking Dataset Version Number Provided
evidenceversion: 3.0.0; distribution_dates and updates document version progression: v1.0 (Jan 17, 2025, 306 participants), v1.1 (Jan 17, 2025, added MFCCs), v2.0.0 (Apr 16, 2025), v2.0.1 (Aug 18, 2025), v3.0.0 (2025, 833 participants)
qualityCurrent version (3.0.0) clearly stated, comprehensive version history with release dates and participant counts, semantic versioning pattern followed
semanticVersion numbering semantically meaningful - major version increments (1.0→2.0→3.0) correspond to significant participant count increases
✓ 1/1 R10 6.Data Provenance and Version Tracking Change Descriptions and Errata Provided
evidenceupdates and distribution_dates document version changes: v1.0 (306 participants, 12,523 recordings), v1.1 (added MFCC features), v2.0.0, v2.0.1, v3.0.0 (833 adults, ~61,937 recordings); users notified through platform news items
qualityVersion change descriptions provided: participant count increases, feature additions (MFCCs in v1.1), notification mechanism (platform news items); pediatric v1.0 as separate release
semanticChange documentation semantically informative - quantitative details (participant counts, recording counts, feature types) enable impact assessment
✓ 1/1 R10 6.Data Provenance and Version Tracking Update Schedule or Frequency Indicated
evidenceupdates: semi-annual releases planned, data collection ongoing through November 30, 2026, target 10,000 participants by 2027; distribution_dates show Jan-Aug 2025 releases (v1.0, v1.1, v2.0.0, v2.0.1, v3.0.0)
qualityUpdate schedule explicitly stated: semi-annual releases, collection end date (Nov 30, 2026), enrollment target (10,000 by 2027), observed release pattern confirms semi-annual cadence
semanticUpdate frequency semantically consistent with observed release pattern - 5 versions in ~8 months (Jan-Aug 2025) aligns with semi-annual commitment
5/5 R20 Q13 (Technical Documentation) Version History Documentation
levelComprehensive versioning with errata, updates, and release notes
evidenceversion: 3.0.0; version_access: all versions available via PhysioNet with version-specific DOIs, older versions on Health Data Nexus; updates: semi-annual releases, v1.0 Jan 17 2025 (306 participants, 12,523 recordings), v1.1 added MFCC features, v2.0.0 Apr 16 2025, v2.0.1 Aug 18 2025, v3.0.0 2025 (833 adults, ~61,937 recordings), pediatric v1.0 separate, target 10,000 by 2027; distribution_dates: detailed release timeline with participant counts per version
qualityExemplary version tracking with specific release dates, participant/recording counts per version, version-specific DOIs, and clear update roadmap
correctnessVersion progression plausible (306→833 participants over 1 year with semi-annual releases); version numbering follows semantic versioning (1.0→1.1→2.0.0→2.0.1→3.0.0)
consistencyVersion information consistent across version field, updates, version_access, and distribution_dates
5/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelStructured version control with timestamps
evidenceupdates: detailed version progression with timestamps (v1.0 Jan 17 2025, v1.1 Jan 17 2025, v2.0.0 Apr 16 2025, v2.0.1 Aug 18 2025, v3.0.0 2025) and change descriptions (v1.1 added MFCC features), participant counts per version, semi-annual release schedule; version_access: all versions maintained with version-specific DOIs; distribution_dates: release timeline documented
qualityExemplary provenance tracking with version-specific DOIs, timestamped releases, participant count evolution, and feature addition logs
correctnessVersion dates chronologically consistent; version numbering follows semantic versioning conventions; participant growth plausible (306→833 over ~1 year)
consistencyProvenance information consistent across updates, version_access, and distribution_dates fields
retention_limit
retention_limit:
  id: voice:retention:1
  name: Data retention and disposition policy
  description: 'Data Transfer and Use Agreement (DTUA) specifies two-year agreement term from start date.
    Upon termination or expiration, data shall be destroyed per provider instructions with written certification
    required within 30 days. Recipient may retain one copy to comply with records retention requirements
    under law, regulation, institutional policy, and for research integrity and verification. Ongoing
    restrictions apply to retained copies. Provider may unilaterally amend agreement if federal sponsor
    requires revision. Health Data Nexus retains dataset as long as useful for research purposes, possibly
    indefinitely. Version-specific DOIs maintained for all historical versions.

    '
5/5 R20 Q18 (FAIRness & Accessibility) Reusability (License Clarity)
levelLicense explicitly defines reuse terms
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License specifies commercial and non-commercial research use permitted for Authorized Researchers, no geographic restriction, no research type restriction, recipients encouraged to publish in open-access journals, Fort Lauderdale Agreement principles apply; ip_restrictions: no IP protections that would block access to dataset or conclusions; retention_limit: 2-year DTUA term with data destruction requirements upon expiration
qualityExceptionally detailed license with explicit reuse permissions (commercial/non-commercial research), restrictions (no sale, no third-party sharing), and IP protections favoring open science
correctnessLicense terms semantically coherent: registered access allows broad research use while prohibiting commercialization of the data itself
consistencyLicense restrictions consistent with ip_restrictions (Fort Lauderdale) and regulatory_restrictions (Certificate of Confidentiality)
version_access
version_access:
  id: voice:versionaccess:1
  name: Version access policy
  description: 'All dataset versions available through PhysioNet at https://physionet.org/content/b2ai-voice/
    with version-specific DOIs. DOI for latest version: https://doi.org/10.13026/37yb-1t42. Earlier versions
    also available on Health Data Nexus at https://healthdatanexus.ai/content/b2ai-voice/1.0/. By default,
    older versions continue to be supported, hosted, and made available. Dataset publishers reserve right
    to remove access to older versions. Each version has unique DOI.

    '
✓ 1/1 R10 10.Cross-Platform and Community Integration Citation and DOI for Cross-referencing
evidencedoi: 10.13026/37yb-1t42; version_access documents version-specific DOIs for all releases; external_resources reference Interspeech 2024 publication (doi.org/10.21437/Interspeech.2024-1926), Zenodo REDCap archive (doi.org/10.5281/zenodo.13834653)
qualityDOI provided for dataset (10.13026/37yb-1t42), version-specific DOIs documented, related publications have DOIs (Interspeech 2024, Zenodo archive); no formal citation string but publisher and version enable construction
semanticCitation infrastructure semantically complete - dataset DOI enables permanent reference, version-specific DOIs enable precise replication, related resource DOIs support comprehensive citation
✓ 1/1 R10 6.Data Provenance and Version Tracking Version Access Methods Documented
evidenceversion_access: all versions available at https://physionet.org/content/b2ai-voice/ with version-specific DOIs; earlier versions also at Health Data Nexus; DOI for latest version: 10.13026/37yb-1t42; older versions continue to be supported and hosted
qualityExplicit version access policy: PhysioNet hosts all versions with unique DOIs, Health Data Nexus provides alternative access, commitment to ongoing support of older versions
semanticVersion access semantically robust - unique DOIs enable permanent citation of specific versions, multiple platforms reduce single-point-of-failure risk
5/5 R20 Q13 (Technical Documentation) Version History Documentation
levelComprehensive versioning with errata, updates, and release notes
evidenceversion: 3.0.0; version_access: all versions available via PhysioNet with version-specific DOIs, older versions on Health Data Nexus; updates: semi-annual releases, v1.0 Jan 17 2025 (306 participants, 12,523 recordings), v1.1 added MFCC features, v2.0.0 Apr 16 2025, v2.0.1 Aug 18 2025, v3.0.0 2025 (833 adults, ~61,937 recordings), pediatric v1.0 separate, target 10,000 by 2027; distribution_dates: detailed release timeline with participant counts per version
qualityExemplary version tracking with specific release dates, participant/recording counts per version, version-specific DOIs, and clear update roadmap
correctnessVersion progression plausible (306→833 participants over 1 year with semi-annual releases); version numbering follows semantic versioning (1.0→1.1→2.0.0→2.0.1→3.0.0)
consistencyVersion information consistent across version field, updates, version_access, and distribution_dates
5/5 R20 Q19 (FAIRness & Accessibility) Data Integrity and Provenance
levelStructured version control with timestamps
evidenceupdates: detailed version progression with timestamps (v1.0 Jan 17 2025, v1.1 Jan 17 2025, v2.0.0 Apr 16 2025, v2.0.1 Aug 18 2025, v3.0.0 2025) and change descriptions (v1.1 added MFCC features), participant counts per version, semi-annual release schedule; version_access: all versions maintained with version-specific DOIs; distribution_dates: release timeline documented
qualityExemplary provenance tracking with version-specific DOIs, timestamped releases, participant count evolution, and feature addition logs
correctnessVersion dates chronologically consistent; version numbering follows semantic versioning conventions; participant growth plausible (306→833 over ~1 year)
consistencyProvenance information consistent across updates, version_access, and distribution_dates fields
extension_mechanism
extension_mechanism:
  id: voice:extension:1
  name: Dataset extension and contribution mechanisms
  description: 'Derivative datasets can be published on Health Data Nexus referencing original source
    under same access conditions. Open-source repositories (b2aiprep, SenseLab) have discussion forums,
    issue pages, and pull request mechanisms for contributing improvements to preprocessing code. REDCap
    data dictionary available at https://github.com/eipm/bridge2ai-redcap with MIT license. Bridge2AI-Voice
    documentation at https://github.com/eipm/bridge2ai-docs. Future augmentations to voice collection
    protocol coordinated through consortium.

    '
no field-level feedback matched
ethical_reviews
ethical_reviews:
- id: voice:ethicalreview:1
  description: 'IRB review and approval obtained from University of South Florida Single IRB with subsite
    IRB approvals through Single IRB process. Study reviewed and approved for human subjects research.
    Bioethics guidance integrated throughout study design and conduct by consortium bioethicists (Belisle-Pipon,
    Ravitsky). Ethics module develops new guidelines for consenting to voice data collection, voice data
    sharing, and utilization in context of voice AI technology. Project addresses ethical issues from
    voice data generation through clinical adoption and downstream health decisions.

    '
⚠ low R20 · consistency
issuehuman_subject_research.involves_human_subjects=true and ethical_reviews present - CONSISTENT
fieldshuman_subject_research, ethical_reviews
fixNo action needed - ethics documentation properly aligned
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards IRB or Ethics Review Documented
evidenceethical_reviews: IRB approval from University of South Florida Single IRB with subsite approvals; human_subject_research: involves_human_subjects=true, IRB approval documented, bioethics guidance integrated
qualityComprehensive ethics documentation: University of South Florida Single IRB, subsite approvals, bioethics consortium members (Belisle-Pipon, Ravitsky) identified, ethics module for voice AI consent guidelines
semanticIRB documentation semantically complete - Single IRB model appropriate for multi-site study, bioethics team integration demonstrates ongoing oversight
✓ 1/1 R10 9.Dataset Evaluation and Limitations Disclosure Ethical Review Details Including Conflicts
evidenceethical_reviews: University of South Florida Single IRB with subsite approvals, bioethics guidance by Belisle-Pipon and Ravitsky, ethics module develops voice AI consent guidelines, addresses ethical issues from data generation through clinical adoption; no conflicts of interest documented but Certificate of Confidentiality and Fort Lauderdale principles referenced
qualityEthical review process detailed: Single IRB mechanism, bioethics team integration, ethics module development, lifecycle ethics consideration (generation → adoption); Fort Lauderdale principles and Open Science commitment address potential conflicts
semanticEthical review semantically robust - Single IRB appropriate for multi-site study, bioethics team integration demonstrates ongoing oversight, ethics module shows proactive guidance development; Fort Lauderdale principles mitigate publication race conflicts
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive (all human subjects protections documented)
evidenceethical_reviews: USF Single IRB with subsite approvals; human_subject_research: involves_human_subjects=true with IRB details; is_deidentified: HIPAA Safe Harbor applied with direct/indirect identifier removal; informed_consent: prospective consent for all participants; at_risk_populations: pediatric protections (ages 2-18, SickKids exclusive); Certificate of Confidentiality coverage
qualityExemplary comprehensive ethics documentation covering all protection areas with specific regulatory frameworks (HIPAA, 45 CFR 46, OMB M-07-16, Certificate of Confidentiality)
correctnessHIPAA Safe Harbor deidentification method appropriate for health voice data; IRB approval plausible for multi-site NIH-funded project
consistencyEthics framework consistent: human subjects=true → IRB approval present → informed consent documented → deidentification applied → vulnerable populations protected
human_subject_research
human_subject_research:
  id: voice:hsr:1
  name: Bridge2AI-Voice Human Subjects Research
  description: 'Observational study (cross-sectional design) involving direct collection of voice recordings,
    questionnaire responses, and EHR linkage from human participants. Prospective informed consent obtained
    from all participants. IRB approved by University of South Florida Single IRB with subsite IRBs through
    Single IRB process. Not a drug or medical device study. No data monitoring committee appointed. Study
    ID: OT2OD032720. Involves human subjects: Yes. IRB approval: University of South Florida Institutional
    Review Board (Single IRB). Special populations: Pediatric participants (aged 2-18 years) at Hospital
    for Sick Children. Regulatory compliance: HIPAA Safe Harbor, 45 CFR 46 (Common Rule), Certificate
    of Confidentiality, OMB Memorandum M-07-16.

    '
  involves_human_subjects: true
⚠ low R20 · consistency
issuehuman_subject_research.involves_human_subjects=true and ethical_reviews present - CONSISTENT
fieldshuman_subject_research, ethical_reviews
fixNo action needed - ethics documentation properly aligned
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards IRB or Ethics Review Documented
evidenceethical_reviews: IRB approval from University of South Florida Single IRB with subsite approvals; human_subject_research: involves_human_subjects=true, IRB approval documented, bioethics guidance integrated
qualityComprehensive ethics documentation: University of South Florida Single IRB, subsite approvals, bioethics consortium members (Belisle-Pipon, Ravitsky) identified, ethics module for voice AI consent guidelines
semanticIRB documentation semantically complete - Single IRB model appropriate for multi-site study, bioethics team integration demonstrates ongoing oversight
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive (all human subjects protections documented)
evidenceethical_reviews: USF Single IRB with subsite approvals; human_subject_research: involves_human_subjects=true with IRB details; is_deidentified: HIPAA Safe Harbor applied with direct/indirect identifier removal; informed_consent: prospective consent for all participants; at_risk_populations: pediatric protections (ages 2-18, SickKids exclusive); Certificate of Confidentiality coverage
qualityExemplary comprehensive ethics documentation covering all protection areas with specific regulatory frameworks (HIPAA, 45 CFR 46, OMB M-07-16, Certificate of Confidentiality)
correctnessHIPAA Safe Harbor deidentification method appropriate for health voice data; IRB approval plausible for multi-site NIH-funded project
consistencyEthics framework consistent: human subjects=true → IRB approval present → informed consent documented → deidentification applied → vulnerable populations protected
at_risk_populations
at_risk_populations:
  id: voice:atrisk:1
  name: Pediatric participants protection
  description: 'Pediatric participants (aged 2-18) require additional protections. Pediatric data collected
    exclusively at Hospital for Sick Children (SickKids) with age-appropriate protocols. Distinct IRB
    considerations for pediatric cohort. Pediatric questionnaires and acoustic tasks specifically designed
    for different age groups (2-4, 4-6, 6-10, 10+ years). Minimum age for adult dataset is 18 years. Pediatric
    dataset released separately with additional privacy precautions.

    '
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Vulnerable Populations and Compensation Documented
evidenceat_risk_populations: pediatric participants (aged 2-18) with age-appropriate protocols, distinct IRB considerations, Hospital for Sick Children exclusive collection site, separate dataset release; data_collectors specify $40-120 compensation via electronic gift cards
qualityVulnerable population protections for pediatric cohort: age-appropriate tasks, distinct IRB review, specialized questionnaires, separate data release; compensation structure documented ($40 <90min, $80 >90min, max $120)
semanticPediatric protections semantically appropriate - separate protocols, single trusted institution (SickKids), age-stratified tasks (2-4, 4-6, 6-10, 10+); compensation reasonable and transparent
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive (all human subjects protections documented)
evidenceethical_reviews: USF Single IRB with subsite approvals; human_subject_research: involves_human_subjects=true with IRB details; is_deidentified: HIPAA Safe Harbor applied with direct/indirect identifier removal; informed_consent: prospective consent for all participants; at_risk_populations: pediatric protections (ages 2-18, SickKids exclusive); Certificate of Confidentiality coverage
qualityExemplary comprehensive ethics documentation covering all protection areas with specific regulatory frameworks (HIPAA, 45 CFR 46, OMB M-07-16, Certificate of Confidentiality)
correctnessHIPAA Safe Harbor deidentification method appropriate for health voice data; IRB approval plausible for multi-site NIH-funded project
consistencyEthics framework consistent: human subjects=true → IRB approval present → informed consent documented → deidentification applied → vulnerable populations protected
license_and_use_terms
license_and_use_terms:
  id: voice:license:1
  name: Bridge2AI Voice Registered Access License
  description: 'Public access dataset distributed through PhysioNet under Bridge2AI Voice Registered Access
    License. Only registered users who sign the specified Data Use Agreement (Bridge2AI Voice Registered
    Access Agreement) can access files. Data covered under Certificate of Confidentiality which must be
    asserted against compulsory legal demands. Raw audio available through controlled access only via
    Data Access Compliance Office (DACO) requiring distinct application and DTUA signed by institutional
    official. Recipient must adhere to PhysioNet requirements managed by MIT Laboratory for Computational
    Physiology. Recipient encouraged to publish results in open-access journals. No export controls apply.
    Commercial and non-commercial research use permitted for Authorized Researchers.

    '
✓ 1/1 R10 2.Dataset Access and Retrieval Access Policy and IP Restrictions Defined
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License with Data Use Agreement requirement; ip_restrictions: no intellectual property protections allowed that would block access, Fort Lauderdale principles apply
qualityComprehensive access policy specifying registered access model, Data Use Agreement requirements, and IP restrictions preventing proprietary barriers
semanticAccess policy semantically appropriate for sensitive health data - registered access balances openness with participant protection
✓ 1/1 R10 3.Data Reuse and Interoperability License Terms Allow Reuse
evidencelicense: Bridge2AI Voice Registered Access License; license_and_use_terms specifies commercial and non-commercial research use permitted, no geographic restrictions, no restriction to specific research type
qualityClear license allowing broad research reuse (commercial + non-commercial), Open Science principles, publication in open-access journals encouraged
semanticLicense terms semantically permissive for research while maintaining participant protections through registered access mechanism
5/5 R20 Q1 (Structural Completeness) Field Completeness
level≥90% fields populated
evidenceid: https://doi.org/10.13026/37yb-1t42, title: Bridge2AI-Voice, description: 1148 chars comprehensive, keywords: 73 keywords, license_and_use_terms: detailed registered access license
qualityAll mandatory fields present with exceptional completeness and detail
correctnessDOI format valid, PhysioNet prefix 10.13026 verified
consistencyAll mandatory fields semantically coherent and appropriate for voice biomarker dataset
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path (platform, login, policy)
evidencedistribution_formats: PhysioNet registered access for features (TSV/JSON/Parquet), controlled access for raw audio (WAV/GZ via DACO); license_and_use_terms: registered users sign Bridge2AI Voice Registered Access Agreement, raw audio requires DACO application + institutional DTUA; download_url: https://physionet.org/content/b2ai-voice/; external_resources: DACO contact mailto:DACO@b2ai-voice.org
qualityExceptionally clear two-tier access mechanism: (1) PhysioNet registered access with DUA for features, (2) DACO-controlled access with institutional DTUA for raw audio biometrics
correctnessAccess tiers semantically appropriate for biometric data: lower barrier for de-identified features, higher barrier with institutional oversight for re-identifiable raw audio
consistencyAccess mechanism consistent across distribution_formats, license_and_use_terms, and confidential_elements documentation
5/5 R20 Q18 (FAIRness & Accessibility) Reusability (License Clarity)
levelLicense explicitly defines reuse terms
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License specifies commercial and non-commercial research use permitted for Authorized Researchers, no geographic restriction, no research type restriction, recipients encouraged to publish in open-access journals, Fort Lauderdale Agreement principles apply; ip_restrictions: no IP protections that would block access to dataset or conclusions; retention_limit: 2-year DTUA term with data destruction requirements upon expiration
qualityExceptionally detailed license with explicit reuse permissions (commercial/non-commercial research), restrictions (no sale, no third-party sharing), and IP protections favoring open science
correctnessLicense terms semantically coherent: registered access allows broad research use while prohibiting commercialization of the data itself
consistencyLicense restrictions consistent with ip_restrictions (Fort Lauderdale) and regulatory_restrictions (Certificate of Confidentiality)
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License with DUA requirement; ip_restrictions: Fort Lauderdale Agreement principles, no IP blocking; regulatory_restrictions: Certificate of Confidentiality, HIPAA Safe Harbor, OMB M-07-16, 45 CFR 46; confidential_elements: raw audio biometric identifiers, controlled access via DACO
qualityExceptionally detailed governance framework with multi-tier access (registered for features, controlled for raw audio), IP protections, and regulatory compliance
correctnessAccess tiers semantically appropriate: registered access for de-identified features, controlled access (DACO + DTUA) for biometric raw audio
consistencyLicensing terms consistent with confidentiality requirements and regulatory restrictions across all documentation
ip_restrictions
ip_restrictions:
  id: voice:ip:1
  name: Intellectual property restrictions
  description: 'Recipient shall not develop or use any intellectual property protections, database rights,
    or related rights in any element of the dataset or in conclusions derived from it that would prevent
    or block access to any element of the dataset or those conclusions. Fort Lauderdale Agreement principles
    apply. Recipients must respect Open Science principles. Recipient shall not disclose, release, sell,
    rent, or lease data to third parties without prior written consent of Provider.

    '
✓ 1/1 R10 2.Dataset Access and Retrieval Access Policy and IP Restrictions Defined
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License with Data Use Agreement requirement; ip_restrictions: no intellectual property protections allowed that would block access, Fort Lauderdale principles apply
qualityComprehensive access policy specifying registered access model, Data Use Agreement requirements, and IP restrictions preventing proprietary barriers
semanticAccess policy semantically appropriate for sensitive health data - registered access balances openness with participant protection
5/5 R20 Q18 (FAIRness & Accessibility) Reusability (License Clarity)
levelLicense explicitly defines reuse terms
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License specifies commercial and non-commercial research use permitted for Authorized Researchers, no geographic restriction, no research type restriction, recipients encouraged to publish in open-access journals, Fort Lauderdale Agreement principles apply; ip_restrictions: no IP protections that would block access to dataset or conclusions; retention_limit: 2-year DTUA term with data destruction requirements upon expiration
qualityExceptionally detailed license with explicit reuse permissions (commercial/non-commercial research), restrictions (no sale, no third-party sharing), and IP protections favoring open science
correctnessLicense terms semantically coherent: registered access allows broad research use while prohibiting commercialization of the data itself
consistencyLicense restrictions consistent with ip_restrictions (Fort Lauderdale) and regulatory_restrictions (Certificate of Confidentiality)
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License with DUA requirement; ip_restrictions: Fort Lauderdale Agreement principles, no IP blocking; regulatory_restrictions: Certificate of Confidentiality, HIPAA Safe Harbor, OMB M-07-16, 45 CFR 46; confidential_elements: raw audio biometric identifiers, controlled access via DACO
qualityExceptionally detailed governance framework with multi-tier access (registered for features, controlled for raw audio), IP protections, and regulatory compliance
correctnessAccess tiers semantically appropriate: registered access for de-identified features, controlled access (DACO + DTUA) for biometric raw audio
consistencyLicensing terms consistent with confidentiality requirements and regulatory restrictions across all documentation
regulatory_restrictions
regulatory_restrictions:
  id: voice:regulatory:1
  name: Regulatory restrictions and compliance requirements
  description: 'Dataset covered under Certificate of Confidentiality per 45 CFR 46 (Common Rule), which
    must be asserted against compulsory legal demands such as court orders and subpoenas. HIPAA Safe Harbor
    de-identification standards applied. Data covered as Personally Identifiable Information per OMB Memorandum
    M-07-16. No export control restrictions apply. Authorized Researchers must comply with all applicable
    laws and regulations. DTUA specifies two-year term with data destruction requirements.

    '
✓ 1/1 R10 2.Dataset Access and Retrieval Regulatory Restrictions and Confidentiality Level Specified
evidenceregulatory_restrictions: Certificate of Confidentiality per 45 CFR 46, HIPAA Safe Harbor, OMB M-07-16 PII protections; confidential_elements: raw audio as biometric identifier, open-response features restricted
qualityDetailed regulatory framework including Certificate of Confidentiality, HIPAA compliance, and specific confidentiality level distinctions (registered vs controlled access)
semanticRegulatory restrictions semantically consistent with data sensitivity - Certificate of Confidentiality appropriate for identifiable health research data
5/5 R20 Q9 (Metadata Quality & Content) Access Requirements and Governance Documentation
levelLicense + restrictions + confidentiality classification
evidencelicense_and_use_terms: Bridge2AI Voice Registered Access License with DUA requirement; ip_restrictions: Fort Lauderdale Agreement principles, no IP blocking; regulatory_restrictions: Certificate of Confidentiality, HIPAA Safe Harbor, OMB M-07-16, 45 CFR 46; confidential_elements: raw audio biometric identifiers, controlled access via DACO
qualityExceptionally detailed governance framework with multi-tier access (registered for features, controlled for raw audio), IP protections, and regulatory compliance
correctnessAccess tiers semantically appropriate: registered access for de-identified features, controlled access (DACO + DTUA) for biometric raw audio
consistencyLicensing terms consistent with confidentiality requirements and regulatory restrictions across all documentation
is_deidentified
is_deidentified:
  id: voice:deident:1
  name: HIPAA Safe Harbor de-identification
  description: 'HIPAA Safe Harbor de-identification applied. All direct identifiers removed (names, civic
    addresses, social security numbers). Indirect identifiers removed where creating significant re-identification
    risk (geographic/demographic identifiers, household composition, cultural identity). Non-identifying
    sensitive information removed (household income, mental health status, traumatic life experiences).
    All raw voice data removed from public release. Sensitive fields per REDCap data dictionary removed.
    Direct identifiers removed: Yes. HIPAA de-identification rules applied: Yes. Dates rebased: Yes. Geographic
    information removed/generalized: Yes. Narrative text fields removed: Yes. K-anonymization: No.

    '
⚠ low R20 · consistency
issueis_deidentified present with HIPAA Safe Harbor method documented - CONSISTENT
fieldsis_deidentified
fixNo action needed - deidentification method appropriate for health data
✓ 1/1 R10 4.Ethical Use and Privacy Safeguards Deidentification Method Described
evidenceis_deidentified: HIPAA Safe Harbor de-identification applied, direct identifiers removed (names, addresses, SSN), indirect identifiers removed where creating re-identification risk, dates rebased, geographic information generalized
qualitySpecific deidentification method (HIPAA Safe Harbor) with detailed implementation: direct identifier removal, indirect identifier assessment, geographic generalization, narrative text removal
semanticHIPAA Safe Harbor method semantically appropriate for health research dataset; detailed implementation description enables reproducibility
5/5 R20 Q8 (Metadata Quality & Content) Ethical and Privacy Declarations
levelComprehensive (all human subjects protections documented)
evidenceethical_reviews: USF Single IRB with subsite approvals; human_subject_research: involves_human_subjects=true with IRB details; is_deidentified: HIPAA Safe Harbor applied with direct/indirect identifier removal; informed_consent: prospective consent for all participants; at_risk_populations: pediatric protections (ages 2-18, SickKids exclusive); Certificate of Confidentiality coverage
qualityExemplary comprehensive ethics documentation covering all protection areas with specific regulatory frameworks (HIPAA, 45 CFR 46, OMB M-07-16, Certificate of Confidentiality)
correctnessHIPAA Safe Harbor deidentification method appropriate for health voice data; IRB approval plausible for multi-site NIH-funded project
consistencyEthics framework consistent: human subjects=true → IRB approval present → informed consent documented → deidentification applied → vulnerable populations protected
external_resources
external_resources:
- id: voice:resource:1
  name: PhysioNet Dataset Landing Page
  description: Primary registered access distribution platform for adult and pediatric datasets
  external_resources:
  - https://physionet.org/content/b2ai-voice/
- id: voice:resource:2
  name: Bridge2AI-Voice Project Documentation
  description: Comprehensive project documentation, collection methods, governance, healthsheet
  external_resources:
  - https://docs.b2ai-voice.org
- id: voice:resource:3
  name: Bridge2AI-Voice GitHub Documentation Repository
  description: Source code for docs and dashboard at docs.b2ai-voice.org (MIT license)
  external_resources:
  - https://github.com/eipm/bridge2ai-docs
- id: voice:resource:4
  name: b2aiprep Software Library
  description: 'Open source library (Apache-2.0 license) for preprocessing raw audio waveforms into parquet
    files and merging source data into phenotype files

    '
  external_resources:
  - https://github.com/sensein/b2aiprep
- id: voice:resource:5
  name: Bridge2AI REDCap Data Dictionary
  description: REDCap data dictionary and metadata (MIT license)
  external_resources:
  - https://github.com/eipm/bridge2ai-redcap
- id: voice:resource:6
  name: NIH RePORTER Project Details
  description: Federal grant information for project 3OT2OD032720-01S3
  external_resources:
  - https://reporter.nih.gov/project-details/11376382
- id: voice:resource:7
  name: Health Data Nexus
  description: Alternative platform providing cloud compute alongside earlier dataset versions
  external_resources:
  - https://healthdatanexus.ai/content/b2ai-voice/1.0/
- id: voice:resource:8
  name: Zenodo Archive (REDCap data dictionary)
  description: 'Bensoussan Y. et al. (2024). eipm/bridge2ai-redcap. Zenodo archive.

    '
  external_resources:
  - https://doi.org/10.5281/zenodo.13834653
- id: voice:resource:9
  name: Interspeech 2024 Protocol Publication
  description: 'Bensoussan et al. "Developing Multi-Disorder Voice Protocols: A team science approach
    involving clinical expertise, bioethics, standards, and DEI." Proc. Interspeech 2024.

    '
  external_resources:
  - https://doi.org/10.21437/Interspeech.2024-1926
- id: voice:resource:10
  name: Data Access Compliance Office (DACO)
  description: Contact for controlled access to raw audio data and institutional DTUA
  external_resources:
  - mailto:DACO@b2ai-voice.org
- id: voice:resource:11
  name: Bridge2AI Program
  description: Parent NIH Common Fund program supporting AI-ready biomedical datasets
  external_resources:
  - https://bridge2ai.org
- id: voice:resource:12
  name: Bridge2AI Voice Scholars Training
  description: Training opportunities for using the dataset
  external_resources:
  - https://www.b2aivoicescholars.org/
✓ 1/1 R10 1.Dataset Discovery and Identification Landing Page and Resources (page, hierarchical resources)
evidencepage: https://docs.b2ai-voice.org; download_url: https://physionet.org/content/b2ai-voice/; external_resources: 12 documented resources including PhysioNet, GitHub repos, documentation
qualityMultiple accessible landing pages (project docs and PhysioNet), hierarchical resources structure with 12 external resources providing documentation, code, and access points
semanticURLs semantically appropriate - docs.b2ai-voice.org for documentation, physionet.org for data access, github.com for code repositories
✓ 1/1 R10 1.Dataset Discovery and Identification Hierarchical Structure (parent datasets, relationships)
evidenceMultiple distribution formats and access tiers: registered access featurized dataset (PhysioNet), controlled access raw audio (DACO), pediatric v1.0 as separate dataset; external_resources links to parent Bridge2AI program
qualityClear hierarchical structure with parent program (Bridge2AI), tiered access (registered vs controlled), separate pediatric dataset, and linked resources
semanticHierarchy semantically coherent - dataset part of Bridge2AI program, multiple access tiers based on re-identification risk, pediatric data as related sub-dataset
✓ 1/1 R10 10.Cross-Platform and Community Integration Citation and DOI for Cross-referencing
evidencedoi: 10.13026/37yb-1t42; version_access documents version-specific DOIs for all releases; external_resources reference Interspeech 2024 publication (doi.org/10.21437/Interspeech.2024-1926), Zenodo REDCap archive (doi.org/10.5281/zenodo.13834653)
qualityDOI provided for dataset (10.13026/37yb-1t42), version-specific DOIs documented, related publications have DOIs (Interspeech 2024, Zenodo archive); no formal citation string but publisher and version enable construction
semanticCitation infrastructure semantically complete - dataset DOI enables permanent reference, version-specific DOIs enable precise replication, related resource DOIs support comprehensive citation
✗ 0/1 R10 10.Cross-Platform and Community Integration Outreach Materials and Documentation Links
evidenceexternal_resources include documentation (docs.b2ai-voice.org, github.com/eipm/bridge2ai-docs), training program (b2aivoicescholars.org), but no webinars, tutorials, or user guides explicitly documented
qualityDocumentation site (docs.b2ai-voice.org) and training program (Bridge2AI Voice Scholars) exist, but no explicit outreach materials (webinars, video tutorials, quickstart guides) documented in metadata
semanticLimited outreach documentation - project docs and training program indicate community engagement, but absence of explicit tutorial/webinar links in metadata prevents full credit
✗ 0/1 R10 10.Cross-Platform and Community Integration Related Datasets with Typed Relationships
evidencePediatric v1.0 mentioned as separate dataset release but no explicit typed relationship documented; external_resources link to Bridge2AI program but no formal related_datasets field with relationship types (supplements, derives_from, is_version_of)
qualityPediatric dataset relationship implied (part of same collection) but not formalized with typed relationship; Bridge2AI program context provided but no explicit related_datasets linkage
semanticRelated dataset connections semantically present but not formally typed - pediatric v1.0 clearly related (same project, separate release) but relationship not structured as supplements/part_of; Bridge2AI program linkage informal
✓ 1/1 R10 2.Dataset Access and Retrieval Related Datasets and External Resources Linked
evidenceexternal_resources: 12 documented resources including related datasets (pediatric v1.0), code repositories (b2aiprep, Bridge2AI-redcap), documentation (docs.b2ai-voice.org), publications (Interspeech 2024)
qualityComprehensive external resource documentation with distinct types: software tools, documentation, publications, alternative platforms, parent program links
semanticExternal resources semantically organized by function - data access (PhysioNet, Health Data Nexus), processing (b2aiprep), documentation (docs, GitHub), validation (publications)
✗ 0/1 R10 7.Scientific Motivation and Funding Transparency Grant IDs or Award Numbers Present
evidencefunders reference grant numbers: 3OT2OD032720-01S3 (primary), R01EB030362 (PhysioNet); external_resources links NIH RePORTER project 11376382
qualityGrant numbers present but format non-standard - 3OT2OD032720-01S3 uses '3' prefix atypical for NIH (standard patterns: 1/2/5 for original/competitive renewal/non-competing continuation); R01EB030362 follows standard R01 mechanism format
semanticPrimary grant 3OT2OD032720-01S3 has non-standard prefix '3' - NIH typically uses 1 (new/renewal), 2 (competitive renewal), 5 (continuation); -01S3 suffix suggests supplement. Secondary grant R01EB030362 follows standard NIH R01 pattern. Scoring 0 due to non-standard format despite presence.
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) Software and Tools Documented
evidencepreprocessing_strategies and machine_annotation_tools reference b2aiprep, SenseLab, OpenSMILE, Parselmouth, Praat, torchaudio, OpenAI Whisper; external_resources link GitHub repos: b2aiprep (github.com/sensein/b2aiprep), Bridge2AI-redcap (github.com/eipm/bridge2ai-redcap), documentation (github.com/eipm/bridge2ai-docs)
qualitySoftware tools named with specific functions: b2aiprep (preprocessing), OpenSMILE (eGeMaps features), Parselmouth/Praat (phonetic features), torchaudio (audio processing), Whisper (transcription); GitHub repos linked for reproducibility
semanticSoftware documentation semantically complete - tools specified with versions/licenses where applicable (b2aiprep Apache-2.0, REDCap dictionary MIT), GitHub links enable code access
✓ 1/1 R10 8.Technical Transparency (Data Collection and Processing) External Standards and Resources Referenced
evidenceexternal_resources: 12 documented including BIDS v1.9.0 specification (referenced 6+ times), Interspeech 2024 protocol publication (doi.org/10.21437/Interspeech.2024-1926), NIH RePORTER grant details, Zenodo REDCap archive (doi.org/10.5281/zenodo.13834653), Bridge2AI program site
qualityExternal standards explicitly cited: BIDS v1.9.0 (data structure), ICD-10 (diagnoses), Bridge2AI protocols; publications provide methodological validation (Interspeech 2024); code/documentation archived (Zenodo, GitHub)
semanticExternal resources semantically comprehensive - standards (BIDS), publications (Interspeech), repositories (GitHub, Zenodo), program context (Bridge2AI), grant information (NIH RePORTER)
4/5 R20 Q14 (Technical Documentation) Associated Publications
levelMultiple references
evidenceexternal_resources: 12 entries including Interspeech 2024 protocol publication (DOI: 10.21437/Interspeech.2024-1926), Zenodo REDCap archive (DOI: 10.5281/zenodo.13834653), NIH RePORTER project details, GitHub repos (bridge2ai-docs, b2aiprep, bridge2ai-redcap with MIT/Apache-2.0 licenses)
qualityStrong publication documentation with 2 DOI-linked publications and comprehensive external resources; minor gap: no dedicated citation field for dataset itself in D4D-core schema
correctnessDOIs valid format: 10.21437 (ISCA Interspeech), 10.5281 (Zenodo); publication venues appropriate (Interspeech for voice research)
consistencyExternal resources align with project scope: protocol papers, data dictionaries, software tools, documentation sites
1/1 R20 Q16 (FAIRness & Accessibility) Findability (Persistent Links)
levelPass
evidencepage: https://docs.b2ai-voice.org, download_url: https://physionet.org/content/b2ai-voice/, external_resources: 12 entries with URLs to PhysioNet, GitHub (bridge2ai-docs, b2aiprep, bridge2ai-redcap), NIH RePORTER, Health Data Nexus, Zenodo, Interspeech DOI, Bridge2AI program, training site
qualityExcellent persistent link coverage with documentation, download, and external resource URLs across multiple platforms
correctnessAll URLs follow valid format with appropriate domains (physionet.org, github.com, doi.org, nih.gov, zenodo.org)
consistencyURLs consistent with described distribution platforms (PhysioNet primary, Health Data Nexus alternative)
5/5 R20 Q17 (FAIRness & Accessibility) Accessibility (Access Mechanism)
levelFully defined access path (platform, login, policy)
evidencedistribution_formats: PhysioNet registered access for features (TSV/JSON/Parquet), controlled access for raw audio (WAV/GZ via DACO); license_and_use_terms: registered users sign Bridge2AI Voice Registered Access Agreement, raw audio requires DACO application + institutional DTUA; download_url: https://physionet.org/content/b2ai-voice/; external_resources: DACO contact mailto:DACO@b2ai-voice.org
qualityExceptionally clear two-tier access mechanism: (1) PhysioNet registered access with DUA for features, (2) DACO-controlled access with institutional DTUA for raw audio biometrics
correctnessAccess tiers semantically appropriate for biometric data: lower barrier for de-identified features, higher barrier with institutional oversight for re-identifiable raw audio
consistencyAccess mechanism consistent across distribution_formats, license_and_use_terms, and confidential_elements documentation
1/1 R20 Q20 (FAIRness & Accessibility) Interlinking Across Platforms
levelPass
evidenceexternal_resources: cross-platform links verified - PhysioNet (primary distribution), Health Data Nexus (alternative with cloud compute), GitHub (3 repos: bridge2ai-docs, b2aiprep, bridge2ai-redcap), Zenodo (archive), NIH RePORTER (grant tracking), Bridge2AI program site, training portal; page: https://docs.b2ai-voice.org; download_url: https://physionet.org/content/b2ai-voice/
qualityExcellent cross-platform interlinking with 7+ distinct platforms (PhysioNet, Health Data Nexus, GitHub, Zenodo, NIH, Bridge2AI, training site) providing complementary access points
correctnessPlatform links semantically appropriate: PhysioNet for biomedical data distribution, GitHub for code/docs, Zenodo for archival, NIH RePORTER for grant transparency
consistencyCross-platform links consistent with described distribution strategy (PhysioNet primary, Health Data Nexus alternative, GitHub for open-source tools)

Unmatched feedback

Feedback that referenced multiple fields or no specific field.
⚠ low R20 · semantic_understanding
issueD4D-core schema lacks fields for software_and_tools, conforms_to_schema - these are expected in full schema but not exchange layer
fieldssoftware_and_tools, conforms_to_schema
fixAccept as expected limitation of core schema; full schema evaluation would score higher on technical documentation

Recommendations

  1. R10 · Verify grant number format 3OT2OD032720-01S3 - NIH typically uses 1/2/5 prefixes; confirm if '3' prefix is valid for supplements or administrative adjustments
  2. R10 · Add explicit typed relationships for pediatric dataset v1.0 using related_datasets field with relationship type (e.g., is_part_of or supplements)
  3. R10 · Document outreach materials explicitly: add links to any webinars, video tutorials, quickstart guides, or user documentation to external_resources
  4. R10 · Consider adding formal citation string in citation field to facilitate proper attribution (e.g., 'Bensoussan, Y. et al. (2025). Bridge2AI-Voice v3.0.0. PhysioNet. doi:10.13026/37yb-1t42')
  5. R10 · Clarify distinction between dataset collection funding (3OT2OD032720-01S3) and infrastructure funding (R01EB030362) to prevent confusion about grant purposes
  6. R10 · Add RRID identifier if Bridge2AI-Voice project or PhysioNet distribution has RRID:SCR_XXXXX registration to enhance tool/resource citability
  7. R10 · Document any analysis notebooks, code examples, or analysis tutorials as external_resources to support dataset reuse
  8. R10 · Consider adding related_datasets links to other Bridge2AI flagship datasets (if applicable) with typed relationships to enhance cross-project discovery
  9. R10 · Add explicit re-identification risk assessment results when formal analysis is completed (noted in limitation:5 as future work)
  10. R10 · Document Spanish protocol development progress in future versions to address known_biases:2 (English-only limitation)
  11. R20 · Enhance interoperability documentation by adding explicit conforms_to_schema field with BIDS v1.9.0 URI when using full D4D schema (currently only in descriptions)
  12. R20 · Add RRID identifiers for major software tools (e.g., OpenSMILE: RRID:SCR_014784, Praat: RRID:SCR_002591) to improve tool citation and reproducibility tracking
  13. R20 · Include software version numbers in structured fields when using full schema: b2aiprep version, OpenSMILE version, Whisper model version (large-v3 mentioned in descriptions)
  14. R20 · Consider adding dataset citation field in future D4D schema versions with BibTeX/RIS export for easier academic citation (complement to DOI)
  15. R20 · Document any errata or known issues in structured updates field (currently only version additions documented, no bug fixes or corrections mentioned)
  16. R20 · Add explicit bytes field for total dataset size alongside instance counts to help users plan storage/compute requirements
  17. R20 · Consider adding missing_data_documentation expansion: currently mentions missingness tables but could detail % completeness per questionnaire/task type
  18. R20 · Future releases could document sampling_strategies refinement: currently non-probability purposive, but could add demographic quotas for representativeness improvements