AI READI all combined d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • Name
    primary-purpose
    Response
    Create a harmonized, multi-modal flagship dataset to enable AI/ML research on Type 2 Diabetes Mellitus (T2DM), with a focus on salutogenesis (the pathway from disease to health).
📊

Composition

What do the instances represent?

  1. Name
    participant-level-multimodal-records
    Representation
    Participant-centric records with linked multi-domain measurements and observations
    Instance Type
    Participants and their associated multi-modal data (surveys, clinical assessments, labs, imaging, wearables, environmental)
    Data Type
    Mixed data modalities including tabular variables (surveys, labs, clinical), imaging (retinal), physiological signals (ECG), continuous glucose monitor streams, and derived features; harmonized across sites.
    Label
    No universal target; labels/variables vary by domain and analysis
    Sampling Strategies
    1. Name
      recruitment-and-balance
      Is Sample
      • True
      Is Random
      • False
      Source Data
      • Prospective recruitment across three clinical sites
      Is Representative
      • Partially; enrollment ongoing
      Representative Verification
      • Recruitment targeted approximate balance by diabetes severity; ongoing monitoring via study dashboard
      Why Not Representative
      • Pilot and early releases may not fully achieve balance across groups as enrollment continues
      Strategies
      • Targeted stratified recruitment by diabetes severity across sites
  • Name
    diabetes-status-and-severity
    Identification
    • Participants with and without T2DM; severity strata targeted during recruitment.
    Distribution
    • Approximate balance targeted; may vary in pilot/early releases due to ongoing enrollment.
  • Name
    distribution
    Description
    • Distributed via the FAIRhub data portal; gated controlled-access workflow; mixed modalities and file types across imaging, signals, and tabular data.
DescriptionName
{'2024-05-03 (DOI': '10.60775/fairhub.1)'}version-1.0.0
{'2024-11-08 (DOI': '10.60775/fairhub.2)'}version-2.0.0
🔍

Collection Process

How was the data acquired?

ai-readi-flagship-dataset-t2dm
Flagship Dataset of Type 2 Diabetes from the AI-READI Project
The AI-READI Flagship Dataset is a harmonized, multi-modal human health dataset collected across three data collection sites from individuals with and without Type 2 Diabetes Mellitus (T2DM). The dataset was designed with future AI/ML applications in mind, including a recruitment and sampling approach intended to approximate equal distribution across diabetes severity levels and a standardized data acquisition protocol spanning multiple domains: survey data, physical measurements, clinical assessments, imaging, wearable device data, and more. The overarching scientific goal is to support research into salutogenesis (the pathway from disease to health) in T2DM. The dataset is distributed via the FAIRhub data portal and has both a public component (de-identified and non-sensitive health data, available under a license) and a controlled-access component (additional sensitive elements, accessible via a data use agreement and an approval workflow). As enrollment is ongoing, early/pilot releases and periodic updates may not fully achieve the target balanced distribution across groups. Documentation for each dataset version is maintained and includes clinical context, acquisition methods, variable summaries, domain-specific processing and formats, and additional resources for AI-readiness, FAIR principles, and ethical considerations.
English
2024-11-08
2024-05-03
  • AI-READI Consortium
  • Diabetes mellitus
  • Machine Learning
  • Artificial Intelligence
  • Electrocardiography
  • Continuous Glucose Monitoring
  • Retinal imaging
  • Eye exam
mixed (tabular, imaging, waveform, sensor-derived)
  • Name
    AI-READI Consortium
  • Name
    addressing-data-gaps
    Response
    Provide integrated, multi-domain data not feasible to obtain from single sources such as claims or EHR alone; harmonize data across multiple collection sites to support robust AI/ML development and evaluation.
DescriptionIs Data SplitIs SubpopulationName
De-identified, non-sensitive health data available for download upon agreement to the dataset licens...FalseFalsePublic subset
Additional sensitive elements available under a data use agreement and approval process. Includes 5-...FalseFalseControlled-access subset
  • Name
    distribution-notes
    Description
    • Early releases and periodic updates may not fully achieve balanced distribution across groups as enrollment is ongoing.
  1. Name
    linked-external-context
    External Resources
    • Environmental variables (e.g., home air quality) and traffic/accident reports may be derived from or linked to external sources.
    Future Guarantees
    • Not specified; see dataset documentation and FAIRhub records for archival details.
    Archival
    • Versioned releases with documentation are maintained; prior versions retain DOIs.
    Restrictions
    • External datasets, if used, may carry their own licenses/terms; users must review terms where applicable.
  • Name
    controlled-elements
    Description
    • Controlled-access data include quasi-identifiers and sensitive health/genetic data (e.g., 5-digit ZIP, sex, race, ethnicity, genetic sequencing, past health records, medications, traffic/accident reports).
  • Name
    potential-sensitive-health-content
    Warnings
    • Contains medical and genetic information; exposure to sensitive health topics.
  • Name
    sensitive-health-and-demographic-elements
    Description
    • Health data (labs, clinical assessments), genetic data, demographic attributes (sex, race, ethnicity), location-derived data (5-digit ZIP).
Name
deidentification-summary
Description
  • Public subset is de-identified and excludes sensitive personal health data.
  • Controlled-access subset includes additional sensitive/quasi-identifying elements and is governed by a DUA and access controls.
  1. Name
    mixed-acquisition
    Description
    • Multi-domain acquisition across surveys, clinical measurements, labs, imaging devices, ECG, CGM, wearables, and linked environmental/contextual sources.
    Was Directly Observed
    yes (e.g., imaging, ECG, CGM, physical measurements)
    Was Reported By Subjects
    yes (e.g., survey responses)
    Was Inferred Derived
    yes (e.g., derived features, harmonized variables)
    Was Validated Verified
    Not specified in provided sources
  • Name
    instruments-and-procedures
    Description
    • Standardized data acquisition protocols across three sites, including clinical instruments, imaging hardware, physiological sensors, wearables, survey tools, and EHR extraction; harmonization applied across sites.
  • Name
    collection-sites-and-study-staff
    Description
    • Data collected at three clinical sites by study staff under a common protocol.
  • Name
    enrollment-and-releases
    Description
    • Pilot phase data comprise v1.0.0 (2024-05-03). Enrollment continued with v2.0.0 (2024-11-08). Ongoing releases are planned as data collection proceeds.
  • Name
    ethical-sourcing-statement
    Description
    • Described as an ethically-sourced dataset; consult site-specific documentation for IRB and ethics review details.
  • Name
    data-protection-and-access-controls
    Description
    • Public data under license; controlled-access data require a data use agreement and justification of research purpose; multi-step access workflow includes login, training, purpose statement, license acceptance, and data selection.
  • Name
    harmonization
    Description
    • Harmonized across three collection sites; domain-specific processing described in versioned documentation.
  • Name
    quality-and-standardization
    Description
    • Standardization and quality checks applied per domain; refer to domain-specific docs for details.
  • Name
    primary-data-sources
    Description
    • Surveys, clinical measurements and tests, laboratory results (blood, urine), imaging (retinal), ECG, CGM and wearable device data, EHR-derived variables, environmental/traffic context.
Name
third-party-ip
Description
  • Users should review terms for any linked external datasets (e.g., environmental, traffic) that may impose additional restrictions.
  • Name
    dataset-maintenance
    Description
    • AI-READI Consortium and FAIRhub platform maintain the dataset and documentation; "Changelog" and versioned docs are provided.
  • Name
    changelog
    Description
    • A versioned changelog is maintained in the documentation site.
Name
community-and-contact
Description
  • Contribution and inquiry via the project documentation site and "Contact Us"; project GitHub links are provided in the docs.
🚀

Uses

What (other) tasks could the dataset be used for?

  • Name
    intended-ai-ml-uses
    Response
    Broad AI/ML analyses across multiple domains (survey, clinical, lab, imaging, wearable/device, environmental) to support prediction, discovery, and integrative modeling in T2DM.
  • Name
    usage-and-stats
    Description
    • FAIRhub provides usage statistics and a "Dataset Uses" section; see portal for details.
  • Name
    potential-uses
    Description
    • Multi-modal integrative modeling, phenotype characterization, prognostic/risk modeling, and cross-domain exploratory analyses in T2DM and related conditions (subject to license/DUA).
  • Name
    fairness-and-generalizability
    Description
    • Potential impacts from evolving cohort composition (e.g., interim imbalance across groups) should be considered to mitigate fairness and generalizability risks in downstream models.
Name
licensing-and-dua
Description
  • Public subset available under a Health Data License; users must accept the license.
  • Controlled-access subset requires a Data Use Agreement and approval.
  • Access workflow includes login, diabetes research eligibility, required training, research purpose statement, license acceptance, and data selection.
📤

Distribution

How will the dataset be distributed?

Health Data License
Name
versioning-and-persistence
Description
  • Older Versions Remain Accessible And Citable Via Their Dois (e.g., V1.0.0
    10.60775/fairhub.1).
🔄

Maintenance

How will the dataset be maintained?

2.0.0
2024-11-08
Name
update-plan
Description
  • Periodic updates to add new participants/modalities and to correct issues; communicated via FAIRhub versioning and documentation updates.
Generated on 2025-11-09 10:17:34 using Bridge2AI Data Sheets Schema