docs aireadi org docs-2 d4d

Datasheet for Dataset - Human Readable Format

🎯

Motivation

Why was the dataset created?

  • Response
    Enable AI/ML research to better understand salutogenesis (pathway from disease to health) in T2DM using harmonized, multi-domain data.
📊

Composition

What do the instances represent?

  1. Representation
    Human participants (individuals with and without T2DM) with multi-domain measurements.
    Instance Type
    Single cohort with multiple data modalities per participant (e.g., surveys, clinical exams, labs, imaging, wearables, environmental measures).
    Data Type
    Mixed data: tabular (surveys, labs, clinical measurements, environmental), time-series signals (e.g., ECG, activity), and images (e.g., retinal).
    Label
    No single pre-defined target; resource intended for diverse AI/ML tasks.
    Sampling Strategies
    1. Is Sample
      • True
      Is Random
      • False
      Source Data
      • Prospective recruitment across three data collection sites.
      Is Representative
      • Not explicitly; recruitment targeted diabetes severity balance.
      Why Not Representative
      • Enrollment is ongoing; pilot and periodic releases may not be balanced across groups.
      Strategies
      • Recruitment aimed at approximately equal distribution across diabetes severity.
  • Identification
    • People with T2DM
    • People without T2DM
    • Diabetes severity strata
    Distribution
    • Recruitment aimed at approximately equal distribution across diabetes severity; pilot/updates may not be balanced.
  • Description
    • Public dataset downloadable under an agreement-defined license; full dataset accessible via Data Use Agreement (DUA) through FAIRhub data portal.
  • Description
    • Pilot release completed; periodic updates planned (versions referenced include v1.0.0 and v2.0.0 in documentation).
🔍

Collection Process

How was the data acquired?

ai-readi-flagship-t2dm
AI-READI Flagship Dataset of Type 2 Diabetes
AI-READI Dataset
A harmonized, multi-modal dataset collected across three data collection sites from individuals with and without Type 2 Diabetes Mellitus (T2DM). The dataset was designed to enable future AI/Machine Learning studies, with recruitment and sampling aimed at approximately equal distribution across diabetes severity and a standardized data acquisition protocol spanning multiple domains (survey data, physical measurements, clinical data, imaging data, wearable device data, etc.). The goal is to better understand salutogenesis (the pathway from disease to health) in T2DM. A public dataset (non-sensitive) is downloadable under an agreement-defined license; the full dataset, which includes sensitive elements, is available via a data use agreement (DUA). Enrollment is ongoing; the pilot data release and subsequent updates may not achieve balanced distribution across all groups.
  • AI
  • Machine Learning
  • Type 2 Diabetes
  • T2DM
  • Multi-modal
  • Survey data
  • Clinical measurements
  • Imaging data
  • Retinal images
  • ECG
  • Wearable devices
  • Blood glucose
  • Laboratory results
  • Environmental data
  • Data harmonization
  • FAIR
  • AI-READI Project
  • Response
    Provides an AI/ML-ready, harmonized, multi-domain cohort resource that enables analyses not feasible with existing sources such as claims or EHR data alone.
RoleNameORCIDAffiliation
Contributor-AI-READI Project
  1. External Resources
    • Dataset landing and access via FAIRhub data portal (documentation references this portal).
    Future Guarantees
    • Not specified.
    Archival
    • Not specified.
    Restrictions
    • Full dataset requires entering into a Data Use Agreement (DUA).
  • Description
    • Dataset includes sensitive personal health data under controlled access (via DUA).
  • Description
    • 5-digit ZIP code
    • Sex
    • Race
    • Ethnicity
    • Genetic sequencing data
    • Past health records
    • Medications
    • Traffic and accident reports
  1. Description
    • Multi-domain protocol including direct measurements, surveys, imaging, and device data.
    • Data include directly observed, participant-reported, and derived elements.
    Was Directly Observed
    True
    Was Reported By Subjects
    True
    Was Inferred Derived
    True
    Was Validated Verified
    Not specified; harmonized protocol across three sites.
  • Description
    • Standardized clinical examinations and measurements.
    • Surveys/questionnaires.
    • Imaging (e.g., retinal photography).
    • Wearable/device data collection.
    • Laboratory assays (blood and urine).
  • Description
    • Data collected and harmonized across three data collection sites (site personnel not specified).
  • Description
    • Pilot study phase; enrollment is ongoing with periodic data release updates.
  • Description
    • Documentation highlights AI-readiness and ethical practices; specific IRB details not provided on this page.
  • Description
    • Documentation provides domain-specific processing details (file formats, standards, metadata, example outputs) per data domain.
  • Description
    • Domain sections describe data processing and harmonization; specific cleaning steps not detailed on this page.
  • Description
    • Not a labeled benchmark; domain documentation may include derived variables/annotations where applicable.
  • Description
    • AI-READI Project (contact via project documentation "Contact Us"; GitHub repository referenced for docs).
Description
  • Public release excludes sensitive personal health information; controlled-access data include quasi-identifiers and sensitive clinical/genomic elements.
Mixed; multi-modal (tabular, images, time-series signals)
🚀

Uses

What (other) tasks could the dataset be used for?

  • Response
    Downstream AI/ML analyses across multiple domains (e.g., modeling from survey, clinical, lab, imaging, wearable, and environmental signals); not limited to tasks feasible with claims/EHR alone.
  • Description
    • Multimodal risk prediction, phenotype discovery, causal inference, and digital biomarker research in T2DM and metabolic health.
  • Description
    • Ongoing enrollment and potential imbalance in early releases may affect model fairness/generalization across demographic and severity groups.
Description
  • Some non-sensitive data are publicly available for download upon agreeing to a license that defines permitted use.
  • Full dataset requires execution of a Data Use Agreement (DUA).
📤

Distribution

How will the dataset be distributed?

Description
  • Documentation provides a version selector to navigate datasets/documentation versions (e.g., v1.0.0 and v2.0.0).
🔄

Maintenance

How will the dataset be maintained?

2.0.0
2025-09-08
Description
  • Pilot data release with periodic updates to subsequent releases; documentation supports version navigation (e.g., v1.0.0, v2.0.0).
Generated on 2025-11-09 10:17:35 using Bridge2AI Data Sheets Schema