Create a harmonized, multi-modal flagship dataset to enable AI/ML research
on Type 2 Diabetes Mellitus (T2DM), with a focus on salutogenesis (the
pathway from disease to health).
📊
Composition
What do the instances represent?
Name
participant-level-multimodal-records
Representation
Participant-centric records with linked multi-domain measurements and observations
Instance Type
Participants and their associated multi-modal data (surveys, clinical assessments, labs, imaging, wearables, environmental)
Data Type
Mixed data modalities including tabular variables (surveys, labs, clinical),
imaging (retinal), physiological signals (ECG), continuous glucose monitor
streams, and derived features; harmonized across sites.
Label
No universal target; labels/variables vary by domain and analysis
Sampling Strategies
Name
recruitment-and-balance
Is Sample
True
Is Random
False
Source Data
Prospective recruitment across three clinical sites
Is Representative
Partially; enrollment ongoing
Representative Verification
Recruitment targeted approximate balance by diabetes severity; ongoing monitoring via study dashboard
Why Not Representative
Pilot and early releases may not fully achieve balance across groups as enrollment continues
Strategies
Targeted stratified recruitment by diabetes severity across sites
Name
diabetes-status-and-severity
Identification
Participants with and without T2DM; severity strata targeted during recruitment.
Distribution
Approximate balance targeted; may vary in pilot/early releases due to ongoing enrollment.
Name
distribution
Description
Distributed via the FAIRhub data portal; gated controlled-access workflow; mixed modalities and file types across imaging, signals, and tabular data.
Flagship Dataset of Type 2 Diabetes from the AI-READI Project
The AI-READI Flagship Dataset is a harmonized, multi-modal human health dataset
collected across three data collection sites from individuals with and without
Type 2 Diabetes Mellitus (T2DM). The dataset was designed with future AI/ML
applications in mind, including a recruitment and sampling approach intended
to approximate equal distribution across diabetes severity levels and a
standardized data acquisition protocol spanning multiple domains: survey data,
physical measurements, clinical assessments, imaging, wearable device data,
and more. The overarching scientific goal is to support research into
salutogenesis (the pathway from disease to health) in T2DM.
The dataset is distributed via the FAIRhub data portal and has both a public
component (de-identified and non-sensitive health data, available under a
license) and a controlled-access component (additional sensitive elements,
accessible via a data use agreement and an approval workflow). As enrollment is
ongoing, early/pilot releases and periodic updates may not fully achieve the
target balanced distribution across groups. Documentation for each dataset
version is maintained and includes clinical context, acquisition methods,
variable summaries, domain-specific processing and formats, and additional
resources for AI-readiness, FAIR principles, and ethical considerations.
Provide integrated, multi-domain data not feasible to obtain from single
sources such as claims or EHR alone; harmonize data across multiple
collection sites to support robust AI/ML development and evaluation.
Description
Is Data Split
Is Subpopulation
Name
De-identified, non-sensitive health data available for download upon agreement
to the dataset licens...
False
False
Public subset
Additional sensitive elements available under a data use agreement and
approval process. Includes 5-...
False
False
Controlled-access subset
Name
distribution-notes
Description
Early releases and periodic updates may not fully achieve balanced distribution across groups as enrollment is ongoing.
Name
linked-external-context
External Resources
Environmental variables (e.g., home air quality) and traffic/accident reports may be derived from or linked to external sources.
Future Guarantees
Not specified; see dataset documentation and FAIRhub records for archival details.
Archival
Versioned releases with documentation are maintained; prior versions retain DOIs.
Restrictions
External datasets, if used, may carry their own licenses/terms; users must review terms where applicable.
Name
controlled-elements
Description
Controlled-access data include quasi-identifiers and sensitive health/genetic data (e.g., 5-digit ZIP, sex, race, ethnicity, genetic sequencing, past health records, medications, traffic/accident reports).
Name
potential-sensitive-health-content
Warnings
Contains medical and genetic information; exposure to sensitive health topics.
Name
sensitive-health-and-demographic-elements
Description
Health data (labs, clinical assessments), genetic data, demographic attributes (sex, race, ethnicity), location-derived data (5-digit ZIP).
Name
deidentification-summary
Description
Public subset is de-identified and excludes sensitive personal health data.
Controlled-access subset includes additional sensitive/quasi-identifying elements and is governed by a DUA and access controls.
Name
mixed-acquisition
Description
Multi-domain acquisition across surveys, clinical measurements, labs, imaging devices, ECG, CGM, wearables, and linked environmental/contextual sources.
Standardized data acquisition protocols across three sites, including clinical instruments, imaging hardware, physiological sensors, wearables, survey tools, and EHR extraction; harmonization applied across sites.
Name
collection-sites-and-study-staff
Description
Data collected at three clinical sites by study staff under a common protocol.
Name
enrollment-and-releases
Description
Pilot phase data comprise v1.0.0 (2024-05-03). Enrollment continued with v2.0.0 (2024-11-08). Ongoing releases are planned as data collection proceeds.
Name
ethical-sourcing-statement
Description
Described as an ethically-sourced dataset; consult site-specific documentation for IRB and ethics review details.
Name
data-protection-and-access-controls
Description
Public data under license; controlled-access data require a data use agreement and justification of research purpose; multi-step access workflow includes login, training, purpose statement, license acceptance, and data selection.
Name
harmonization
Description
Harmonized across three collection sites; domain-specific processing described in versioned documentation.
Name
quality-and-standardization
Description
Standardization and quality checks applied per domain; refer to domain-specific docs for details.
Users should review terms for any linked external datasets (e.g., environmental, traffic) that may impose additional restrictions.
Name
dataset-maintenance
Description
AI-READI Consortium and FAIRhub platform maintain the dataset and documentation; "Changelog" and versioned docs are provided.
Name
changelog
Description
A versioned changelog is maintained in the documentation site.
Name
community-and-contact
Description
Contribution and inquiry via the project documentation site and "Contact Us"; project GitHub links are provided in the docs.
🚀
Uses
What (other) tasks could the dataset be used for?
Name
intended-ai-ml-uses
Response
Broad AI/ML analyses across multiple domains (survey, clinical, lab,
imaging, wearable/device, environmental) to support prediction, discovery,
and integrative modeling in T2DM.
Name
usage-and-stats
Description
FAIRhub provides usage statistics and a "Dataset Uses" section; see portal for details.
Multi-modal integrative modeling, phenotype characterization, prognostic/risk modeling, and cross-domain exploratory analyses in T2DM and related conditions (subject to license/DUA).
Name
fairness-and-generalizability
Description
Potential impacts from evolving cohort composition (e.g., interim imbalance across groups) should be considered to mitigate fairness and generalizability risks in downstream models.
Name
licensing-and-dua
Description
Public subset available under a Health Data License; users must accept the license.
Controlled-access subset requires a Data Use Agreement and approval.
Access workflow includes login, diabetes research eligibility, required training, research purpose statement, license acceptance, and data selection.