================================================================================
CONCATENATED DOCUMENT
================================================================================
Input Directory: data/preprocessed/individual/AI_READI
Total Files: 11
Extensions: ['.txt']
Recursive: False
Selection Manifest: data/preprocessed/source_manifest.yaml
================================================================================

TABLE OF CONTENTS
--------------------------------------------------------------------------------
  1. bmjopen-2024-097449_row2.txt
  2. s42255-024-01165-x_row3.txt
  3. reporter_nih_gov_project-details-10471118_row7.txt
  4. docs_aireadi_org_docs-2_row10.txt
  5. docs_aireadi_org_docs-3_2026-07-24.txt
  6. AI-READI-LICENSE-v2.0_2026-08-12.txt
  7. fairhub_dataset_2_row12.txt
  8. fairhub_dataset_3_2026-07-24.txt
  9. fairhub_api_dataset_3_2026-07-27.txt
 10. aireadi_ro_crate_metadata_2026-08-12.txt
 11. gdrive_1rJsa5kySlBRRNhsO_WY7N3bfSKtqDi-Q_row13.txt
================================================================================

FILE: bmjopen-2024-097449_row2.txt
PATH: data/preprocessed/individual/AI_READI/bmjopen-2024-097449_row2.txt
SIZE: 50174 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: AI_READI
Source ID: bmj_protocol_publication
Source type: publication
Source URL: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11800295/pdf/bmjopen-2024-097449.pdf
Raw file: data/raw/AI_READI/bmjopen-2024-097449_row2.pdf
--------------------------------------------------------------------------------
Open access

Protocol

Cross- sectional design and protocol for
Artificial Intelligence Ready and
Equitable Atlas for Diabetes
Insights (AI- READI)

Cynthia Owsley      ,1 Dawn S Matthies,1 Gerald McGwin,1,2 Jeffrey C Edberg,3
Sally L Baxter      ,4 Linda M Zangwill,4 Julia P Owen,5 Cecilia S Lee,5 AI- READI
Consortium

To cite: Owsley C, Matthies DS,
McGwin G, et al.  Cross-
sectional design and protocol
for Artificial Intelligence Ready
and Equitable Atlas for Diabetes
Insights (AI- READI). BMJ Open
2025;15:e097449. doi:10.1136/
bmjopen-2024-097449

 ► Prepublication history for
this paper is available online.
To view these files, please visit
the journal online (https://doi.
org/10.1136/bmjopen-2024-
097449).

Received 02 December 2024
Accepted 22 January 2025

© Author(s) (or their
employer(s)) 2025. Re- use
permitted under CC BY- NC. No
commercial re- use. See rights
and permissions. Published by
BMJ Group.
1Ophthalmology and Visual
Sciences, The University of
Alabama at Birmingham,
Birmingham, Alabama, USA
2Epidemiology, The University of
Alabama at Birmingham School
of Public Health, Birmingham,
Alabama, USA
3Medicine, University of Alabama
at Birmingham, Birmingham,
Alabama, USA
4Ophthalmology, University of
California San Diego, La Jolla,
California, USA
5Ophthalmology, University
of Washington, Seattle,
Washington, USA

Correspondence to
Dr Cynthia Owsley;
 cynthiaowsley@ uabmc. edu

ABSTRACT
Introduction  Artificial Intelligence Ready and Equitable
for Diabetes Insights (AI- READI) is a data collection
project on type 2 diabetes mellitus (T2DM) to facilitate
the widespread use of artificial intelligence and machine
learning (AI/ML) approaches to study salutogenesis
(transitioning from T2DM to health resilience). The
fundamental rationale for promoting health resilience in
T2DM stems from its high prevalence of 10.5% of the
world’s adult population and its contribution to many
adverse health events.
Methods  AI- READI is a cross- sectional study whose
target enrollment is 4000 people aged 40 and older, triple-
balanced by self- reported race/ethnicity (Asian, black,
Hispanic, white), T2DM (no diabetes, pre- diabetes and
lifestyle- controlled diabetes, diabetes treated with oral
medications or non- insulin injections and insulin- controlled
diabetes) and biological sex (male, female) ( Clinicaltrials. org
approval number STUDY00016228). Data are collected in a
multivariable protocol containing over 10 domains, including
vitals, retinal imaging, electrocardiogram, cognitive function,
continuous glucose monitoring, physical activity, home air
quality, blood and urine collection for laboratory testing and
psychosocial variables including social determinants of
health. There are three study sites: Birmingham, Alabama;
San Diego, California; and Seattle, Washington.
Ethics and dissemination  AI- READI aims to establish
standards, best practices and guidelines for collection,
preparation and sharing of the data for the purposes of
AI/ML, including guidance from bioethicists. Following
Findable, Accessible, Interoperable, Reusable principles,
AI- READI can be viewed as a model for future efforts
to develop other medical/health data sets targeted
for AI/ML. AI- READI opens the door for novel insights
in understanding T2DM salutogenesis. The AI- READI
Consortium are disseminating the principles and processes
of designing and implementing the AI- READI data set
through publications. Those who download and use AI-
READI data are encouraged to publish their results in the
scientific literature.

INTRODUCTION
Artificial  Intelligence  Ready  and  Equitable
for  Diabetes  Insights  (AI- READI)1  is  one  of

STRENGTHS AND LIMITATIONS OF THIS STUDY
 ⇒ The  targeted  sample  size  of  4000  persons  is  the
largest publicly accessible data set currently avail-
able  containing  many  multidomain  variables  rele-
vant to type 2 diabetes mellitus (T2DM).

 ⇒ The  sample  is  designed  to  be  approximately  bal-
anced  with  respect  to  sex  and  race/ethnicities
(Asians,  blacks,  Hispanics  and  whites),  a  demo-
graphic improvement over many previous epidemi-
ological studies and clinical trials on T2DM.

 ⇒ The study population is designed to be balanced in
T2DM severity groups of no diabetes, pre- diabetes,
non- insulin controlled and insulin controlled, to fa-
cilitate artificial intelligence/machine learning model
developments.

 ⇒ The study design process enacted ethical and equi-
table data collection and management practices as
well as data sharing with adherence to the Findable,
Accessible, Interoperable, Reusable principles.
 ⇒ A limitation of our biorepository is that there are a
finite  number  of  samples  to  share  with  scientists
interested  in  using  them  in  research.  Procedures
for  reviewing  and  prioritising  written  requests  will
be developed before the biorepository is complete.

four  National  Institutes  of  Health- funded
Bridge2AI
(https://bridge2ai.org)  proj-
ects  that  aim  to  generate  flagship  biomed-
ical  and  behavioural  data  sets  that  are
ethically  sourced,  trustworthy,  well- defined
and  publicly  available.  AI- READI  is  a  data
generation project focused on type 2 diabetes
mellitus (T2DM) to facilitate the widespread
use  of  artificial  intelligence  and  machine
learning (AI/ML) approaches to study saluto-
genesis.2  Salutogenesis  in  this  context  refers
to the pathway from T2DM to health, which is
the opposite of studying how healthy people
proceed  through  pathogenesis  resulting  in
T2DM. The data set generated by AI- READI,
to  be  described  below,  is  multimodal,  given

1

Owsley C, et al. BMJ Open 2025;15:e097449. doi:10.1136/bmjopen-2024-097449
Open access

the multifactorial and multisystemic nature of T2DM. The
fundamental  rationale  for  promoting  health  resilience
in T2DM stems from its high prevalence of 10.5% of the
world’s adult population.3 Persons with T2DM are at risk
for  many  types  of  adverse  health  consequences  such  as
stroke, kidney disease, heart disease, vision impairment,
cognitive  decline,  peripheral  neuropathy  and  physical
inactivity.4 Social determinants of health weigh heavily in
threatening medical status in T2DM,5 including reduced
access to healthcare, decline in treatment adherence and
challenges  in  seeking  follow- up  preventative  care.6  The
prevalence of T2DM is higher in certain racial and ethnic
populations,7  8  and  the  deleterious  consequences  are
exacerbated in these populations.9 10

There  are  certain  unique  features  of  AI- READI’s
design  which  make  it  ideally  suited  for  the  use  of  AI/
ML  analytic  approaches,  which  opens  the  door  for  crit-
ical novel insights into understanding salutogenesis. First,
the AI- READI team, a multidisciplinary group of investi-
gators,  including  clinicians,  data  scientists,  vision  scien-
tists,  computer  scientists,  organisational  scientists  and
ethicists,  designed  the  protocol  such  that  it  is  agnostic
to hypotheses. Rather, the team assembled a data collec-
tion protocol composed of assessments from many health
domains  that  are  likely  impacted  by  T2DM.  Second,  to
address inherent limitations of many existing large data
sets  for  AI/ML  training,  the  cohort  is  designed  to  be
balanced in race/ethnicity, T2DM severity and biological
sex. Third, the data set is large, with a target sample size
of 4000 persons from three geographical regions in the
USA. Fourth, AI- READI places special emphasis on estab-
lishing standards, best practices and guidelines for collec-
tion, preparation and sharing of the data that can be used
for future efforts to develop other data sets targeted for
AI/ML.  Fifth,  AI- READI  addresses  challenges  that  have
compromised  AI  implementation  in  clinical  research
by  including  ethical  and  equitable  data  collection  and
management,  and  adherence  to  Findable,  Accessible,
Interoperable,  Reusable  (FAIR)  principles.  These  issues
have been discussed previously1 but are incorporated in
administering the design and protocol described below.

METHODS
AI- READI was approved by the Institutional Review Board
(IRB) of the University of Washington (approval number
STUDY00016228),  including  reliance  agreements  with
the  IRBs  of  University  of  Alabama  at  Birmingham  and
University  of  California,  San  Diego.  Written  informed
consent is provided by all participants.

Patient and public involvement
A  Community  Advisory  Board  of  11  persons  from  the
three  sites  including  diversity  in  race  and  ethnicity  as
represented  in  the  AI- READI  sample  contributes  to  the
development of the protocol.

Design
AI- READI  is  a  cross- sectional  study  whose  target  enrol-
ment is 4000 people aged 40 and older, triple- balanced by

2

diabetes,

self- reported race/ethnicity, T2DM presence and severity,
and  biological  sex  (figure  1).  Building  balanced  data
sets is critical for the development of unbiased machine
learning  models,  so  rather  than  targeting  the  demo-
graphic  distribution  of  the  US  population,  the  study  is
triple- balanced by recruiting the following populations in
equal proportions: four race/ethnic groups (Asian, black,
white, Hispanic), four categories of T2DM (no diabetes,
pre- diabetes/lifestyle- controlled
diabetes
treated  with  oral  medications  or  non- insulin- injectable
medications, and insulin- controlled diabetes) and biolog-
ical males and females. Participants are recruited across
three data collection sites in different geographical loca-
tions:  Birmingham,  Alabama  (University  of  Alabama  at
Birmingham  (UAB)),  San  Diego,  California  (University
of  California,  San  Diego  (UCSD)),  and  Seattle,  Wash-
ington  (University  of  Washington  (UW)).  All  study
groups are recruited from each site to capture diversity,
but the proportion of each group will vary depending on
the  demographic  prevalence  of  each  group  in  the  site’s
geographical  area.  Participants  are  required  to  speak,
read  and  understand  English.  Pregnancy  and  type  1
diabetes are exclusionary for participation.

Source population
The study base is all patients aged ≥40 years of age who
had a medical encounter within each health system site
(UAB, UCSD, UW) between 2020 and 2025. Patients with
T2DM and pre- diabetes are identified by screening elec-
tronic health records for ICD- 10 diagnosis codes R73.09
and E11.X, respectively. Patients without diabetes will not
have encounters with these ICD codes.

Recruitment
Enrolment began on 18 July 2023 and will continue until
30  November  2026.  Participants  are  recruited  in  waves
to facilitate the efficient sampling of the study base. The
composition and size of each wave are influenced by the
observed  participation  characteristics  of  race/ethnicity,
gender,  severity  of  T2DM  and  site,  according  to  health
records.  As  recruitment  progresses,  the  composition  of
participants is monitored in terms of diversity and inclu-
sion  and  will  be  adjusted  by  under-  and  oversampling
groups as needed. Recruitment in waves also allows suffi-
cient  time  for  coordinators  to  respond  to  persons  who
are interested and follow- up on mailings with expediency.
For  each  recruitment  wave,  a  contact  pool  is  identified
by screening electronic health records at each site. Indi-
viduals  in  each  pool  are  mailed  a  hardcopy  invitation
letter and sent an invitation email, both personalised to
direct  them  into  our  online  REDCap  recruitment  inter-
face using links, access codes and QR codes. Once in the
REDCap  interface,  individuals  can  read  an  overview  of
the  research  programme,  expectations  for  participation
and answers to Frequently Asked Questions. They are also
given the opportunity to download the informed consent
document, request a call back from study staff, complete
a screening survey for qualification and enroll by signing

Owsley C, et al. BMJ Open 2025;15:e097449. doi:10.1136/bmjopen-2024-097449
Open access

Figure 1  This project is generating an accessible, shared data set that includes a diverse set of health and behavioural
domains and is harmonised across all variables to be artificial intelligence and machine learning ready. Descriptions of variables
are available in tables 1–4. BMI, body mass index; OCTA, optical coherence tomography angiography.

our electronic consent. Once enrolled, they have access
to  study  questionnaires  (see  below).  Those  who  do  not
have access to the internet may call research staff for alter-
native methods of enrolment.

Protocol
The protocol is performed as a single visit lasting between
2.5  and  4 hours,  depending  on  the  participant.  Partic-
ipants  are  volunteers;  therefore,  there  is  selection  bias
known  as  volunteer  bias  which  may  limit  the  generalis-
ability  of  the  results  among  those  who  are  not  repre-
sented  in  the  study  population.  For  most  participants,
informed  consent  is  performed  remotely  by  computer,
tablet or smart phone through a link in the hardcopy invi-
tation letter or the invitation email which directed them
into  our  REDCap  database.  Once  enrolled,  participants
are presented with questionnaires for online completion.
If participants did not select the option to consent elec-
tronically,  it  is  performed  in- person  at  the  start  of  the
visit, followed by the completion of all questionnaires. All
research personnel at the three data sites who recruited,
enrolled  and  collected  data  underwent  training  by  the
data  manager  at  UAB  and  a  detailed  Manual  of  Proce-
dures  (MOP)  for  reference.  Successful  completion  of
a  certification  process  is  required  for  all  coordinators,

involved  mandatory

which
institutional  compliance
training  for  human  subject  research.  In  addition,  prior
to enrolling and testing participants, all coordinators are
required to successfully administer all protocol elements
to at least three volunteer practice subjects, meeting the
standards  outlined  in  the  MOP.  Data  managers  at  each
site  oversee  the  process.  Preliminary  pilot  enrolment
occurred between 18 July 2023 and 30 November 2023 in
order to secure sufficient familiarity for all coordinators.
The formal data collection process began on 1 December
2023. Preliminary data from the pilot enrollment period
were  included  in  the  first  released  data  set  released
in  May  2024.  A  subsequent  version  of  the  data  set  that
includes all data collected in the first year of the study (up
to 31 July 2024) is planned for release in November 2024.
Information on how data can be accessed is available at
https://fairhub.io/

Figure 1 presents information on all the data domains
collected in AI- READI. Questionnaires addressed several
content domains as listed in table 1. Height, weight and
waist  and  hip  circumference  are  measured,  followed  by
calculation  of  the  waist- hip  ratio  and  body  mass  index.
Systolic  blood  pressure,  diastolic  blood  pressure  and
heart  rate  were  measured  twice  separated  by  2 min  with

3

Owsley C, et al. BMJ Open 2025;15:e097449. doi:10.1136/bmjopen-2024-097449
Open access

Table 1  Questionnaires

Questionnaire

Screening questionnaire

Domain addressed

Eligibility (no type 1 diabetes mellitus or pregnancy), diabetes health status,
diabetes treatments (lifestyle, medications), race and ethnicity, biological sex.

Demographic information
Center for Epidemiological Studies Depression Scale—1024 Screens for depressive symptoms.
Problem Areas in Diabetes525
Diabetes score (self- care)26

Date of birth, gender identification, marital status.

Diabetes- associated injuries, lifestyle changes and medical care.

Focuses on diabetes self- management querying dietary habits, exercise,
healthcare and foot care.

Dietary assessment27

Ophthalmic survey

Asks basic questions about food and drink habits over the past few months.

Asks about difficulties with vision and recent eye care.

Smoking, alcohol use, vaping and marijuana use*

Asks about the history of each of these behaviours.

General health

Social Determinants of Health (surveys28

Current Medications with RxNorm codes29

Asks about health history using this question: ‘Has a doctor or other healthcare
professional ever told you that you have/had?’ followed by a list of chronic health
conditions. Responses are yes/no; if yes, may be asked to specify a condition.

Asks about social determinants of health and access to healthcare. Surveys
on food and job insecurity, educational attainment, health insurance coverage,
access to healthcare, housing and neighbourhood environment, perceptions of
discrimination in medical settings and racial/ethnic discrimination (current and
lifetime).

Asks about all prescription and over- the- counter medications currently used.
Includes pills, injections, creams, salves, sprays, eye drops,and dermal patches.

*Marijuana use not assessed for participants living in Alabama because it is an illegal substance.

an  automatic  oscillometric,  medically  approved  device.
A  12- lead  ECG  is  performed  (Philips  Pagewriter  TC30
Cardiograph,  Amsterdam,  The  Netherlands)  while  the
participant sits in a reclining chair or lies supine; the posi-
tion is recorded (0°, 30°, 60° or 90°) relative to the supine
position.  Peripheral  neuropathy  is  performed  using  the
monofilament test11 assessing touch perception on both
feet  with  shoes  and  socks  removed.  During  testing,  the
participants’  eyes  are  closed  while  responding  yes/no
whether  they  feel  the  10 g  filament  in  three  locations
on  each  foot  (10  times  per  location).  A  general  cogni-
tive screener is carried out using the Montreal Cognitive
Assessment12 (MoCA), administered electronically on an
iPad using the MoCA Duo Application (MoCA Cognition,
Quebec,  Canada).  The  total  possible  score  is  30,  with
higher numbers representing better performance.

Visual acuity and contrast sensitivity are assessed under
both  photopic  (daylight)  conditions  and  mesopic  (dim
light) conditions. The right and left eyes are tested sepa-
rately. Photopic letter visual acuity is measured with the
Electronic Visual Acuity tester13 (M&S Technology, Niles,
Illinois) using the instrument’s routine protocol. Mesopic
acuity is measured by the participant viewing the display
through a 2.0 neutral density (ND) filter which reduces
the light level of the test to a mesopic level.14 Visual acuity
is measured while the participant views the display under
best- corrected conditions, expressed as the logarithm of
the  minimum  angle  of  resolution.  Lower  numbers  are
better  resolution.  Photopic  letter  contrast  sensitivity  is
measured  with  the  Mars  chart  (Mars  Perceptrix,  Chap-
paqua, New York)15 using the routine procedure. Mesopic
contrast sensitivity is assessed by viewing through the 2.0

ND filter. Contrast sensitivity is expressed as log sensitivity,
with higher numbers meaning better  sensitivity. Autore-
fraction  data  for  each  eye  are  obtained  (spherical  and
cylindrical  components  expressed  as  diopter  with  cylin-
drical  axis)  using  the  autorefractor’s  standard  protocol
(Topcon  KR  800,  Topcon  Healthcare,  Oakland,  New
Jersey).

Blood  (non- fasting)  and  urine  samples  are  collected
from  participants  during  the  visit.  Whole  blood,  blood
derivatives  and  urine  are  used  for  clinical  lab  testing,
summarised in table 2. Clinical Laboratory Improvement
Amendments- certified  laboratories  local  to  each  data
collection  site  perform  complete  blood  count  (CBC)
testing on fresh whole blood samples; all other lab tests
are performed by the University of Washington Nutrition
and  Obesity  Research  Center  (NORC)  testing  facility
using  stored,  frozen  samples.  Additionally,  blood  deriv-
atives  are  sent  to  the  biorepository  at  UAB’s  Center  for
Clinical  and  Translational  Science  (CCTS)  and  will  be
available  to  researchers  for  future  ancillary  studies  (see
table  3  for  a  summary  and  detailed  discussion  of  the
biorepository).

The  retinal  imaging  protocol  is  designed  to  capture
imaging  from  multiple  devices.  Both  eyes  are  imaged
on  each  participant.  Table  4  details  the  scans  and  data
formats  used  for  each  retinal  imaging  device.  The
imaging  devices  were  chosen  due  to  their  different,  yet
partially overlapping, modalities. The specific scan types
were chosen to provide widespread coverage of the retina
with the overarching goal of detecting pixel- level differ-
ences in pathology. All retinal images are collected under
dilated conditions, except for the Optomed (see table 4).

4

Owsley C, et al. BMJ Open 2025;15:e097449. doi:10.1136/bmjopen-2024-097449
Open access

Table 2  Clinical laboratory tests and biorepository specimens

Test

EDTA plasma tests

Units

Reference range*

Rationale for inclusion

 N- terminal pro- B- type natriuretic peptide

pg/mL

Varies by age

Severity and outcome predictor of heart
failure

 Troponin- T

 C peptide

 Insulin

Serum tests

ng/L

ng/mL

ng/mL

Female<11; male<16 Marker of myocardial injury

1.1–4.4

0.0–24.9

Indicator of insulin production

Marker for diabetes

 C reactive protein, high sensitivity

mg/L

0.0–10.0

 Total cholesterol

 Triglycerides

 Hign- density lipoprotein cholesterol

 Low- density lipoprotein cholesterol
(calculated)

 Glucose

 Blood urea nitrogen

 Creatinine

Blood urea nitrogen/creatinine ratio

 Sodium

 Potassium

 Chloride

 Carbon dioxide, total

 Calcium

 Protein, total

 Albumin

 Globulin, total (calculated)

 A/G ratio (calculated)

 Bilirubin, total

 Alkaline phosphatase

 Aspartate aminotransferase

 Alanine aminotransferase

Whole blood tests

 HbA1c

 White blood cell

 Red blood cell

 Haemoglobin

 Haematocrit

 MCV

 MCH

 MCHC

 RDW

 Platelets
Urine tests

mg/dL

mg/dL

mg/dL

mg/dL

mg/dL

mg/dL

mg/dL

mEq/L

mEq/L

mEq/L

mEq/L

mg/dL

g/dL

g/dL

g/dL

<200

<150

>39

<130

62–125

8.0–21.0

Female: 0.38–1.02;
male: 0.51–1.18

135–145

3.6–5.2

98–108

22–32

8.9–10.2

6.0–8.2

3.5–5.2

Inflammation marker; risk factor for
T2DM

Cardiovascular disease risk factor

Cardiovascular disease risk factor

Cardiovascular disease risk factor

Cardiovascular disease risk factor

Marker for diabetes

Marker for kidney and liver function

Marker for kidney function

Indicator of kidney function

Marker of metabolic health

Marker of metabolic health

Marker of metabolic health

Marker of metabolic health

Marker of metabolic health

Useful tool for overall health status

Useful tool for overall health status

Useful tool for overall health status

Useful tool for overall health status

mg/dL

0.2–1.3

Marker of liver health

IU/L

IU/L

IU/L

%

×10E3/μL

×10E6/μL

g/dL

%

fL

pg

g/dL

%

Varies by age

Marker of liver health

9–38

Female age 7–33 and
male age 0–49: 10–64;
male age 50+: 10–48

Marker of liver health

Marker of liver health

4.0–6.0

3.4–10.8

3.77–5.80

11.1–17.7

34.0–51.0

79–97

26.6–33.0

31.5–35.7

11.6–15.4

Marker for diabetes

Useful tool for overall health status

Useful tool for overall health status

Useful tool for overall health status

Useful tool for overall health status

Useful tool for overall health status

Useful tool for overall health status

Useful tool for overall health status

Useful tool for overall health status

×10E3/μL

150–450

Useful tool for overall health status

Continued

5

Owsley C, et al. BMJ Open 2025;15:e097449. doi:10.1136/bmjopen-2024-097449



































Open access

Table 2  Continued

Test

 Urine creatinine

 Urine albumin

Biorepository specimens

 Buffy coats (from EDTA anticoagulated
vacutainers)

 Genomic DNA

 Plasma (from EDTA anticoagulated
vacutainers)

 Serum

 PAXgene vacutainers
 Peripheral blood mononuclear cells

Units

mg/dL

mg/dL

–

–

–

–

–
–

Reference range*

Rationale for inclusion

N/A

N/A

–

–

–

–

–
–

Urinary biomarker for early diabetic
nephropathy

Urinary biomarker for early diabetic
nephropathy

DNA isolations for future studies such as
whole genome sequencing

Genomic applications

Future proteomics and metabolomics
studies

Future proteomics and metabolomics
studies

Future RNA isolations
Peripheral blood mononuclear cells for
future immunological studies; generation
of iPSCs

*Reference range is a set of values that represent low and high ends of results that are considered normal. The above
reference ranges were provided by the testing laboratories.
HbA1c, hemoglobin A1c; iPSCs, induced pluripotent stem cells; MCH, mean corpuscular hemoglobin; MCHC, mean
corpuscular hemobglobin concentration; MCV, mean corpuscular volume; RDW, red blood cell distribution width.

All participant images are exported from imaging devices
in  their  raw  format  and  some  images  required  conver-
sion to a DICOM standard format prior to upload to an
external AI- READI data storage site (discussed below).

At  the  end  of  their  in- person  visit,  participants  are
sent  home  with  three  data  monitoring  devices  as  listed
in  table  4:  (1)  home  environmental  sensor,  (2)  Garmin
fitness  tracker  and  (3)  Dexcom  Continuous  Glucose
Monitor  (CGM).  Coordinators  provide  instructions  on
how to use these monitoring devices continuously for 10
days. Participants are instructed to return the devices to
study staff by overnight mail (mailing materials and fees
provided by the site) or in- person after 10 days. The envi-
ronmental  sensor  and  the  Garmin  fitness  tracker  were
specifically  chosen  with  participant  privacy  as  a  primary
concern. These devices do not capture video or audio, are
not equipped with GPS location monitoring and are not
synced  with  participant- owned  devices.  Participants  are
masked to the results from the Dexcom CGM during the
monitoring  period.  Participants  return  the  monitoring
devices with a data sheet indicating which wrist the fitness
tracker was worn on, their dominant hand and where the
environmental sensor was placed in the home.

The  following  results  are  returned  to  participants
either as they leave the visit or later by HIPAA- compliant,
encrypted  email.  Results  from  heart  rate,  blood  pres-
sure  and  visual  acuity  assessments  are  provided  to  the
participant  on  an  exam  card  before  they  leave  the  visit,
along  with  instructions  in  lay  language  for  interpreting
the data and any follow- up recommendations. Following
the 10- day at- home monitoring period, a Dexcom report
providing average blood glucose levels (hourly, daily and

overall) and an overall Glucose Management Indicator is
provided  by  encrypted  email.  At  the  time  of  the  yearly
data  release  (see  below),  laboratory  test  results  are  sent
with  information  on  normative  values  by  Health  Insur-
ance Portabillity and Accountability- compliant, encrypted
email.

Coordinators  alert  participants  to  several  incidental
findings before they leave a visit. Participants are recom-
mended  to  go  for  emergency  care  for  the  following
reasons: individuals with systolic blood pressure readings
of  over  180 mm  Hg  or  less  than  100 mm  Hg  (with  any
symptoms  of  haemodynamic  instability),  diastolic  blood
pressure readings of greater than 120 mm Hg or less than
60 mm Hg (along with any symptoms of haemodynamic
instability) or heart rate readings of greater than 100 bpm
or less than 60 bpm (if not known to be usual for them
and  with  any  symptoms  of  haemodynamic  instability).
Retinal imaging technicians are trained to detect certain
conditions that are potentially life or vision threatening
(retinal detachment, tumour and optic disc oedema). If
one of these conditions is suspected, immediate follow- up
with  the  participant  is  performed  with  a  referral  to  the
emergency  department  (disc  oedema)  or  referral  to  an
ophthalmologist for immediate care (retinal detachment,
tumour).

Data management
REDCap is used for collecting patient- reported data from
questionnaires  and  the  following  clinical  data  from  the
visit: medications assessment, vitals, visual acuity, contrast
sensitivity, monofilament test results and CBC test results.
Data from the following devices are exported in their raw

6

Owsley C, et al. BMJ Open 2025;15:e097449. doi:10.1136/bmjopen-2024-097449












Open access

Table 3  Biospecimen collection including processing and purpose

Specimen collection

Sample type

Location

Processing details

Purpose

EDTA vacutainers

Whole blood

Whole blood

Plasma

Local clinical lab

UW NORC*

UW NORC*

Plasma

Biobanking at UAB CCTS

Buffy coats

Biobanking at UAB CCTS

None

None

Processed locally at data site
on the day of collection

Complete blood count analysis

Testing for HbA1c

Testing for NT- proBNP,
troponin- T, C- peptide and
insulin

Processed locally at data site
on the day of collection

Future proteomics and
metabolomics studies

Processed locally at data site
on the day of collection

DNA extractions for future
genomics

Serum separator
vacutainers

Serum

Serum

Biobanking at UAB CCTS

Processed locally at data site
on the day of collection

Future proteomics and
metabolomics studies

UW NORC*

Processed locally at data site
on the day of collection

Testing for CRP- HS, glucose,
BUN, creatinine, carbon
dioxide, total protein, albumin,
globulin, bilirubin, alkaline
phosphatase, AST, ALT,
electrolytes and lipids

Urine collection kit

Urine

UW NORC*

Processed locally at data site
on the day of collection

Testing for creatinine and
albumin

CPT mononuclear cell
preparation tubes

Peripheral blood
mononuclear cells

Biobanking at UAB CCTS

UAB—processed locally on
day of collection

Future immunological studies;
generation of iPSCs

PAXgene RNA vacutainers Stabilised,

Biobanking at UAB CCTS

unisolated RNA

UCSD/UW—tubes sent to
UAB for next- day processing

UAB—processed locally on
day of collection

UCSD/UW—tubes sent to
UAB for next- day processing

Future gene expression studies

*Samples analysed by the UW NORC laboratory (serum, plasma, whole blood and urine) are first stored locally at UAB and UCSD at −70°C, then
batch- shipped to UW on dry ice approximately once per quarter.
ALT, alanine aminotransferase; AST, aspartate aminotransferase; BUN, blood urea nitrogen; CRP- HS, C reactive protein, high sensitivity; Electrolytes,
sodium, potassium, chloride, calcium; Lipid analysis, total cholesterol, high- density lipoprotein cholesterol, low- density lipoprotein cholesterol,
triglycerides; NT- proBNP, N- terminal pro- B- type natriuretic peptide; UAB, University of Alabama at Birmingham; UAB CCTS, University of Alabama
at Birmingham Centre for Clinical and Translational Science; UCSD, University of California, San Diego; UW, University of Washington; UW NORC,
University of Washington Nutrition Obesity Research Center.

format to local storage: ECG (.xml), MoCA Duo applica-
tion (total score, section subscores, Memory Index Score
and task completion times; .csv), retinal imaging (table 4)
and at- home monitoring devices (environmental sensors,
CGMs, fitness trackers; Table 3). Blood and urine testing
by the NORC lab is provided in .csv format. All data are
mapped to applicable data standard formats, such as the
Observational  Medical  Outcomes  Partnership  Common
Data Model for clinical data and the DICOM format for
retinal  imaging.  All  data  are  uploaded  at  regular  inter-
vals to the AI- READI- specific data management platform
called FAIRhub (https://fairhub.io/), a Microsoft Azure
cloud- based platform developed for this project. For some
devices,  the  data  are  transformed  from  its  proprietary
state  into  a  standard  model  format  prior  to  upload  to
FAIRhub (tables 3 and 4). The Data Manager at each data
site is responsible for the quality control of all data before
uploading  to  FAIRhub.  All  data  are  stored  and  shared
using  FAIRhub  as  ‘AI- ready’  which  enables  immediate

AI/ML research without reformatting and preprocessing.
More  details  about  the  AI- readiness  of  the  data  set  are
provided  in  the  data  set  documentation  (https://docs.
aireadi.org/).

Data availability
A  more  detailed  explanation  of  FAIRhub  was  provided
previously1  and  at  https://aireadi.org.  Deidentified
data  are  released  approximately  yearly  in  two  data  sets:
a  controlled  access  set  and  publicly  accessible  set  with
fewer requirements for access. The controlled access data
set  contains  all  data  available  in  the  publicly  accessible
set,  plus  more  sensitive  data  that  are  only  available  for
use by scientists whose institutions have completed legal
and privacy use agreements with the AI- READI research
programme.  The  publicly  accessible  data  set  from  pilot
data collection was made available in May 2024. All data
collected  through  31  July  2024  were  made  available
in  November  2024.16  Information  regarding  data  set

7

Owsley C, et al. BMJ Open 2025;15:e097449. doi:10.1136/bmjopen-2024-097449


Open access

Table 4  Retinal imaging and at- home monitoring devices

Imaging device

Scans collected

Data format

Manufacturer

Aurora IQ

EIDON widefield truecolor
confocal fundus system

DICOM

DICOM

Fundus: macula and disc centred
(CFP) colour fundus photo (undilated)

Ultra- wide Field Central CFP
Ultra- wide Field Nasal CFP
Ultra- wide Field Temporal CFP
Ultra- wide Field Central IR
Ultra- ide Field Central FAF
Mosaic of Ultra- wide Field images

Number of images
per participant

4

Optomed
(Oulu, Finland)

iCare USA, Inc (Raleigh, North
Carolina)

12

Spectralis HRA optical
coherence tomography
(OCT) and optical coherence
tomography angiography
(OCTA)

Optic Nerve Head- Radial Circle- High
Resolution (OCT)
Posterior Pole Macula- HR- 61 lines
(OCT)
Macula- 20×20- HS- 512 lines (OCTA)

Maestro2 3D OCT- 1

3D Wide (H) 12×9–512×128 (OCT)
3D Macula 6×6–512×128 (OCT)
Macula 6×6–360×360 (OCTA)

Triton DRI OCT

Cirrus 5000

3D(H)+Rad 12×9–512×256 (OCT)
Macula 6×6–320×320 (OCTA)
Macula 12×12–512×512 (OCTA)

Disc Cube, −200×200 (OCT)
Macula Cube, −512×128 (OCT)
Macula, −6×6 (OCTA)
Disc, −6×6 (OCTA)

Fluorescence Lifetime Imaging
Ophthalmoscopy

Macula- centred

DICOM

Heidelberg Engineering GmbH
(Heidelberg, Germany)

6

Topcon Healthcare (Itabashi-
ku, Tokyo, Japan)

Topcon Healthcare (Itabashi-
ku, Tokyo, Japan)

6

6

Carl Zeiss Meditec AG (Jena,
Germany)

8

Heidelberg Engineering GmbH
(Heidelberg, Germany)

2

Exported in
.fda file format;
converted to
DICOM for the
data set

Exported in
.fda file format;
converted to
DICOM for the
data set

DICOM

Exported in
.sdt file format;
converted to
DICOM for the
data set

Monitoring device

Data variables collected

Data format

Manufacturer

Dexcom G6 Continuous
Glucose Monitor

Blood glucose measurements (mg/dL)
every 5 min

.csv

Dexcom, Inc. (San Diego, California)

Garmin VivoSmart 5 Physical
Activity Monitor

Environmental Sensor

Exported in
.FIT file format;
converted
to mHealth
standard for the
data set

.csv

Number of steps
Heart rate
Sleep duration (circadian and diurnal
rhythm)
Oxygen saturation

Ambient temperature
Relative humidity
Nitrogen oxides (NO and NO2)
Volatile organic compounds
Particulate matter (PM1.0, PM2.5,
PM4, PM10)
Multi- spectral light intensity
measurements (11)

Garmin Ltd (Schaffhausen, Switzerland)

Custom designed, Karalis Johnson Retina Center,
University of Washington (Seattle, Washington)30

availability  and  access  may  be  found  at  https://aireadi.
org/goals/data-sharing. The requirements for controlled
access data sets are being currently developed by the Data
Access Committee.

Biospecimen processing and repository
Blood  (53 mL)  is  collected  for  clinical  lab  assays  and
biobanking  purposes,  and  a  urine  specimen  is  also
collected  during  the  study  encounter  (table  2).  The
collection  materials,  sample  type,  specimen  handling

and  location  details  are  summarised  in  table  3.  Local
processing  for  plasma,  serum  and  buffy  coats
is
performed  using  standardised  operating  procedures,
ensuring consistent handling of these biospecimens for
subsequent  biobanking  and  clinical  lab  analyses.  One
EDTA vacutainer is dedicated for delivery to a local clin-
ical  testing  lab  for  CBC  analysis.  Whole  blood,  plasma,
serum and urine collected for central clinical lab assays
are  batch- shipped  to  the  UW  NORC  on  a  regular  basis

8

Owsley C, et al. BMJ Open 2025;15:e097449. doi:10.1136/bmjopen-2024-097449
(table  3).  Likewise,  plasma,  serum  and  buffy  coats
collected at UW and UCSD are batch- shipped to the UAB
CCTS for integration into a central study- wide biobank.
Genomic  DNA  extractions  from  stored  buffy  coats  are
performed at UAB CCTS. Because of the complexity in
isolating and cryopreserving peripheral blood mononu-
clear  cells  (PBMCs),  centralised  processing  of  the  CPT
tubes occurs at the UAB CCTS. While this strategy does
introduce  an  additional  time  variable  for  isolation  and
cryopreservation of PBMC from participants recruited at
UW and UCSD, the ability to pre- centrifuge these tubes
locally  followed  by  overnight  delivery  to  UAB  allows
for  consistency  in  initial  vacutainer  handling.  Finally,
the  PAXgene  RNA  vacutainers  are  stable  for  up  to  72
hours  after  collection,  allowing  for  overnight  shipment
of  vacutainers  from  UW  and  UCSD  to  the  UAB  CCTS.
Biobanked samples will eventually be available to scien-
tists  according  to  procedures  and  policies  that  are  in
development.

AI- READI  has  several  strengths  which  make  it  highly
suitable  for  AI/ML  studies  on  T2DM.  The  targeted
sample size of 4000 persons is the largest publicly acces-
sible data set currently available containing many multi-
domain  variables  relevant  to  T2DM.  Furthermore,  after
1 year  of  enrolment,  the  sites  are  currently  on  schedule
for reaching a final sample of 4000 persons within 3 years
of  data  collection.  The  sample  will  be  approximately
balanced  with  respect  to  Asians,  blacks,  Hispanics  and
whites, a demographic improvement over many previous
epidemiological  and  clinical  trials  on  T2DM.  A  meta-
epidemiological review of T2DM studies published from
2000  to  2020  indicated  that  samples  lacked  racial  and
ethnic  diversity,17  thus  making  AI- READI  an  important
step forward in addressing these demographic inequities
in research. The study design process enacted ethical and
equitable  data  collection  and  management  practices,  as
well as data sharing with adherence to the FAIR principles.
Challenges  will  also  be  encountered.  It  is  a  well-
documented  observation
that
recruiting  black  men  into  studies  is  more  difficult  than
recruiting  other  populations,18  yet  this  population
represents  a  disproportionate  percentage  of  burden  in
T2DM.19  Factors  such  as  medical  mistrust  and  fear  of
safety  stem  from  historical  events  such  as  the  Tuskegee
syphilis study.20 Researchers have identified strategies to
encourage  interest  among  black  men  to  participate  in
research such as tailoring printed materials and using a
personalised,  participatory  approach  to  recruitment.21
These  trends  for  under- recruitment  and  for  disease
burden  have  also  been  reported  for  Hispanic22  and
Asian  persons.23  If  sample  balancing  on  race/ethnicity
emerges as a challenge in AI- READI, we will implement
new  recruitment  strategies  to  overcome  the  imbalance.
Requiring that eligible participants speak, read and under-
stand English may have created challenges in recruiting
Hispanic and/or Asian persons. A limitation of our biore-
pository  is  that  there  are  a  finite  number  of  samples  to
share with scientists interested in using them in research.

in  health  research

Open access

Procedures for reviewing and prioritising written requests
will be developed before the biorepository is complete.

In  summary,  AI- READI  is  an  NIH  Common  Fund
Bridge2AI- supported data collection effort to generate a
4000- person data set with T2DM that is suitable for AI/
ML.  It  is  triple- balanced  with  respect  to  race/ethnicity,
biological  sex  and  T2DM  severity  (including  those  who
do not have T2DM). It is multimodal with over 10 variable
domains that can be associated with adverse conditions in
T2DM, yet it is hypothesis agnostic. AI- READI emphasises
making  data  ‘AI- ready’  and  establishing  standards,  best
practices and guidelines for collection, preparation and
sharing  of  the  data.  This  approach  can  be  viewed  as  an
example  for  future  efforts  to  develop  other  health  data
sets  targeted  for  AI/ML.  AI- READI  opens  the  door  for
novel insights into understanding T2DM salutogenesis.

ETHICS AND DISSEMINATION
AI- READI aims to establish standards, best practices and
guidelines for collection, preparation and sharing of the
data for the purposes of AI/ML, including guidance from
bioethicists.  Following  FAIR  principles,  AI- READI  can
be viewed as a model for future efforts to develop other
medical/health data sets targeted for AI/ML. AI- READI
opens  the  door  for  novel  insights  into  understanding
T2DM  salutogenesis.  The  AI- READI  Consortium  are
disseminating  the  principles  and  processes  of  designing
and implementing the AI- READI data set through publi-
cations.  Those  who  download  and  use  AI- READI  data
are  encouraged  to  publish  their  results  in  the  scientific
literature.

X Cynthia Owsley @cynthiaowsley

Collaborators  AI- READI Consortium. Amir Bahmani, Sally Baxter, Edward Boyko,
Christopher Chute, Aaron Cohen, Jorge Contreras, Garrison Cottrell, Virginia de
Sa, Jeffrey Edberg, Nicole Ehrhardt, Nicholas Evans, Irl Hirsch, Michelle Hribar,
Samantha Hurst, Aaron Lee, Cecilia Lee, T.Y. Alvin Liu, Bonnie Maldenado, Gerald
McGwin Jr., Shannon McWeeney, Cynthia Owsley, Bhavesh Patel, Sara Singer,
Michael Snyder, Bradley Voytek, Joseph Yracheta, Linda Zangwill.

Contributors  CO is the guarantor. All authors are submitting authors. They meet
the four criteria of the IJMJE Recommendations 2019 as follows: substantial
contributions to the conception or design of the work; drafting the work or revising
it critically for important intellectual content; final approval of the version to be
published; and agreement to be accountable for all aspects of the work in ensuring
that questions related to the accuracy or integrity of any part of the work are
appropriately investigated and resolved. The AI- READI Consortium are collaborators
(group authorship). CO is the corresponding author.

Funding  This research is supported by National Institutes of Health grants
OT2OD032644, P30DK035816, and UL1TR003096 and Research to Prevent
Blindness.

Competing interests  CO: Johnson & Johnson Vision, Sanofi (consultant). SLB:
Topcon (consultant, travel). LZ: AbbVie, Topcon Medical Systems (consultant);
AISight Health Inc. (stock or stock options). DM, GMcG, JE, JO, CL: none.

Patient and public involvement  Patients and/or the public were involved in the
design, or conduct, or reporting, or dissemination plans of this research. Refer to
the Methods section for further details.

Patient consent for publication  Not applicable.

Provenance and peer review  Not commissioned; externally peer reviewed.

Open access  This is an open access article distributed in accordance with the
Creative Commons Attribution Non Commercial (CC BY- NC 4.0) license, which

9

Owsley C, et al. BMJ Open 2025;15:e097449. doi:10.1136/bmjopen-2024-097449
Open access

permits others to distribute, remix, adapt, build upon this work non- commercially,
and license their derivative works on different terms, provided the original work is
properly cited, appropriate credit is given, any changes made indicated, and the use
is non- commercial. See: http://creativecommons.org/licenses/by-nc/4.0/.

ORCID iDs
Cynthia Owsley http://orcid.org/0000-0003-3424-011X
Sally L Baxter http://orcid.org/0000-0002-5271-7690

REFERENCES
  1  AI- READI Consortium. AI- READI: rethinking AI data collection,

preparation and sharing in diabetes research and beyond. Nat Metab
2024;6:2210–2.

  2  Kalra S, Baruah MP, Sahay R. Salutogenesis in Type 2 Diabetes
Care: A Biopsychosocial Perspective. Indian J Endocrinol Metab
2018;22:169–72.

  3  International Diabetes Federation. IDF diabetes atlas. 2021.
  4  Bonora E, DeFronzo RA. Diabetes complications, comorbidities and

related disorders. Springer, 2020.

  5  Hill- Briggs F, Fitzpatrick SL. Overview of Social Determinants
of Health in the Development of Diabetes. Diabetes Care
2023;46:1590–8.

  6  Khunti N, Khunti N, Khunti K. Adherence to type 2 diabetes

management. Br J Diabetes 2019;19:99–104.

  7  Cheng YJ, Kanaya AM, Araneta MRG, et al. Prevalence of Diabetes
by Race and Ethnicity in the United States, 2011- 2016. JAMA
2019;322:2389–98.

  8  Unnikrishnan R, Pradeepa R, Joshi SR, et al. Type 2 Diabetes:
Demystifying the Global Epidemic. Diabetes 2017;66:1432–42.

  9  Wilson C, Alam R, Latif S, et al. Patient access to healthcare

services and optimisation of self- management for ethnic minority
populations living with diabetes: a systematic review. Health Soc
Care Community 2012;20:1–19.

 10  Office of Minority Health. Diabetes and american Indians/Alaskan
natives (US department of health and human services). 2021.
Available: https://minorityhealth.hhs.gov/diabetes-and-american-
indiansalaska-natives

 11  Boulton AJM, Armstrong DG, Albert SF, et al. Comprehensive foot
examination and risk assessment: a report of the task force of the
foot care interest group of the American Diabetes Association,
with endorsement by the American Association of Clinical
Endocrinologists. Diabetes Care 2008;31:1679–85.

 12  Nasreddine ZS, Phillips NA, Bédirian V, et al. The Montreal Cognitive

Assessment, MoCA: a brief screening tool for mild cognitive
impairment. J Am Geriatr Soc 2005;53:695–9.

 13  Beck RW, Moke PS, Turpin AH, et al. A computerized method
of visual acuity testing: adaptation of the early treatment of

diabetic retinopathy study testing protocol. Am J Ophthalmol
2003;135:194–205.

 14  Owsley C, Swain TA, McGwin G Jr, et al. How Vision Is Impaired
From Aging to Early and Intermediate Age- Related Macular
Degeneration: Insights From ALSTAR2 Baseline. Transl Vis Sci
Technol 2022;11:17.

 15  Arditi A. Improving the design of the letter contrast sensitivity test.

Invest Ophthalmol Vis Sci 2005;46:2225–9.

 16  AI- READI Consortium. Flagship dataset of type 2 diabetes from the
AI- READI project (1.0.0) FAIRhub. Available: https://docs.aireadi.org/

 17  Ahmed R, de Souza RJ, Li V, et al. Twenty years of participation

of racialised groups in type 2 diabetes randomised clinical trials: a
meta- epidemiological review. Diabetologia 2024;67:443–58.

 18  Randolph S, Coakley T, Shears J. Recruiting and engaging African-

American men in health research. Nurse Res 2018;26:8–12.

 19  Liburd LC, Namageyo- Funa A, Jack L Jr. Understanding 'masculinity'
and the challenges of managing type- 2 diabetes among African-
American men. J Natl Med Assoc 2007;99:550–2.

 20  Scharff DP, Mathews KJ, Jackson P, et al. More than Tuskegee:

understanding mistrust about research participation. J Health Care
Poor Underserved 2010;21:879–97.

 21  Woods VD, Montgomery SB, Herring RP. Recruiting Black/African
American men for research on prostate cancer prevention. Cancer
2004;100:1017–25.

 22  O Rojo M, Jing J, Wells C, et al. Hispanics’ Perceptions of

Participation in Research Studies and Solutions for Improvement in
Participation. J Family Med Community Health 2024;11:1–9.
 23  Lim J- W, Paek M- S. Recruiting Chinese- and Korean- Americans in
Cancer Survivorship Research: Challenges and Lessons Learned. J
Cancer Educ 2016;31:108–14.

 24  Andresen EM, Malmgren JA, Carter WB, et al. Screening for

depression in well older adults: evaluation of a short form of the
CES- D (Center for Epidemiologic Studies Depression Scale). Am J
Prev Med 1994;10:77–84.

 25  McGuire BE, Morrison TG, Hermanns N, et al. Short- form measures

of diabetes- related emotional distress: the Problem Areas in Diabetes
Scale (PAID)- 5 and PAID- 1. Diabetologia 2010;53:66–9.

 26  Hashim MJ, Nurulain SM, Riaz M, et al. Diabetes Score questionnaire
for lifestyle change in patients with type 2 diabetes. Clin Diabetes
2020;9:379–86.

 27  Paxton AE, Strycker LA, Toobert DJ, et al. Starting the conversation
performance of a brief dietary assessment and intervention tool for
health professionals. Am J Prev Med 2011;40:67–71.

 28  Hamilton CM, Strader LC, Pratt JG, et al. The PhenX Toolkit: get the
most from your measures. Am J Epidemiol 2011;174:253–60.

 29  Unified Medical Language System (UMLS). RxNorm. Available:

https://www.nlm.nih.gov/research/umls/rxnorm/index.html [Accessed
7 Aug 2024].

 30  Shaffer J, Gim N, Wei R, et al. Portable environmental sensor

enabling studies of exposome on ocular health. Invest Ophthalmol
Vis Sci 2024;65:6370.

10

Owsley C, et al. BMJ Open 2025;15:e097449. doi:10.1136/bmjopen-2024-097449


================================================================================

FILE: s42255-024-01165-x_row3.txt
PATH: data/preprocessed/individual/AI_READI/s42255-024-01165-x_row3.txt
SIZE: 18165 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: AI_READI
Source ID: nature_metabolism_publication
Source type: publication
Source URL: https://www.nature.com/articles/s42255-024-01165-x.pdf
Raw file: data/raw/AI_READI/s42255-024-01165-x_row3.pdf
--------------------------------------------------------------------------------
AI-READI: rethinking AI data collection,
preparation and sharing in diabetes research
and beyond

https://doi.org/10.1038/s42255-024-01165-x

AI-READI Consortium

Here, we introduce Artificial Intelligence
Ready and Equitable Atlas for Diabetes
Insights (AI-READI), a multidisciplinary
data-generation project designed to create
and share a multimodal dataset optimized
for artificial intelligence research in type 2
diabetes mellitus.

The AI-READI project is one of four Data Generation Projects (DGPs)
funded by Bridge2AI (https://commonfund.nih.gov/bridge2ai), a
new NIH Common Fund Program aimed at setting the stage for the
widespread adoption of AI in healthcare research. The primary goal
of AI-READI is to collect and publicly share a multimodal, AI-ready
dataset for studying the pathogenesis and salutogenesis (that is, the
pathway from disease to health) of type 2 diabetes mellitus (T2DM).

T2DM is a growing public health threat, affecting >6% of the
world’s population, and at increasingly younger ages. Certain popu-
lations experience a greater burden of disease, exacerbating health
disparities. Our data collection effort is centred around the prin-
ciple of collecting a large multimodal dataset uniformly balanced
for T2DM severity from a diverse group of participants so as to bet-
ter understand this complex multifactorial disease using AI. Given
the  complexity  of  T2DM,  AI-based  approaches  and  multimodal
model  development  may  improve  our  understanding  of  T2DM.

 Check for updates

However, a major barrier has been the lack of diverse datasets that are
AI ready1 for training AI models, which the AI-READI project is designed
to address.

Here, we present the AI-READI dataset, give preliminary insights
into our blueprint for making data AI-ready and provide our ongoing
strategies for training a diverse workforce at the intersection of AI
and biomedical research.

Introduction to the AI-READI dataset
AI-READI is enrolling 4,000 participants at three study sites over 4 years
(2022–2026) (Supplementary Fig. 1). We plan to balance the racial and
ethnic diversity of participants to address the under-representation in
past research studies of groups that bear a higher burden of the disease.
Having three data collection sites will increase geographic represen-
tation. Participants with a range of health states will be included to
facilitate AI-based research into salutogenesis.

The recruitment process consists of a selection of individuals
based on the available clinical diagnosis using ICD-10 codes and demo-
graphic data from the electronic health records of study sites. Each
participant completes the study protocol, which includes question-
naires, on-site data collection and at-home data collection (Fig. 1).
Blood is being collected to establish a large repository of participant
plasma, serum and buffy coats at the University of Alabama at Birming-
ham. In addition, we are collecting PaxGene RNA tubes for future RNA
isolation and cryopreserving peripheral blood mononuclear cells for
future studies that could include functional immunologic studies or
generation of induced pluripotent stem cells.

Pre-visit
(1 hour, at home)

Self-reporting surveys
• Initial screening
• Demographic
• Center for Epidemiological Studies

Depression Scale (CES-D)-10

• Problem areas in diabetes questionnaire (PAID-5)
• Diabetes score
• Diet
• Smoking history
• Alcohol use, vaping and marijuana use
• General health
• Social determinants of health (SDoH)
• Visual impairment and eye care access

Data collection

On-site visit
(3 to 4 hours)

Current medicine list

Driving record
(accident report)

Monofilament test

Vision testing
(lensometer, autorefraction
best corrected visual acuity,
letter contrast sensitivity)

Retinal imaging
(undilated/dilated fundus
photography, FLIO, OCT,
OCTA)

Blood test
(NT-proBNP, C-peptide, troponin-T,
HbA1c, lipid panel, CRP, CMP12)

Physical assessment
(height, weight, waist and hip
circumferences, blood pressure,
heart rate)

Biospecimen
(blood)

Urine test
(albumin, creatinine)

ECG

Cognitive screening

Post-visit
(10 days, at home)

Continuous glucose monitoring

Physical activity monitoring
(heart rate, respiration rate, activity type, SpO2,
stress level, sleep phases)

Environmental measurements
(temperature, humidity, light spectrometry, PM1.0,
PM2.5, PM4.0, PM10.0, NOx, volatile organic
compounds)

Fig. 1 | Data collection protocol for each participant in AI-READI.
Each participant who enrols in the AI-READI study completes the study
protocol, which includes pre-visit questionnaires, on-site data collection
and post-visit data collection for 10 days using a Dexcom G6 for continuous
glucose monitoring, Garmin Vivosmart 5 for physical activity monitoring and

a custom-built environmental sensor for indoor environmental parameter
measurements. ECG, electrocardiogram; FLIO, fluorescence lifetime imaging
ophthalmoscopy; OCT, optical coherence tomography; OCTA, optical coherence
tomography angiography ; PM, particulate matter.

nature metabolism

Volume 6 | December 2024 | 2210–2212 | 2210

Comment
Each data collection site uploads data at regular intervals to a
cloud-based data management platform (FAIRhub) that is being devel-
oped for this project. Data are released annually, and each version is
accessible through FAIRhub. At each release, two datasets are made
available: a controlled-access set and a publicly accessible set with
fewer requirements for access. The public set is stripped of Protected
Health Information (PHI), as defined by the HIPAA (Health Insurance
Portability and Accountability Act) Privacy Rule via the “Safe Harbor”
method, as well as information related to the sex and race/ethnicity of
the participants to prevent stigmatization of findings. The first version
of the public set was recently shared using a new licence developed
as part of this project that allows reuse for any purpose but includes
restrictions to protect participant privacy2.

Blueprint for future data-generating projects
An additional aim of the AI-READI project is to share our blueprint for
collecting and sharing AI-ready datasets so that future DGPs can follow
them (Supplementary Fig. 2), including guidelines for project manage-
ment, as well as data collection, management and sharing (https://
zenodo.org/communities/aireadi; https://github.com/AI-READI;
https://aireadi.org).

The integration of multiple disciplines has become a cornerstone
for discoveries and innovation in biomedical research, so the successful
preparation of an AI-ready dataset must be a multidisciplinary team
effort3,4. Therefore, an essential component of AI-READI is to imple-
ment team science strategies to understand the interaction patterns
of our multi-team systems and then disseminate our findings (Sup-
plementary Fig. 3).

We believe that, at a high level, making data AI ready involves two
main sets of considerations: technical and ethical. The FAIR (find-
able, accessible, interoperable, reusable) principles provide high-level
instructions for making data technically ready for reuse by AI systems2.
As part of the project, we are developing guidelines for making data
types from the AI-READI project FAIR. We are using existing standards
when available and working with their maintainers to extend them
when necessary (Supplementary Table 1). We are also developing new
standards, such as the Clinical Dataset Structure (CDS), a simple and
intuitive way to organize clinical research datasets and include struc-
tured metadata (https://cds-specification.readthedocs.io).

Data used in developing AI systems are often a major source of
downstream ethical issues5. To prevent such issues, ethical, legal and
social implications are considered at every stage of the project cycle.
Documenting practices regarding the creation, use and maintenance
of clinical research datasets is critically important so that AI developers
can easily understand the provenance and intended use of the data. We
are reviewing existing documentation approaches such as datasheets
and healthsheets to identify the most suitable one for biomedical
data6,7. Al algorithms can infer individual patient characteristics and
could facilitate individual-level identification. To prevent reidentifi-
cation attempts, we are implementing a robust data dissemination
system, a strict licence agreement and data watermarking for both
public and controlled sets.

Preparing and sharing AI-ready datasets can rapidly become time
consuming and difficult for data-collecting researchers. To address this
problem, we are developing a cloud-based data management, curation
and sharing platform called FAIRhub (https://fairhub.io/). Inspired by
existing user-friendly ‘FAIRification’ tools, the platform is being devel-
oped to include intuitive user interfaces and a suite of tools to simplify
and automate the implementation of our guidelines for making data AI

ready8. After testing FAIRhub for the AI-READI dataset, we anticipate
making the platform available for future DGPs desiring to manage and
share AI-ready clinical research datasets.

American Indians and Alaska Natives (AI/AN) are unique and highly
identifiable groups that could benefit from precision health AI studies.
Through engagement with the Native Biodata Consortium, this project
will provide a co-learning opportunity for understanding challenges
specific to AI/AN communities and evaluating tools for addressing
them, such as data usage agreements, contracts negotiations, transpar-
ent access agreements and some form(s) of assured return of benefit
and sustainable input, control and monitoring of Tribal data.

Training future AI researchers
Concerns regarding bias and inequity in AI have been partially attrib-
uted to the lack of diversity in the AI workforce. In particular, women
and racial and ethnic minority groups are under-represented.

The Bridge2AI Program has therefore defined workforce devel-
opment as one of its key pillars. AI-READI hosts a year-long research
internship programme (https://shileyeye.ucsd.edu/research/ai_readi)
to provide immersive training for individuals interested in working at
the nexus of biomedicine and AI. The programme includes a 2-week
data science/programming bootcamp followed by a yearlong curricu-
lum of didactic lectures and mentored research. Interns are mentored
by one or more AI-READI investigators and are involved in various
aspects of the project, including developing healthsheets to describe
the AI-READI data, mapping project data into common data models,
and analysing equity and ethical issues. The inaugural cohort in aca-
demic year 2023–2024 consisted of 10 interns, among whom 70% were
women, 30% were Black and, overall, 50% were from under-represented
backgrounds based on NIH criteria. The programme is continuing
broad outreach and recruitment efforts and is disseminating its strate-
gies to serve other similar programmes.

Discussion
The AI-READI project is ambitious in scope and may be viewed as a
milestone in the field of big data and AI health research. We expect
that the flagship dataset will lead to novel discoveries into the T2DM
pathogenesis and salutogenesis while benefiting a broader population,
given the diversity of the study participants. In addition, our blueprint
has the potential to remove one of the major bottlenecks to the wide-
spread use of AI in healthcare: the availability of data that is off-the-shelf
ready for AI-based analysis.

AI-READI Consortium*
*A list of authors and their affiliations appears at the end of the paper.

Published online: 8 November 2024

References
1.  Wilkinson, M. D. et al. Sci. Data 3, 160018 (2016).
2.  Contreras, J. et al. License terms for reusing the AI-READI dataset. Zenodo https://doi.org/

10.5281/zenodo.10642459 (2024).

3.  Hackman, J. R. & Katz, N. Handb. Soc. Psychol. 2, 1208–1251 (2010).
4.  Lemieux-Charles, L. & McGuire, W. L. Med. Care Res. Rev. 63, 263–300 (2006).
5.  Abràmoff, M. D. et al. NPJ Digit. Med. 6, 170 (2023).
6.  Gebru, T. et al. Commun. ACM 64, 86–92 (2021).
7.  Rostamzadeh, N. et al. In Proceedings of the 2022 ACM Conference on Fairness,

Accountability, and Transparency 1943–1961 (Association for Computing Machinery, 2022).

8.  Patel, B., Soundarajan, S., Ménager, H. & Hu, Z. Sci. Data 10, 557 (2023).

Acknowledgements
This work was supported by the US National Institutes of Health (NIH) through grants
OT2OD032644 and P30 DK035816. We thank the Microsoft AI for Good Lab for supporting the

nature metabolism

Volume 6 | December 2024 | 2210–2212 | 2211

Comment
cloud services needed for the project. We thank Topcon Corporation (Tokyo, Japan),
Optomed (Oulu, Finland), iCare World (Raleigh, NC) and Carl Zeiss (Oberkochen, Germany)
for loaning their devices for research purposes at no cost. We thank Heidelberg Engineering
(Heidelberg, Germany), Dexcom (San Diego, CA) and Garmin (Olathe, KS) for research discounts
on study devices. We also thank the study participants and the AI-READI Advisory Council.

Author contributions
The Writing Committee members created the first draft, which was reviewed, edited and
approved by all the authors.

Competing interests
S.B.: funding — NIH, University of California Office of the President, Research to Prevent
Blindness; consultant — Topcon; equipment — Optomed. V.R.d.S.: funding — NSF, UCSD
Social Sciences, Sanford Institute for Empathy and Compassion (Center for Empathy and
Technology), Intel, Mathworks, UCSD instructional improvement grant, equipment funding
from Adobe and NVIDIA, Kavli Institute for Brain and Mind, IBM, past funding from Sony;
member — Cognitive Science Society Governing Board. K.F.: member — Institutional Review
Board for the All of Us Research Program, Digital Ethics Advisory Panel for Merck KGaA
(Merck Germany). C.S.L.: funding — NIH, Alzheimer’s Disease Drug Discovery Foundation,

Gates Ventures, Research to Prevent Blindness. T.Y.A.L.: funding — Research to Prevent
Blindness, Dr. H. James and Carole Free Career Development Award. B.P.: funding — NIH.
L.M.Z: funding — NEI, NIH, The Glaucoma Foundation, Heidelberg Engineering, DRCR Retina
Network/JAEB Center for Health Research, The Krupp Foundation; receipt of equipment,
materials, software — Optomed, ICare, Topcon, Heidelberg Engineering, Carl Zeiss Meditec,
Optovue/Visionix; consultant — Abbvie, Topcon Medical Systems; co-founder, inventor,
board member, equity holder — AISight Health Inc. S.H.: funding — NIH, NIH/NARCH, RWJF,
UCSD Herbert Wertheim School of Public Health. H.I.: funding — NIH; founder, stock holder
— Gobiquity, Inc. A.Y.L.: funding — Santen, Topcon, Carl Zeiss Meditec, Regeneron, Amazon,
Meta, Research to Prevent Blindness; personal fees — Genentech, Sanofi, US FDA, Johnson
and Johnson, Boehringer Ingelheim, Gyroscope; non-financial support — iCareWorld,
Optomed, Heidelberg, Microsoft. S.M.: Funding — NIH, Edward P. Evans Foundation, OHSU
Knight Cancer Institute; receipt of in-kind contribution — Nike. C.N.: funding — NIH, NSF,
PCORI. C.O.: consultant — Johnson and Johnson. L.H.: funding — NIH Grant UL1TR001442;
consultant — Bristol Myers Squibb. The remaining authors declare no competing interests.

Additional information
Supplementary information The online version contains supplementary material available at
https://doi.org/10.1038/s42255-024-01165-x.

AI-READI Consortium

Writing Committee

Sally L. Baxter
T. Y. Alvin Liu2, Julia P. Owen4, Bhavesh Patel

  1, Virginia R. de Sa

  5, Qilu Yu

  6 & Linda M. Zangwill1

  1, Kadija Ferryman2, Prachee Jain3, Cecilia S. Lee

  4, Jennifer Li-Pook-Than3,

Principal Investigators

Amir Bahmani
Hiroshi Ishikawa
Camille Nebeker

  3, Sally L. Baxter1, Christopher G. Chute

  2, Jeffrey C. Edberg7, Kadija Ferryman2, Samantha Hurst

  1,

  8, Cecilia S. Lee4, Aaron Y. Lee
  1, Cynthia Owsley7, Bhavesh Patel5, Sara J. Singer3 & Linda M. Zangwill1

, T. Y. Alvin Liu2, Gerald McGwin7, Shannon McWeeney

  4

  8,

Research, Technical and Clinical Staff

Riddhiman Adib8, Mohammad Adibuzzaman8, Arash Alavi
  8,
Marian Blazes4, Aaron Cohen8, Benjamin Cordier8, Katie Crist1, Colleen Cuddy3, Virginia R. de Sa1, Aydan Gasimova5,
Nayoon Gim
Jessica Mitchell2, Caitlyn Ngadisastra4, Victoria Patronilo

  2, Prachee Jain3, Trina Kim4, Jennifer Li-Pook-Than3, Wei-Chun Lin
  4, Sanjay Soundarajan
  1, Jamie Shaffer

  8,
  5 & Kevin Zhao

  4, Adrienne Baer3, Erik Benton

  3, Catherine Ashley

  4, Stephanie Hong

  4

Project Managers

Caroline Drolet

  4, Abigail Lucero

  8, Dawn Matthies7, Julia P. Owen4, Hanna Pittock

  3, Kate Watkins3 & Brittany York1

Interns

Charles E. Amankwa1, Monique Bangudi1, Nada Haboudal
  1, Apoorva Karsolia1, Hadi Khazaei
Fritz Gerald P. Kalaw

NIH Program Scientists

Xujing Wang11 & Qilu Yu6

  1, Shahin Hallaj
  8,9, Muna Mohammed

  1, Anna Heinke

  1, Lingling Huang

  1,

  10 & Kyongmi Simpkins1

1University of California San Diego, La Jolla, CA, USA. 2Johns Hopkins University, Baltimore, MD, USA. 3Stanford University, Stanford, CA, USA.
4University of Washington, Seattle, WA, USA. 5FAIR Data Innovations Hub, California Medical Innovations Institute, San Diego, CA, USA. 6National Center
for Complementary and Integrative Health, NIH, Bethesda, MD, USA. 7University of Alabama at Birmingham, Birmingham, AL, USA. 8Oregon Health &
Science University, Portland, OR, USA. 9Portland State University, Portland, OR, USA. 10Meharry Medical College, Nashville, TN, USA. 11National Institute of
Diabetes and Digestive and Kidney Diseases (NIDDK), NIH, Bethesda, MD, USA.

 e-mail: leeay@uw.edu

nature metabolism

Volume 6 | December 2024 | 2210–2212 | 2212

Comment


================================================================================

FILE: reporter_nih_gov_project-details-10471118_row7.txt
PATH: data/preprocessed/individual/AI_READI/reporter_nih_gov_project-details-10471118_row7.txt
SIZE: 4792 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: AI_READI
Source ID: nih_reporter_project
Source type: NIH project page
Source URL: https://reporter.nih.gov/project-details/10471118
Raw file: data/raw/AI_READI/reporter_nih_gov_project-details-10471118_row7.txt
--------------------------------------------------------------------------------
NIH RePORTER Project
Source: https://reporter.nih.gov/project-details/10471118
Application ID: 10471118
Project number: 1OT2OD032644-01
Core project number: OT2OD032644
Title: Bridge2AI: Salutogenesis Data Generation Project
Principal investigator: LEE, AARON
Organization: UNIVERSITY OF WASHINGTON
Fiscal year: 2022
Award amount: 5026499
Project start: 2022-09-01T00:00:00
Project end: 2025-08-31T00:00:00

Abstract Text
The Artificial Intelligence Ready and Exploratory Atlas for Diabetes Insights (AI-READI) project is one of the data generation projects in the NIH Common Fund’s Bridge2AI program. The project seeks to create a flagship ethically-sourced dataset to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. The ability to understand and affect the course of complex, multi-organ diseases such as T2DM has been limited by a lack of well-designed, high quality, and large multimodal datasets. The team of investigators will aim to collect a cross-sectional dataset of 4,000+ people and longitudinal data from 10% of the study cohort across the US. The study cohort will be balanced for diabetes disease stage. Data collection will be specifically designed to permit downstream pseudotime manifold analysis, an approach used to predict disease trajectories by collecting and learning from complex, multimodal data from participants with differing disease severity (normal to insulin-dependent T2DM). The long-term objective for this project is to develop a foundational dataset in diabetes, agnostic to existing classification criteria, which can be used to reconstruct a temporal atlas of T2DM development and reversal towards health (i.e., salutogenesis). Six cross-disciplinary project modules involving teams located across eight institutions will work together to develop this flagship dataset. All data will be optimized for downstream AI/ML research and made publicly available. . The AI-READI project will also engage in a tribal consultation to address barriers and facilitators of participation with the goal of collecting similar data within a Native American cohort in an ethical and respectful manner. Specific aims include 1) Collect and share the dataset for AI/ML research according to the Findable, Accessible, Interoperable, Reusable (FAIR) data principles, 2) Create a model for developing large scalable datasets, and 3) Increase access to and quality of AI/ML research by recruiting and training personnel.

Public Health Relevance Statement
Recent advances in artificial intelligence (AI) research are poised to provide breakthrough discoveries, but have been limited by the lack of large, well-characterized comprehensive datasets that capture molecular, physiological, pathological, and clinical at various stages of illness. To address these challenges, the AI-READI team of investigators will generate an ethically-sourced and unique dataset with many types of data collected from patients with different severities of type 2 diabetes mellitus (T2DM), which will enable key discoveries about the trajectory of this disease and how improvements to health (i.e., salutogenesis) can be promoted over time. The project will train future scientists in AI-based research and establish best practices for the generation of future datasets that are ethically sourced and accessible for responsible and scientifically valid use by the greater research community.

Preferred terms:
Address;Affect;Artificial Intelligence;Asian;Atlases;Awareness;Behavioral;Black race;Bridge to Artificial Intelligence;Classification;Clinical;Cohort Studies;Collaborations;Communities;Complex;Consultations;Data;Data Collection;Data Set;Development;Diabetes Mellitus;Disease;Ethics;FAIR principles;Foundations;Funding;Future;Future Generations;Generations;Goals;Health;Health Promotion;Hispanic;Human Resources;Institution;Insulin;Latinx;Learning;Machine Learning;Modeling;Molecular;Native Americans;Non-Insulin-Dependent Diabetes Mellitus;Participant;Pathologic;Pathway interactions;Patients;Persons;Physiological;Population Heterogeneity;Process;Research;Research Design;Research Personnel;Scientist;Severity of illness;Source;Text;Time;Training;United States National Institutes of Health;Work;base;cohort;data reuse;design;ethical legal social implication;insight;multimodal data;programs;recruit;social


================================================================================

FILE: docs_aireadi_org_docs-2_row10.txt
PATH: data/preprocessed/individual/AI_READI/docs_aireadi_org_docs-2_row10.txt
SIZE: 4734 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: AI_READI
Source ID: dataset_documentation
Source type: documentation
Source URL: https://docs.aireadi.org/docs/2/about
Raw file: data/raw/AI_READI/docs_aireadi_org_docs-2_row10.html
--------------------------------------------------------------------------------
About | Documentation for the AI-READI Dataset
Skip to main content
AI-READI Dataset Docs
Dataset v2.0.0
Dataset v3.0.0
Dataset v2.0.0
Dataset v1.0.0
GitHub
Contact Us
Search
About
Citation
Preliminary Information
Dataset Documentation
Controlled Variables
AI-readiness
Additional Resources
Changelog
Contact Us
This documentation is for v
2.0.0
of the dataset, which is no longer accessible.
Refer to the documentation for the latest version of the dataset v
3.0.0
.
About
Version: 2.0.0
On this page
About
About this documentation
â
This is the documentation for the AI-READI dataset called
Flagship Dataset of Type 2 Diabetes from the AI-READI Project
. It is intended to complement the information provided on the dataset landing page on the FAIRhub data portal. It is highly suggested to read this documentation first before accessing the dataset.
About the AI-READI dataset
â
The AI-READI is a dataset consisting of data collected from individuals with and without Type 2 Diabetes Mellitus (T2DM) and harmonized across 3 data collection sites. The composition of the dataset was designed with future studies using AI/Machine Learning in mind. This included recruitment sampling procedures aimed at achieving approximately equal distribution of participants across diabetes severity, as well as the design of a data acquisition protocol across multiple domains (survey data, physical measurements, clinical data, imaging data, wearable device data, etc.) to enable downstream AI/ML analyses that may not be feasible with existing data sources such as claims or electronic health records data.
The goal is to better understand salutogenesis (the pathway from disease to health) in T2DM. Some data that are not considered to be sensitive personal health data will be available to the public for download upon agreement with a license that defines how the data can be used. The full dataset will be accessible by entering into a data use agreement. The public dataset will include survey data, blood and urine lab results, fitness activity levels, clinical measurements (e.g. monofilament and cognitive function testing), retinal images, ECG, blood sugar levels, and environmental variables such as home air quality. The data held under controlled access include 5-digit zip code, sex, race, ethnicity, genetic sequencing data, past health records, medications, and traffic and accident reports. As enrollment is ongoing, the pilot data release and periodic updates to data releases may not have achieved balanced distribution across groups.
About the structure of this documentation
â
There is one version of this documentation associated with each version of the AI-READI dataset. You can navigate between the different versions from the dropdown in the upper right corner.
This documentation is associated with v2.0 of the dataset, which contains data from the participants of the pilot study phase. This version of the documentation is structured as follows
The
Citation
section explains how the dataset needs to be cited if used.
The
Preliminary Information
section contains information about accessing and using the dataset so you can understand what is expected from users of the dataset. This includes information about the license associated with the dataset and training prerequisites.
The
Dataset Documentation
section includes an overall description of the dataset using the Healthsheet template, and this section also provides details about each of the data domains included in the dataset. For each data domain, we have provided a general overview of the clinical context and how the data were acquired in the study, relevant individual variables, and data processing details including file formats, data standards, metadata, and example outputs.
The
AI-readiness
section describes additional considerations that were integrated into the study design, including consideration of FAIR principles and Ethical practices.
The
Additional resources
section provides a list of resources related to the dataset that could be useful for understanding and using the dataset.
Edit this page
Last updated
on
Jun 4, 2026
by
Eamon Dysinger
Next
Citation
About this documentation
About the AI-READI dataset
About the structure of this documentation
Docs
Changelog
Community
Homepage
More
GitHub
Copyright Â© 2026 AI-READI.
This repository is under review for potential modification in compliance with Administration directives.


================================================================================

FILE: docs_aireadi_org_docs-3_2026-07-24.txt
PATH: data/preprocessed/individual/AI_READI/docs_aireadi_org_docs-3_2026-07-24.txt
SIZE: 4638 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: AI_READI
Source ID: dataset_documentation_v3
Source type: documentation
Source URL: https://docs.aireadi.org/docs/3/about
Raw file: data/raw/AI_READI/docs_aireadi_org_docs-3_2026-07-24.html
--------------------------------------------------------------------------------
About | Documentation for the AI-READI Dataset
Skip to main content
AI-READI Dataset Docs
Dataset v3.0.0
Dataset v3.0.0
Dataset v2.0.0
Dataset v1.0.0
GitHub
Contact Us
Search
About
Citation
Preliminary Information
Dataset Documentation
Controlled Variables
AI-readiness
Additional Resources
Changelog
Contact Us
About
Version: 3.0.0
On this page
About
About this documentation
​
This is the documentation for the AI-READI dataset called
Flagship Dataset of Type 2 Diabetes from the AI-READI Project
. It is intended to complement the information provided on the dataset landing page on the FAIRhub data portal. It is highly suggested to read this documentation first before accessing the dataset.
About the AI-READI dataset
​
The AI-READI is a dataset consisting of data collected from individuals with and without Type 2 Diabetes Mellitus (T2DM) and harmonized across 3 data collection sites. The composition of the dataset was designed with future studies using AI/Machine Learning in mind. This included recruitment sampling procedures aimed at achieving approximately equal distribution of participants across diabetes severity, as well as the design of a data acquisition protocol across multiple domains (survey data, physical measurements, clinical data, imaging data, wearable device data, etc.) to enable downstream AI/ML analyses that may not be feasible with existing data sources such as claims or electronic health records data.
The goal is to better understand salutogenesis (the pathway from disease to health) in T2DM. Some data that are not considered to be sensitive personal health data will be available to the public for download upon agreement with a license that defines how the data can be used. The full dataset will be accessible by entering into a data use agreement. The public dataset will include survey data, blood and urine lab results, fitness activity levels, clinical measurements (e.g. monofilament and cognitive function testing), retinal images, ECG, blood sugar levels, and environmental variables such as home air quality. The data held under controlled access include 5-digit zip code, sex, race, ethnicity, genetic sequencing data, past health records, medications, and traffic and accident reports. As enrollment is ongoing, the pilot data release and periodic updates to data releases may not have achieved balanced distribution across groups.
About the structure of this documentation
​
There is one version of this documentation associated with each version of the AI-READI dataset. You can navigate between the different versions from the dropdown in the upper right corner.
This documentation is associated with v3.0 of the dataset, which contains data from the participants of the pilot study phase. This version of the documentation is structured as follows
The
Citation
section explains how the dataset needs to be cited if used.
The
Preliminary Information
section contains information about accessing and using the dataset so you can understand what is expected from users of the dataset. This includes information about the license associated with the dataset, training prerequisites, the mini-subset version of the dataset, and Azure Storage access.
The
Dataset Documentation
section includes an overall description of the dataset using the Healthsheet template, and this section also provides details about each of the data domains included in the dataset. For each data domain, we have provided a general overview of the clinical context and how the data were acquired in the study, relevant individual variables, and data processing details including file formats, data standards, metadata, and example outputs.
The
AI-readiness
section describes additional considerations that were integrated into the study design, including consideration of FAIR principles and Ethical practices.
The
Additional resources
section provides a list of resources related to the dataset that could be useful for understanding and using the dataset.
Edit this page
Last updated
on
Jun 4, 2026
by
Eamon Dysinger
Next
Citation
About this documentation
About the AI-READI dataset
About the structure of this documentation
Docs
Changelog
Community
Homepage
More
GitHub
Copyright © 2026 AI-READI.
This repository is under review for potential modification in compliance with Administration directives.


================================================================================

FILE: AI-READI-LICENSE-v2.0_2026-08-12.txt
PATH: data/preprocessed/individual/AI_READI/AI-READI-LICENSE-v2.0_2026-08-12.txt
SIZE: 14633 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: AI_READI
Source ID: dataset_license
Source type: license
Source URL: https://zenodo.org/records/17555036/files/AI-READI-LICENSE-v2.0.pdf?download=1
Raw file: data/raw/AI_READI/AI-READI-LICENSE-v2.0_2026-08-12.pdf
--------------------------------------------------------------------------------
WASHINGTON UNIVERSITY IN ST. LOUIS (“Licensor”)

AI-READI DATA LICENSE AGREEMENT (Version 2.0)

BY  INDICATING  ASSENT,  THE  LICENSEE  IDENTIFIED  IN THE DATA REQUEST WORKFLOW
(“LICENSEE”  OR  “YOU”),  AGREES  TO  THE  TERMS  AND  CONDITIONS  OF  THIS  DATA
LICENSE  AGREEMENT  WITH  LICENSOR  (“AGREEMENT”)  WITH  RESPECT  TO   THE
CONTENTS  OF  THE  ACCOMPANYING  DATA  FILES  (COLLECTIVELY,  THE  "DATA").  THE
IN  THE  DATA  REQUEST  WORKFLOW
INFORMATION  THAT  YOU  HAVE  PROVIDED
CONSTITUTES AN INTEGRAL PART OF THIS AGREEMENT.

YOU SHOULD SAVE OR PRINT A COPY OF THIS AGREEMENT FOR YOUR RECORDS.

IF  YOU  DO  NOT  AGREE  TO  ALL  OF  THE  TERMS  OF  THIS  AGREEMENT,  YOU  MUST  NOT
DOWNLOAD, INSTALL OR USE THE DATA.

1.

PARTIES; AUTHORIZED USERS.

A.

B.

C.

If, in Your Data Request Workflow, You indicated that you are entering into this Agreement
in  your  individual  capacity,  then  You  are  the  “Licensee”  and  no  other  person  will  be
authorized to access or use the Data under this Agreement. If you wish to share Data with
members  of  your  internal  group  or  team  or  other  employees  or  contractors  of  your
employer, please initiate a new Data Request Workflow and indicate this information when
requested,  upon which a new license agreement will be generated and provided for your
acceptance. References to “Authorized Group” and “Authorized Users” in this Agreement,
and the provisions of Paragraphs 1.B through 1.E below, do not apply to You.

If, in Your Data Request Workflow, You indicated that you are entering into this Agreement
on  behalf  of  an  internal  group,  lab,  or  business  unit  identified  in  the  Data  Request
Workflow  (“Authorized  Group”)  that  is  a  part  of the Institution/Employer specified in your
Data  Request  Workflow  (“Institution/Employer”),  then  this  Agreement  authorizes access,
downloading and use of the Data by You, as Licensee, as well as Authorized Users, on the
terms set forth below.

“Authorized Users” means individuals who are legal members of the Authorized Group via
contract, employment status or student status. The Authorized Group must be an officially
recognized  subunit  within  the  Institution/Employer  identified  in  the  Data  Request
Workflow,  as  evidenced  by  a  public  web  page  or  other  official  and  publicly  available
Institution/Employer information source. An individual’s status as an Authorized User, and
their  rights  under  this  Agreement,  terminate  automatically  upon  the  severance  of  their
relationship or employment with the Authorized Group or Institution/Employer.

D.  You, as Licensee, are permitted to sublicense your rights to Authorized Users for so long
as they are members of the Authorized Group. Authorized Users are entitled to exercise
all  rights  granted  to  you  as  Licensee  under  this  Agreement.  It  is  your  responsibility  to
ensure  that  each  Authorized  User  is  provided  with  a  copy  of  this  Agreement  and
understands and agrees to comply with the terms and conditions of this Agreement.

E.  You  must  ensure  that  each  Authorized  User  complies  fully  with  the  terms  of  this
Agreement  and  you  agree  that  you  will be fully liable for all acts and omissions of each
Authorized User. You represent and warrant to Licensor that you have all necessary legal
rights and authority to enter into this Agreement on behalf of all Authorized Users.

2.
LICENSE  GRANT.  Subject  to  Licensee’s  and  all  Authorized  Users’  compliance  with  the
terms  and  conditions  of  this  Agreement,  Licensor  grants  to  Licensee  a  non-exclusive  and
non-transferable license to download, reproduce and use the Data, and to create derivative works
















of the Data, for research,   commercial and non-commercial purposes. All full and partial copies of
the Data made by Licensee shall be subject to the terms of this Agreement.

3.

LIMITATIONS ON DATA SHARING; STORAGE; AND USAGE.

A.  Permitted  Sharing  with  Other  Licensees.  Licensee  shall  not  transfer,  license,
sublicense, sell, assign, display, share or otherwise convey any portion of the Data or any
derivative  work  to  any  third  party  other  than  another  licensee  (“Other Licensee”) that is
bound by the terms of an agreement with Licensor on terms identical to those contained in
this  Agreement, in which case Licensee shall be permitted to give access to the Data to
such  Other  Licensee  and  its  employees,  agents  and  contractors  that  are  bound  under
such  agreement  for  the  purpose  of  collaborating with Licensee on one or more projects
involving the Data.

B.  Permitted Data Storage. Licensee may use and store the data only on (i) servers
and  devices  maintained  by  and  located  within  Licensee’s  Institution/Employer,  or (ii) on
cloud  or  remote  storage  and  backup  services  (e.g.,  Dropbox,  Google  Drive,  AWS,
Microsoft  Azure) that have a HIPAA-approved Business Associate Agreement (“BAA”) in
place with Licensee’s Institution/Employer.

C.   Interaction  with  Third  Party  Models.  Licensee  shall not share or distribute Data with
any  third  party  model  vendor or developer for training or development purposes, even if
that  vendor  is  a  party  to  a  BAA  with  Licensee’s  Institution/Employer,  where  training
includes model weight modification and other adjustments to a model’s logic or operation.
Notwithstanding the foregoing, Licensee may use a third party model to analyze the
Data if the model vendor is a party to a BAA with Licensee’s Institution/Employer, where
the model’s interaction with the Data is limited to short-term interaction (e.g., prompting or
querying), but is not used for training purposes.

D.
Licensee  Models.  Licensee is permitted to make, reproduce and distribute models,
algorithms and programs that are developed, trained or adapted using the Data, but which
do  not  themselves  contain  the  Data  or  any  modified  version  of  the  Data  (“Licensee
Models”),  provided  that  Licensee,  prior  to  dissemination  of  any  such  Licensee  Models,
undertakes all reasonable efforts to minimise the likelihood that Data can be memorized,
derived,  reconstructed  or reconstituted through the use or construction of such Licensee
Models.

E.   Derivative  Data.  “Derivative  data”  is  Data  that  has  been  modified,  excerpted,
encrypted,  condensed,  encoded,  translated  or  otherwise  altered,  such  that  it  contains
Data  or  Data  may  be  derived  from  it.  “Synthetic  Data”  is  artificially  generated data that
mimics  real-world  data  characteristics.  Synthetic  Data  that  is  created  using  Data  or
Derivative  Data  is  also  considered  Derivative  Data.  For  purposes  of  this  Agreement,
Derivative Data is considered to be Data subject to all restrictions described herein.

F.   Publications.  Without  limiting  the  generality  of  the  foregoing,  Data  may  not  be
reproduced in papers, articles, presentations, analyses, reports or publications (“Papers”)
except that small representative samples of Data may be reproduced in up to five images
or  figures  per  Paper for illustrative purposes only. Notwithstanding journal or conference
requirements,  larger  amounts  of  data  shall  not  be published, posted or otherwise made
available  via  supplemental  files,  zip  archives,  code  packages  or other means. Licensee
may  refer  publishers  and  conference  organizers  to  Licensor  if  they  wish  to  obtain  a
separate license to the Data for such purposes.

4.
Licensee shall not:

ADDITIONAL  USE  RESTRICTIONS.  Without  limiting  the  generality  of  the  foregoing,

A.  Make  clinical  treatment  decisions  based  on  the  Data,  as  it is intended solely as a
research resource, or














Use  or  attempt  to  use  the  Data,  alone  or  in  concert  with  other  information,  (i)  to
B.
compromise  or  otherwise  infringe  the  confidentiality  of  information  about  an  individual
person who is the source of any Data or any clinical data or biological sample from which
Data  has been generated (a “Data Subject”), (ii) to invade or compromise the   privacy of
any Data Subject, (iii) to attempt to identify or contact any Data Subject or group of Data
Subjects,(iv)   to extract or extrapolate any identifying information about a Data Subject, to
establish  a  Data  Subject's  membership  in  a  particular  group  of persons, or otherwise to
cause harm or injury to any Data Subject.

ACKNOWLEDGEMENT.  Licensee  agrees  to  acknowledge  Licensor  and  the  source  and
5.
any funder of the Data in any Papers reporting use of the Data. The current citation can be found
here: docs.aireadi.org.

SECURITY.  Licensee  agrees  to  comply  with  all  data  security  and  privacy  standards
6.
established by the U.S. National Institutes of Health under its Genomic Data Sharing (GDS) Policy
from  time  to  time,  the  current  version  of  which  is  located  at  NIH  Security  Best  Practices  for
Controlled-Access  Data  Subject
(GDS)  Policy
(https://sharing.nih.gov/sites/default/files/flmngr/NIH_Best_Practices_for_Controlled-
Access_Data_Subject_to_the_NIH_GDS_Policy.pdf). Licensee acknowledges that the Data may be
statically  watermarked  to  identify  Licensee  for  security  purposes, and Licensee agrees that it will
take no action to remove, obscure, alter or mask such watermarking.

the  NIH  Genomic  Data  Sharing

to

TERMINATION. This Agreement will terminate automatically upon any breach of any term
7.
of this Agreement by Licensee or any Authorized User. Upon termination, Licensee shall delete all
copies  of  the  Data  in  its  possession  and  control,  including  in  the  possession  or  control  of  all
Authorized Users, and cease all use of the Data.

8.
PROPRIETARY RIGHTS. Title to the Data, and all industrial and intellectual property rights
therein, shall at all times remain solely and exclusively with Licensor and its suppliers, and Licensee
shall not take any action inconsistent with such ownership. Any rights not expressly granted herein
are reserved to Licensor and its suppliers.

9.
DISCLAIMER  OF  WARRANTY.  THE  DATA  IS  PROVIDED  ON  AN  "AS  IS"  BASIS,
WITHOUT  WARRANTY  OF ANY KIND, INCLUDING WITHOUT LIMITATION THE WARRANTIES
THAT IT IS FREE FROM DEFECTS, MERCHANTABLE, FIT FOR A PARTICULAR PURPOSE OR
NON-INFRINGING.  THE  ENTIRE  RISK  AS  TO  THE  QUALITY  AND  PERFORMANCE  OF  THE
DATA  IS  BORNE  BY  LICENSEE.  SHOULD  THE  DATA PROVE DEFECTIVE IN ANY RESPECT,
LICENSEE  AND NOT LICENSOR OR ITS SUPPLIERS ASSUMES THE ENTIRE COST OF ANY
SERVICE  AND  REPAIR.  THIS  DISCLAIMER  OF  WARRANTY  CONSTITUTES  AN  ESSENTIAL
PART OF THIS AGREEMENT. NO USE OF THE DATA IS AUTHORIZED HEREUNDER EXCEPT
UNDER THIS DISCLAIMER.

10.
LIMITATIONS OF LIABILITY. TO THE MAXIMUM EXTENT PERMITTED BY APPLICABLE
LAW,  IN  NO  EVENT WILL LICENSOR OR ITS SUPPLIERS BE LIABLE TO LICENSEE OR ANY
AUTHORIZED  USER  OR  OTHER  PARTY  CLAIMING  THROUGH  LICENSEE  FOR  ANY
PUNITIVE,  EXEMPLARY, MULTIPLE, INDIRECT, SPECIAL, INCIDENTAL OR CONSEQUENTIAL
DAMAGES  ARISING  OUT  OF  THE  USE  OF  OR  INABILITY  TO  USE  THE  DATA,  INCLUDING,
WITHOUT  LIMITATION,  DAMAGES  FOR  LOSS  OF  GOODWILL,  WORK  STOPPAGE,
COMPUTER  FAILURE  OR  MALFUNCTION,  OR  ANY  AND  ALL  OTHER  COMMERCIAL
DAMAGES  OR  LOSSES,  EVEN
IF  ADVISED  OF  THE  POSSIBILITY  THEREOF,  AND
REGARDLESS OF THE LEGAL OR EQUITABLE THEORY (CONTRACT, TORT OR OTHERWISE)
UPON WHICH THE CLAIM IS BASED.

IN  ANY  CASE,  LICENSOR'S  ENTIRE  LIABILITY  UNDER  ANY  PROVISION  OF  THIS
AGREEMENT AND WITH RESPECT TO THE DATA SHALL NOT EXCEED IN THE AGGREGATE
ONE  U.S.  DOLLAR,  WITH  THE  EXCEPTION OF DEATH OR PERSONAL INJURY CAUSED BY
THE  NEGLIGENCE  OF  LICENSOR  TO  THE  EXTENT  APPLICABLE  LAW  PROHIBITS  THE
LIMITATION  OF  DAMAGES  IN  SUCH  CASES.  SOME  JURISDICTIONS  DO  NOT  ALLOW  THE
EXCLUSION  OR  LIMITATION  OF  INCIDENTAL  OR  CONSEQUENTIAL  DAMAGES,  SO  THIS









EXCLUSION AND LIMITATION MAY NOT BE APPLICABLE.

11.
INDEMNIFICATION. To the extent allowed by applicable law, Licensee agrees to indemnify,
defend  and  hold  harmless  Licensor  and  its  suppliers  and  their  respective  employees,  officers,
directors,  contractors  and  agents  from  and  against  any  and  all  claims,  damages,  losses,
settlements,  penalties,  costs,  expenses  and  other  amounts  arising  directly  or  indirectly  from
Licensee’s or any Authorized Users use of the Data and any use, distribution or activity of a Model,
including,  without  limitation,  all  third  party  claims  asserting  violation  of  privacy  rights,  death,
personal  harm  or  injury,  economic  loss,  emotional  distress,  discrimination,  defamation, breach of
security,  national  security,  or  infringement  of  patent,  copyright  or  other  intellectual  or  industrial
property rights.

12.
restrictions relating to the distribution and use of the Data and Models.

COMPLIANCE.  Licensee  agrees  to  comply  with  all  applicable  laws,  regulations  and

13.
GENERAL.  (a)  This  Agreement  constitutes  the  entire  agreement  between  the  parties
concerning  the  subject  matter  hereof.  (b)  Subject  to  the Licensor’s right to update and modify its
security  policies  as  provided  in  Paragraph  6,  this  Agreement  may  be  amended only by a writing
signed by both parties. (c) If any provision in this Agreement should be held illegal or unenforceable
by a court having jurisdiction, such provision shall be modified to the extent necessary to render it
enforceable  without  losing  its  intent,  or  severed  from  this  Agreement  if  no  such  modification  is
possible,  and  other  provisions  of  this  Agreement  shall  remain  in  full  force  and  effect.  (d)  The
language of this Agreement is English. (e) A waiver by either party of any term or condition of this
Agreement  or  any  breach  thereof,  in  any  one  instance, shall not waive such term or condition or
any  subsequent  breach  thereof.  (f)  This  Agreement  shall  be  binding  upon  and  shall  inure  to the
benefit of the parties, their successors and permitted assigns.


================================================================================

FILE: fairhub_dataset_2_row12.txt
PATH: data/preprocessed/individual/AI_READI/fairhub_dataset_2_row12.txt
SIZE: 1631 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: AI_READI
Source ID: fairhub_dataset
Source type: data resource
Source URL: https://fairhub.io/datasets/2
Raw file: data/raw/AI_READI/fairhub_dataset_2_row12.html
--------------------------------------------------------------------------------
Flagship Dataset of Type 2 Diabetes from the AI-READI Project
FAIRhub
Open main menu
Find Datasets
Share Datasets
About
Contact
My Requests
Flagship Dataset of Type 2 Diabetes from the AI-READI Project
AI-READI Consortium
The AI-READI project seeks to create and share a flagship ethically-sourced dataset of type 2 diabetes
This version of the dataset is no longer accessible. Please refer to the
latest version
.
About
Healthsheet
Study Dashboard
Study Metadata
Dataset Metadata
Dataset Structure Preview
Dataset Quality Dashboard
Dataset Uses
Usage statistics
Views
Cited by
Access approved
All versions
Current version
More info on how stats are collected....
2.01 TB
165,051
Files
License
Health Data License
Keywords
Diabetes mellitus
Machine Learning
Artificial Intelligence
Electrocardiography
Continuous Glucose Monitoring
Retinal imaging
Eye exam
Citation
When using this resource, please cite:
When using this resource, please follow the citation instructions provided at
https://docs.aireadi.org/docs/2/citation
Versions
Version 3.0.0
10.60775/fairhub.3
Nov 17, 2025
Version 2.0.0
10.60775/fairhub.2
Nov 8, 2024
Version 1.0.0
10.60775/fairhub.1
May 3, 2024
Dataset Impact
Dataset Index
FAIR score
Citations
Mentions
Platform is currently in beta
This repository is under review for potential modification in compliance with Administration directives.


================================================================================

FILE: fairhub_dataset_3_2026-07-24.txt
PATH: data/preprocessed/individual/AI_READI/fairhub_dataset_3_2026-07-24.txt
SIZE: 1571 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: AI_READI
Source ID: fairhub_dataset_v3
Source type: data resource
Source URL: https://fairhub.io/datasets/3
Raw file: data/raw/AI_READI/fairhub_dataset_3_2026-07-24.html
--------------------------------------------------------------------------------
Flagship Dataset of Type 2 Diabetes from the AI-READI Project
FAIRhub
Open main menu
Find Datasets
Share Datasets
About
Contact
My Requests
Flagship Dataset of Type 2 Diabetes from the AI-READI Project
AI-READI Consortium
The AI-READI project seeks to create and share a flagship ethically-sourced dataset of type 2 diabetes.
Access this dataset
View the dataset documentation
About
Healthsheet
Study Dashboard
Study Metadata
Dataset Metadata
Dataset Structure Preview
Dataset Quality Dashboard
Dataset Uses
Usage statistics
Views
Cited by
Access approved
All versions
Current version
More info on how stats are collected....
3.82 TB
356,343
Files
A smaller version is available for pipeline development...
License
Health Data License
Keywords
Diabetes mellitus
Machine Learning
Artificial Intelligence
Electrocardiography
Continuous Glucose Monitoring
Retinal imaging
Eye exam
Citation
When using this resource, please cite:
When using this resource, please follow the citation instructions provided at
https://docs.aireadi.org/docs/3/citation
Versions
Version 3.0.0
10.60775/fairhub.3
Nov 17, 2025
Version 2.0.0
10.60775/fairhub.2
Nov 8, 2024
Version 1.0.0
10.60775/fairhub.1
May 3, 2024
This repository is under review for potential modification in compliance with Administration directives.


================================================================================

FILE: fairhub_api_dataset_3_2026-07-27.txt
PATH: data/preprocessed/individual/AI_READI/fairhub_api_dataset_3_2026-07-27.txt
SIZE: 167476 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: AI_READI
Source ID: fairhub_dataset_v3_api
Source type: structured metadata
Source URL: https://fairhub.io/api/datasets/3
Raw file: data/raw/AI_READI/fairhub_api_dataset_3_2026-07-27.json
--------------------------------------------------------------------------------
{
  "id": "3",
  "title": "Flagship Dataset of Type 2 Diabetes from the AI-READI Project",
  "created_at": 1763366400,
  "data": {
    "size": 3815969779678,
    "fileCount": 356343,
    "viewCount": 24636,
    "cited": 0,
    "mini": false,
    "parent": null,
    "child": 4
  },
  "dataset_id": "d894862f-0795-4ba6-b40b-fae14eb77813",
  "description": "The AI-READI project seeks to create and share a flagship ethically-sourced dataset of type 2 diabetes.",
  "doi": "10.60775/fairhub.3",
  "files": [],
  "metadata": {
    "datasetDescription": {
      "schema": "https://schema.aireadi.org/v0.1.0/dataset_description.json",
      "identifier": {
        "identifierValue": "10.60775/fairhub.3",
        "identifierType": "DOI"
      },
      "title": [
        {
          "titleValue": "Flagship Dataset of Type 2 Diabetes from the AI-READI Project"
        },
        {
          "titleValue": "AI-READI dataset",
          "titleType": "AlternativeTitle"
        }
      ],
      "version": "3.0.0",
      "creator": [
        {
          "creatorName": "AI-READI Consortium",
          "nameType": "Organizational"
        }
      ],
      "publicationYear": "2025",
      "date": [
        {
          "dateValue": "2025-11-17",
          "dateType": "Available",
          "dateInformation": "Date dataset made available on FAIRhub"
        },
        {
          "dateValue": "2023-07-19/2025-05-01",
          "dateType": "Collected",
          "dateInformation": "Period when the data was collected"
        }
      ],
      "resourceType": {
        "resourceTypeValue": "Type 2 Diabetes",
        "resourceTypeGeneral": "Dataset"
      },
      "datasetDeIdentLevel": {
        "deIdentType": "NoDeIdentification",
        "deIdentDirect": true,
        "deIdentHIPAA": true,
        "deIdentDates": false,
        "deIdentNonarr": false,
        "deIdentKAnon": false,
        "deIdentDetails": "No identifiers were collected so no active de-identification was necessary but we checked that no identifiable data per US HIPAA were present in the data."
      },
      "datasetConsent": {
        "consentType": "ConsentSpecifiedNotElsewhereCategorised",
        "consentNoncommercial": false,
        "consentGeogRestrict": false,
        "consentResearchType": false,
        "consentGeneticOnly": false,
        "consentNoMethods": false,
        "consentsDetails": "The public version of the dataset can only be used for type 2 diabetes related research. A private version will allow for more generic use."
      },
      "description": [
        {
          "descriptionValue": "This dataset contains data from 2280 participants that was collected between from July 19, 2023 and May 01, 2025. Data from multiple modalities are included. The data in this dataset contain no protected health information (PHI). Information related to the sex and race/ethnicity of the participants as well as medication used has also been removed. A detailed description of the dataset is available in the AI-READI documentation for v3.0.0 of the dataset at https://docs.aireadi.org",
          "descriptionType": "Abstract"
        }
      ],
      "language": "en",
      "relatedIdentifier": [
        {
          "relatedIdentifierValue": "https://docs.aireadi.org/",
          "relatedIdentifierType": "URL",
          "relationType": "IsDocumentedBy",
          "resourceTypeGeneral": "Other"
        },
        {
          "relatedIdentifierValue": "https://aireadi.org/",
          "relatedIdentifierType": "URL",
          "relationType": "IsDocumentedBy",
          "resourceTypeGeneral": "Other"
        }
      ],
      "subject": [
        {
          "subjectValue": "Diabetes mellitus",
          "subjectIdentifier": {
            "classificationCode": "45636-8",
            "subjectScheme": "Logical Observation Identifier Names and Codes (LOINC)",
            "schemeURI": "https://loinc.org/",
            "valueURI": "https://loinc.org/45636-8"
          }
        },
        {
          "subjectValue": "Machine Learning",
          "subjectIdentifier": {
            "classificationCode": "D000069550",
            "subjectScheme": "Medical Subject Headings (MeSH)",
            "schemeURI": "https://meshb.nlm.nih.gov/",
            "valueURI": "https://meshb.nlm.nih.gov/record/ui?ui=D000069550"
          }
        },
        {
          "subjectValue": "Artificial Intelligence",
          "subjectIdentifier": {
            "classificationCode": "D001185",
            "subjectScheme": "Medical Subject Headings (MeSH)",
            "schemeURI": "https://meshb.nlm.nih.gov/",
            "valueURI": "https://meshb.nlm.nih.gov/record/ui?ui=D001185"
          }
        },
        {
          "subjectValue": "Electrocardiography",
          "subjectIdentifier": {
            "classificationCode": "D004562",
            "subjectScheme": "Medical Subject Headings (MeSH)",
            "schemeURI": "https://meshb.nlm.nih.gov/",
            "valueURI": "https://meshb.nlm.nih.gov/record/ui?ui=D004562"
          }
        },
        {
          "subjectValue": "Continuous Glucose Monitoring",
          "subjectIdentifier": {
            "classificationCode": "D000095583",
            "subjectScheme": "Medical Subject Headings (MeSH)",
            "schemeURI": "https://meshb.nlm.nih.gov/",
            "valueURI": "https://meshb.nlm.nih.gov/record/ui?ui=D000095583"
          }
        },
        {
          "subjectValue": "Retinal imaging"
        },
        {
          "subjectValue": "Eye exam"
        }
      ],
      "managingOrganization": {
        "name": "Washington University in St. Louis",
        "managingOrganizationIdentifier": {
          "managingOrganizationIdentifierValue": "https://ror.org/01yc7t268",
          "managingOrganizationScheme": "ROR",
          "schemeURI": "https://ror.org"
        }
      },
      "accessType": "PublicDownloadSelfAttestationRequired",
      "accessDetails": {
        "description": "Accessing the dataset requires several steps, including: Login in through a verified ID system, Agreeing to use the data only for type 2 diabetes related research, Agreeing to the license terms which set certain restrictions and obligations for data usage (see 'rights' property)"
      },
      "rights": [
        {
          "rightsName": "AI-READI custom license v2.0",
          "rightsURI": "https://doi.org/10.5281/zenodo.17555036"
        }
      ],
      "publisher": {
        "publisherName": "FAIRhub"
      },
      "size": [
        "3.82 TB",
        "356343 files"
      ],
      "fundingReference": [
        {
          "funderName": "National Institutes of Health",
          "funderIdentifier": {
            "funderIdentifierValue": "https://ror.org/01cwqze88",
            "funderIdentifierType": "ROR",
            "schemeURI": "https://ror.org"
          },
          "awardNumber": {
            "awardNumberValue": "OT2OD032644",
            "awardURI": "https://reporter.nih.gov/search/yatARMM-qUyKAhnQgsCTAQ/project-details/10885481"
          },
          "awardTitle": "Bridge2AI: Salutogenesis Data Generation Project"
        }
      ],
      "format": [
        "application/dicom",
        "text/markdown",
        "text/csv",
        "application/json"
      ]
    },
    "datasetStructureDescription": {
      "schema": "https://schema.aireadi.org/v0.1.1/dataset_structure_description.json",
      "directoryList": [
        {
          "directoryName": "cardiac_ecg",
          "directoryType": "dataType",
          "directoryDescription": "This directory contains electrocardiogram data collected by a 12 lead protocol (the current standard), Holter monitor, or smartwatch. The terms ECG and EKG are often used interchangeably.",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://docs.aireadi.org",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other",
              "relatedIdentifierDescription": "This is the documentation of the AI-READI dataset that contains additional information about the data contained in this directory."
            }
          ],
          "relatedTerm": [
            {
              "relatedTermValue": "Electrocardiogram",
              "relatedTermIdentifier": [
                {
                  "relatedTermClassificationCode": "C168186",
                  "relatedTermScheme": "NCI Thesaurus (NCIT)",
                  "relatedTermSchemeURI": "https://ncim.nci.nih.gov/",
                  "relatedTermValueURI": "https://ncit.nci.nih.gov/ncitbrowser/pages/concept_details.jsf?dictionary=NCI%20Thesaurus&code=C168186"
                }
              ]
            },
            {
              "relatedTermValue": "Electrocardiography",
              "relatedTermIdentifier": [
                {
                  "relatedTermClassificationCode": "D004562",
                  "relatedTermScheme": "Medical Subject Headings (MeSH)",
                  "relatedTermSchemeURI": "https://meshb.nlm.nih.gov/",
                  "relatedTermValueURI": "https://meshb.nlm.nih.gov/record/ui?ui=D004562"
                }
              ]
            }
          ],
          "relatedStandard": [
            {
              "standardName": "Clinical Data Structure (CDS) v0.1.1",
              "standardDescription": "Standard for consistently structuring and describing clinical research datasets",
              "standardUse": "This directory and its sub-directories are named and organized following the specification from this standard.",
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            },
            {
              "standardName": "WaveForm DataBase (WFDB)",
              "standardDescription": "Set of file standards designed for reading and storing physiologic signal data, and associated annotations.",
              "standardUse": "All the data files within this directory follow the format specified in this standard.",
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://wfdb.readthedocs.io/en/latest/wfdb.html",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            }
          ],
          "directoryList": [
            {
              "directoryName": "ecg_12lead",
              "directoryType": "modality",
              "directoryDescription": "This directory contains ECG data collected using the 12 lead protocol",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://en.wikipedia.org/wiki/Electrocardiography",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy",
                  "resourceTypeGeneral": "Other",
                  "relatedIdentifierDescription": "This page describes 3 types of ECG protocol including the 12 lead (standard) protocol."
                }
              ],
              "directoryList": [
                {
                  "directoryName": "philips_tc30",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains ECG data collected using the 12 lead protocol using the Philips PageWriter TC30 device.",
                  "relatedIdentifier": [
                    {
                      "relatedIdentifierValue": "https://www.documents.philips.com/doclib/enc/fetch/2000/4504/577242/577243/577246/581601/711562/DXL_ECG_Algorithm_Physician_s_Guide_(ENG)_Ed.2.pdf",
                      "relatedIdentifierType": "URL",
                      "relationType": "IsDocumentedBy",
                      "resourceTypeGeneral": "Other",
                      "relatedIdentifierDescription": "The 'Philips DXL ECG Algorithm Physician’s Guide' contains information on the fields that are printed on an ECG report."
                    }
                  ]
                }
              ]
            }
          ],
          "metadataFileList": [
            {
              "metadataFileName": "manifest.tsv",
              "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS) v0.1.1",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDocumentedBy",
                  "resourceTypeGeneral": "Other"
                }
              ]
            }
          ],
          "size": 302931703,
          "numberOfFiles": 4515
        },
        {
          "directoryName": "clinical_data",
          "directoryType": "dataType",
          "directoryDescription": "This directory contains clinical data collected through REDCap, including blood/urine lab values and survey data. Each CSV file in this directory is a one-to-one mapping to the OMOP CDM tables.",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://docs.aireadi.org",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other",
              "relatedIdentifierDescription": "This is the documentation of the AI-READI dataset that contains additional information about data contained in this directory."
            }
          ],
          "relatedTerm": [
            {
              "relatedTermValue": "Clinical Data",
              "relatedTermIdentifier": [
                {
                  "relatedTermClassificationCode": "C15783",
                  "relatedTermScheme": "NCI Thesaurus (NCIT)",
                  "relatedTermSchemeURI": "https://ncim.nci.nih.gov/",
                  "relatedTermValueURI": "https://ncit.nci.nih.gov/ncitbrowser/pages/concept_details.jsf?dictionary=NCI%20Thesaurus&code=C15783"
                }
              ]
            }
          ],
          "relatedStandard": [
            {
              "standardName": "Clinical Data Structure (CDS) v0.1.1",
              "standardDescription": "Standard for consistently structuring and describing clinical research datasets",
              "standardUse": "This directory is named following the specification of this standard.",
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            },
            {
              "standardName": "The Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM)",
              "standardDescription": "Standard designed to standardize the structure and content of observational data and to enable efficient analyses that can produce reliable evidence.",
              "standardUse": "All the data files within this directory follow this standard.",
              "standardIdentifier": [
                {
                  "identifierValue": "https://doi.org/10.25504/FAIRsharing.qk984b",
                  "identifierType": "DOI"
                }
              ],
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://ohdsi.github.io/TheBookOfOhdsi/CommonDataModel.html",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            }
          ],
          "metadataFileList": [
            {
              "metadataFileName": "dqd_omop.json",
              "metadataFileDescription": "The dqd_omop.json file supports OMOP CDM data quality analysis using the OMOP CDM Data Quality Dashboard (DQD). The OMOP CDM DQD tool (https://ohdsi.github.io/DataQualityDashboard/) runs a set of > 3500 data quality checks against an OMOP CDM instance",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://ohdsi.github.io/DataQualityDashboard/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDocumentedBy",
                  "resourceTypeGeneral": "Other"
                }
              ]
            }
          ],
          "size": 176182781,
          "numberOfFiles": 7
        },
        {
          "directoryName": "environment",
          "directoryType": "dataType",
          "directoryDescription": "This directory contains data collected through an environmental sensor device custom built for the AI-READI project.",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://docs.aireadi.org",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other",
              "relatedIdentifierDescription": "This is the documentation of the AI-READI dataset that contains additional information about the data contained in this directory."
            }
          ],
          "relatedTerm": [
            {
              "relatedTermValue": "Environmental sensor data"
            }
          ],
          "relatedStandard": [
            {
              "standardName": "Clinical Data Structure (CDS) v0.1.1",
              "standardDescription": "Standard for consistently structuring and describing clinical research datasets",
              "standardUse": "This directory and its sub-directories are named and organized following the specification from this standard.",
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            },
            {
              "standardName": "ASCII File Format Guidelines for Earth Science Data",
              "standardDescription": "NASA recommended practices for formatting and describing ASCII encoded data files",
              "standardUse": "All the data files within this directory follow the format specified in this standard.",
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://www.earthdata.nasa.gov/esdis/esco/standards-and-practices/ascii-file-format-guidelines-for-earth-science-data",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            }
          ],
          "directoryList": [
            {
              "directoryName": "environmental_sensor",
              "directoryType": "modality",
              "directoryList": [
                {
                  "directoryName": "leelab_anura",
                  "directoryDescription": "You can learn more about the data in this directory by consulting relevant documentation. Search 'AS7431' with Type set as 'Datasheet' at https://ams-osram.com/support/download-center. Search 'DS3231 Precision RTC' at https://www.adafruit.com, look for product ID 5188, select this link and scroll down to Technical Details to find a link for the Datasheet (as of this writing, you may find the datasheet at https://www.analog.com/media/en/technical-documentation/data-sheets/DS3231.pdf). Search 'Datasheet SEN5x' in the search box at https://sensirion.com/ to get the document titled 'Datasheet SEN5x' (the name of the downloaded file may be 'Sensirion_Datasheet_Environmental_Node_SEN5x.pdf'). Search 'NOx Index' in the search box at https://sensirion.com to get the document titled 'What is Sensirion's NOx Index?' (the name of the downloaded file may be 'Info_Note NOx_Index.pdf').",
                  "directoryType": "device"
                }
              ]
            }
          ],
          "metadataFileList": [
            {
              "metadataFileName": "manifest.tsv",
              "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDocumentedBy",
                  "resourceTypeGeneral": "Other"
                }
              ]
            }
          ],
          "size": 55625676514,
          "numberOfFiles": 2232
        },
        {
          "directoryName": "retinal_flio",
          "directoryType": "dataType",
          "directoryDescription": "This directory contains data collected through fluorescence lifetime imaging ophthalmoscopy (FLIO), an imaging modality for in vivo measurement of lifetimes of endogenous retinal fluorophores.",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://docs.aireadi.org",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other",
              "relatedIdentifierDescription": "This is the documentation of the AI-READI dataset that contains additional information about the data contained in this directory."
            }
          ],
          "relatedTerm": [
            {
              "relatedTermValue": "Fuorescence Lifetime Imaging Ophthalmoscopy"
            }
          ],
          "relatedStandard": [
            {
              "standardName": "Clinical Data Structure (CDS) v0.1.1",
              "standardDescription": "Standard for consistently structuring and describing clinical research datasets",
              "standardUse": "This directory and its sub-directories are named and organized following the specification from this standard.",
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            },
            {
              "standardName": "Digital Imaging and Communications in Medicine (DICOM)",
              "standardDescription": "Standard for the digital storage and transmission of medical images and related information.",
              "standardUse": "All the data files within this directory follow the format specified in this standard.",
              "standardIdentifier": [
                {
                  "identifierValue": "https://doi.org/10.25504/FAIRsharing.b7z8by",
                  "identifierType": "DOI"
                }
              ],
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "http://medical.nema.org/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            }
          ],
          "directoryList": [
            {
              "directoryName": "flio",
              "directoryType": "modality",
              "directoryList": [
                {
                  "directoryName": "heidelberg_flio",
                  "directoryType": "device"
                }
              ]
            }
          ],
          "metadataFileList": [
            {
              "metadataFileName": "manifest.tsv",
              "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDocumentedBy",
                  "resourceTypeGeneral": "Other"
                }
              ]
            }
          ],
          "size": 1069466876718,
          "numberOfFiles": 7969
        },
        {
          "directoryName": "retinal_oct",
          "directoryType": "dataType",
          "directoryDescription": "This directory contains data collected using optical coherence tomography (OCT), an imaging method using lasers that is used for mapping subsurface structure.",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://docs.aireadi.org",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other",
              "relatedIdentifierDescription": "This is the documentation of the AI-READI dataset that contains additional information about the data contained in this directory."
            }
          ],
          "relatedTerm": [
            {
              "relatedTermValue": "Optical Coherence Tomography",
              "relatedTermIdentifier": [
                {
                  "relatedTermClassificationCode": "C20828",
                  "relatedTermScheme": "NCI Thesaurus (NCIT)",
                  "relatedTermSchemeURI": "https://ncim.nci.nih.gov/",
                  "relatedTermValueURI": "https://ncit.nci.nih.gov/ncitbrowser/pages/concept_details.jsf?dictionary=NCI%20Thesaurus&code=C20828"
                }
              ]
            },
            {
              "relatedTermValue": "Tomography, Optical Coherence",
              "relatedTermIdentifier": [
                {
                  "relatedTermClassificationCode": "D041623",
                  "relatedTermScheme": "Medical Subject Headings (MeSH)",
                  "relatedTermSchemeURI": "https://meshb.nlm.nih.gov/",
                  "relatedTermValueURI": "https://meshb.nlm.nih.gov/record/ui?ui=D041623"
                }
              ]
            }
          ],
          "relatedStandard": [
            {
              "standardName": "Clinical Data Structure (CDS) v0.1.1",
              "standardDescription": "Standard for consistently structuring and describing clinical research datasets",
              "standardUse": "This directory and its sub-directories are named and organized following the specification from this standard.",
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            },
            {
              "standardName": "Digital Imaging and Communications in Medicine (DICOM)",
              "standardDescription": "Standard for the digital storage and transmission of medical images and related information.",
              "standardUse": "All the data files within this directory follow the format specified in this standard.",
              "standardIdentifier": [
                {
                  "identifierValue": "https://doi.org/10.25504/FAIRsharing.b7z8by",
                  "identifierType": "DOI"
                }
              ],
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "http://medical.nema.org/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            }
          ],
          "directoryList": [
            {
              "directoryName": "structural_oct",
              "directoryType": "modality",
              "directoryList": [
                {
                  "directoryName": "heidelberg_spectralis",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains OCT data collected from the Spectralis device, manufactured by Heidelberg."
                },
                {
                  "directoryName": "topcon_maestro2",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains OCT data collected from the Maestro2 device, manufactured by Topcon."
                },
                {
                  "directoryName": "topcon_triton",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains OCT data collected from the Triton device, manufactured by Topcon."
                },
                {
                  "directoryName": "zeiss_cirrus",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains OCT data collected from the Cirrus device, manufactured by Zeiss."
                }
              ]
            }
          ],
          "metadataFileList": [
            {
              "metadataFileName": "manifest.tsv",
              "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDocumentedBy",
                  "resourceTypeGeneral": "Other"
                }
              ]
            }
          ],
          "size": 1317625293027,
          "numberOfFiles": 56478
        },
        {
          "directoryName": "retinal_octa",
          "directoryType": "dataType",
          "directoryDescription": "This directory contains data collected using optical coherence tomography angiography (OCTA), a non-invasive imaging technique that generates volumetric angiography images.",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://docs.aireadi.org",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other",
              "relatedIdentifierDescription": "This is the documentation of the AI-READI dataset that contains additional information about the data contained in this directory."
            }
          ],
          "relatedTerm": [
            {
              "relatedTermValue": "Optical Coherence Tomography Angiography"
            },
            {
              "relatedTermValue": "Optical Coherence Tomography Angiography Images"
            },
            {
              "relatedTermValue": "OCTA Images"
            }
          ],
          "relatedStandard": [
            {
              "standardName": "Clinical Data Structure (CDS) v0.1.1",
              "standardDescription": "Standard for consistently structuring and describing clinical research datasets",
              "standardUse": "This directory and its sub-directories are named and organized following the specification from this standard.",
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            },
            {
              "standardName": "Digital Imaging and Communications in Medicine (DICOM)",
              "standardDescription": "Standard for the digital storage and transmission of medical images and related information.",
              "standardUse": "All the data files within this directory follow the format specified in this standard.",
              "standardIdentifier": [
                {
                  "identifierValue": "https://doi.org/10.25504/FAIRsharing.b7z8by",
                  "identifierType": "DOI"
                }
              ],
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "http://medical.nema.org/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            }
          ],
          "directoryList": [
            {
              "directoryName": "enface",
              "directoryType": "modality",
              "directoryDescription": "This directory contains en face data collected using OCTA. En face images, derived from 3D volume scans, are also referred to as C-scan OCT. These images provide a 2D view of the retina layers and have an orientation siimilar to fundus photographs.",
              "directoryList": [
                {
                  "directoryName": "heidelberg_spectralis",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains IR images, specifically from the Spectralis device, manufactured by Heidelberg."
                },
                {
                  "directoryName": "topcon_maestro2",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains OCTA en face data collected from the Maestro2 device, manufactured by Topcon."
                },
                {
                  "directoryName": "topcon_triton",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains OCTA en face data collected from the Triton device, manufactured by Topcon."
                },
                {
                  "directoryName": "zeiss_cirrus",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains OCTA en face data collected from the Cirrus device, manufactured by Zeiss."
                }
              ]
            },
            {
              "directoryName": "flow_cube",
              "directoryType": "modality",
              "directoryDescription": "This directory contains flow cube data collected using OCTA. Flow cube provides information on blood flow in a 3D view.",
              "directoryList": [
                {
                  "directoryName": "heidelberg_spectralis",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains IR images, specifically from the Spectralis device, manufactured by Heidelberg."
                },
                {
                  "directoryName": "topcon_maestro2",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains flow cube data collected from the Maestro2 device, manufacture by Topcon."
                },
                {
                  "directoryName": "topcon_triton",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains flow cube data collected from the Triton device, manufacture by Topcon."
                },
                {
                  "directoryName": "zeiss_cirrus",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains flow cube en face data collected from the Cirrus device, manufactured by Zeiss."
                }
              ]
            },
            {
              "directoryName": "segmentation",
              "directoryType": "modality",
              "directoryDescription": "This directory contains segmentation data collected using OCTA. The segmentation information is presented in the form of heightmaps and includes information about the associated layers.",
              "directoryList": [
                {
                  "directoryName": "heidelberg_spectralis",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains IR images, specifically from the Spectralis device, manufactured by Heidelberg."
                },
                {
                  "directoryName": "topcon_maestro2",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains segmentation data collected from the Maestro device, manufactured by Topcon."
                },
                {
                  "directoryName": "topcon_triton",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains segmentation data collected from the Triton device, manufactured by Topcon."
                },
                {
                  "directoryName": "zeiss_cirrus",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains segmentation data collected from the Cirrus device, manufactured by Zeiss."
                }
              ]
            }
          ],
          "metadataFileList": [
            {
              "metadataFileName": "manifest.tsv",
              "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDocumentedBy",
                  "resourceTypeGeneral": "Other"
                }
              ]
            }
          ],
          "size": 1155908809724,
          "numberOfFiles": 173721
        },
        {
          "directoryName": "retinal_photography",
          "directoryType": "dataType",
          "directoryDescription": "This directory contains retinal photography data, which are 2D images. They are also referred to as fundus photography.",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://docs.aireadi.org",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other",
              "relatedIdentifierDescription": "This is the documentation of the AI-READI dataset that contains additional information about the data contained in this directory."
            },
            {
              "relatedIdentifierValue": "https://en.wikipedia.org/wiki/Fundus_photography",
              "relatedIdentifierType": "URL",
              "relationType": "IsDescribedBy",
              "resourceTypeGeneral": "Other"
            }
          ],
          "relatedTerm": [
            {
              "relatedTermValue": "Eye Fundus Photography",
              "relatedTermIdentifier": [
                {
                  "relatedTermClassificationCode": "C147467",
                  "relatedTermScheme": "NCI Thesaurus (NCIT)",
                  "relatedTermSchemeURI": "https://ncim.nci.nih.gov/",
                  "relatedTermValueURI": "https://ncit.nci.nih.gov/ncitbrowser/pages/concept_details.jsf?dictionary=NCI%20Thesaurus&code=C147467"
                }
              ]
            }
          ],
          "relatedStandard": [
            {
              "standardName": "Clinical Data Structure (CDS) v0.1.1",
              "standardDescription": "Standard for consistently structuring and describing clinical research datasets",
              "standardUse": "This directory and its sub-directories are named and organized following the specification from this standard.",
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            },
            {
              "standardName": "Digital Imaging and Communications in Medicine (DICOM)",
              "standardDescription": "Standard for the digital storage and transmission of medical images and related information.",
              "standardUse": "All the data files within this directory follow the format specified in this standard.",
              "standardIdentifier": [
                {
                  "identifierValue": "https://doi.org/10.25504/FAIRsharing.b7z8by",
                  "identifierType": "DOI"
                }
              ],
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "http://medical.nema.org/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            }
          ],
          "directoryList": [
            {
              "directoryName": "cfp",
              "directoryType": "modality",
              "directoryDescription": "This directory contains retinal photography data, specifically color fundus photographs.",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://ophthalmology.med.ubc.ca/patient-care/ophthalmic-photography/color-fundus-photography/#:~:text=Color%20Fundus%20Retinal%20Photography%20uses,monitor%20their%20change%20over%20time",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy",
                  "resourceTypeGeneral": "Other"
                }
              ],
              "directoryList": [
                {
                  "directoryName": "icare_eidon",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains retinal photography data, specifically color fundus photographs from the Eidon device, manufactured by iCare."
                },
                {
                  "directoryName": "optomed_aurora",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains retinal photography data, specifically color fundus photographs from the Aurora device, manufactured by Optomed."
                },
                {
                  "directoryName": "topcon_maestro2",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains retinal photography data, specifically color fundus photographs from the Maestro2 device, manufactured by Topcon."
                },
                {
                  "directoryName": "topcon_triton",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains retinal photography data, specifically color fundus photographs from the Triton device, manufactured by Topcon."
                }
              ]
            },
            {
              "directoryName": "faf",
              "directoryType": "modality",
              "directoryDescription": "This directory contains retinal photography data, specificallly fundus autofluorescence photographs that uses the fluorescent characteristics of lipofuscin in an non-invasive way.",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://eyewiki.aao.org/Fundus_Autofluorescence",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy",
                  "resourceTypeGeneral": "Other"
                }
              ],
              "directoryList": [
                {
                  "directoryName": "icare_eidon",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains fundus autofluorescence photographs from the Eidon device, manufactured by iCare."
                }
              ]
            },
            {
              "directoryName": "ir",
              "directoryType": "modality",
              "directoryDescription": "This directory contains retinal photography data using near-infrared reflectance (IR).",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8349282/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy",
                  "resourceTypeGeneral": "Other"
                }
              ],
              "directoryList": [
                {
                  "directoryName": "heidelberg_spectralis",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains IR images, specifically from the Spectralis device, manufactured by Heidelberg."
                },
                {
                  "directoryName": "icare_eidon",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains IR images, specifically from the Eidon device, manufactured by iCare."
                },
                {
                  "directoryName": "topcon_maestro2",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains IR images, specifically from the Maestro2 device, manufactured by Topcon."
                },
                {
                  "directoryName": "zeiss_cirrus",
                  "directoryType": "device",
                  "directoryDescription": "This directory contains IR images, specifically from the Cirrus device, manufactured by Zeiss."
                }
              ]
            }
          ],
          "metadataFileList": [
            {
              "metadataFileName": "manifest.tsv",
              "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDocumentedBy",
                  "resourceTypeGeneral": "Other"
                }
              ]
            }
          ],
          "size": 174381046406,
          "numberOfFiles": 93921
        },
        {
          "directoryName": "wearable_activity_monitor",
          "directoryType": "dataType",
          "directoryDescription": "This directory contains data collected through a wearable fitness tracker.",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://docs.aireadi.org",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other",
              "relatedIdentifierDescription": "This is the documentation of the AI-READI dataset that contains additional information about the data contained in this directory."
            }
          ],
          "relatedTerm": [
            {
              "relatedTermValue": "smartwatch"
            },
            {
              "relatedTermValue": "activity monitoring"
            }
          ],
          "relatedStandard": [
            {
              "standardName": "Clinical Data Structure (CDS) v0.1.1",
              "standardDescription": "Standard for consistently structuring and describing clinical research datasets",
              "standardUse": "This directory and its sub-directories are named and organized following the specification from this standard.",
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            },
            {
              "standardName": "Open mHealth",
              "standardDescription": "Open Standard for Mobile Health Data (Open mHealth) is the leading mobile health data interoperability standard.",
              "standardUse": "All the data files within this directory follow the format specified in this standard",
              "standardIdentifier": [
                {
                  "identifierValue": "https://doi.org/10.25504/FAIRsharing.mrpMBj",
                  "identifierType": "DOI"
                }
              ],
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://www.openmhealth.org/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            }
          ],
          "directoryList": [
            {
              "directoryName": "heart_rate",
              "directoryType": "modality",
              "directoryList": [
                {
                  "directoryName": "garmin_vivosmart5",
                  "directoryType": "device"
                }
              ]
            },
            {
              "directoryName": "oxygen_saturation",
              "directoryType": "modality",
              "directoryList": [
                {
                  "directoryName": "garmin_vivosmart5",
                  "directoryType": "device"
                }
              ]
            },
            {
              "directoryName": "physical_activity",
              "directoryType": "modality",
              "directoryList": [
                {
                  "directoryName": "garmin_vivosmart5",
                  "directoryType": "device"
                }
              ]
            },
            {
              "directoryName": "physical_activity_calorie",
              "directoryType": "modality",
              "directoryList": [
                {
                  "directoryName": "garmin_vivosmart5",
                  "directoryType": "device"
                }
              ]
            },
            {
              "directoryName": "respiratory_rate",
              "directoryType": "modality",
              "directoryList": [
                {
                  "directoryName": "garmin_vivosmart5",
                  "directoryType": "device"
                }
              ]
            },
            {
              "directoryName": "sleep",
              "directoryType": "modality",
              "directoryList": [
                {
                  "directoryName": "garmin_vivosmart5",
                  "directoryType": "device"
                }
              ]
            },
            {
              "directoryName": "stress",
              "directoryType": "modality",
              "directoryList": [
                {
                  "directoryName": "garmin_vivosmart5",
                  "directoryType": "device"
                }
              ]
            }
          ],
          "metadataFileList": [
            {
              "metadataFileName": "manifest.tsv",
              "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDocumentedBy",
                  "resourceTypeGeneral": "Other"
                }
              ]
            }
          ],
          "size": 38313536220,
          "numberOfFiles": 15245
        },
        {
          "directoryName": "wearable_blood_glucose",
          "directoryType": "dataType",
          "directoryDescription": "This directory contains data collected through a continuous glucose monitoring (CGM) device.",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://docs.aireadi.org",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other",
              "relatedIdentifierDescription": "This is the documentation of the AI-READI dataset that contains additional information about the data contained in this directory."
            }
          ],
          "relatedTerm": [
            {
              "relatedTermValue": "Continuous Glucose Monitoring System",
              "relatedTermIdentifier": [
                {
                  "relatedTermClassificationCode": "C159776",
                  "relatedTermScheme": "NCI Thesaurus (NCIT)",
                  "relatedTermSchemeURI": "https://ncim.nci.nih.gov/",
                  "relatedTermValueURI": "https://ncit.nci.nih.gov/ncitbrowser/pages/concept_details.jsf?dictionary=NCI%20Thesaurus&code=C159776"
                }
              ]
            }
          ],
          "relatedStandard": [
            {
              "standardName": "Clinical Data Structure (CDS) v0.1.1",
              "standardDescription": "Standard for consistently structuring and describing clinical research datasets",
              "standardUse": "This directory and its sub-directories are named and organized following the specification from this standard.",
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            },
            {
              "standardName": "Open mHealth",
              "standardDescription": "Open Standard for Mobile Health Data (Open mHealth) is the leading mobile health data interoperability standard.",
              "standardUse": "All the data files within this directory follow the format specified in this standard",
              "standardIdentifier": [
                {
                  "identifierValue": "https://doi.org/10.25504/FAIRsharing.mrpMBj",
                  "identifierType": "DOI"
                }
              ],
              "standardRelatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://www.openmhealth.org/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDescribedBy"
                }
              ]
            }
          ],
          "directoryList": [
            {
              "directoryName": "continuous_glucose_monitoring",
              "directoryType": "modality",
              "directoryList": [
                {
                  "directoryName": "dexcom_g6",
                  "directoryType": "device"
                }
              ]
            }
          ],
          "metadataFileList": [
            {
              "metadataFileName": "manifest.tsv",
              "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
              "relatedIdentifier": [
                {
                  "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
                  "relatedIdentifierType": "URL",
                  "relationType": "IsDocumentedBy",
                  "resourceTypeGeneral": "Other"
                }
              ]
            }
          ],
          "size": 4169006971,
          "numberOfFiles": 2246
        }
      ],
      "metadataFileList": [
        {
          "metadataFileName": "CHANGELOG.md",
          "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other"
            }
          ]
        },
        {
          "metadataFileName": "dataset_description.json",
          "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other"
            }
          ]
        },
        {
          "metadataFileName": "dataset_structure_description.json",
          "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other"
            }
          ]
        },
        {
          "metadataFileName": "healthsheet.md",
          "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other"
            }
          ]
        },
        {
          "metadataFileName": "LICENSE.txt",
          "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other"
            }
          ]
        },
        {
          "metadataFileName": "participants.json",
          "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other"
            }
          ]
        },
        {
          "metadataFileName": "participants.tsv",
          "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other"
            }
          ]
        },
        {
          "metadataFileName": "README.md",
          "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other"
            }
          ]
        },
        {
          "metadataFileName": "study_description.json",
          "metadataFileDescription": "This is a metadata file based on the Clinical Dataset Structure (CDS)",
          "relatedIdentifier": [
            {
              "relatedIdentifierValue": "https://cds-specification.readthedocs.io/en/v0.1.1/",
              "relatedIdentifierType": "URL",
              "relationType": "IsDocumentedBy",
              "resourceTypeGeneral": "Other"
            }
          ]
        }
      ]
    },
    "healthsheet": {
      "general_information": [
        {
          "id": 1,
          "question": "Provide a 2 sentence summary of this dataset.",
          "response": "The Artificial Intelligence Ready and Exploratory Atlas for Diabetes Insights (AI-READI) is a dataset consisting of data collected from individuals with and without “Type 2 Diabetes Mellitus (T2DM)” and harmonized across 3 data collection sites. The composition of the dataset was designed with future studies using AI/Machine Learning in mind. This included recruitment sampling procedures aimed at achieving approximately equal distribution of participants across sex, race, and diabetes severity, as well as the design of a data acquisition protocol across multiple domains (survey data, physical measurements, clinical data, imaging data, wearable device data, etc.) to enable downstream AI/ML analyses that may not be feasible with existing data sources such as claims or electronic health records data. The goal is to better understand salutogenesis (the pathway from disease to health) in T2DM. Some data that are not considered to be sensitive personal health data will be available to the public for download upon agreement with a license that defines how the data can be used. The full dataset will be accessible by entering into a data use agreement. The public dataset will include survey data, blood and urine lab results, fitness activity levels, clinical measurements (e.g. monofilament and cognitive function testing), retinal images, ECG, blood sugar levels, and home air quality. The data held under controlled access include 5-digit zip code, sex, race, ethnicity, genetic sequencing data, past health records, and traffic and accident reports. Of note, the overall enrollment goal is to have balanced distribution between different racial groups. As enrollment is ongoing, the periodic updates to data releases may not have achieved balanced distribution across groups."
        },
        {
          "id": 2,
          "question": "Has the dataset been audited before? If yes, by whom and what are the results?",
          "response": "The dataset has not undergone any formal external audits. However, the dataset has been reviewed internally by AI-READI team members for quality checks and to ensure that no personally identifiable information was accidentally included."
        }
      ],
      "versioning": [
        {
          "id": 1,
          "question": "Does the dataset get released as static versions or is it dynamically updated?\n\n   a. If static, how many versions of the dataset exist?\n\n   b. If dynamic, how frequently is the dataset updated?",
          "response": "The dataset gets released as static versions. This is the third version of the dataset and consists of data collected up through the end of the second year of the study, i.e. between July 19, 2023 and May 1st, 2025. There are plans to release new versions of the dataset approximately once a year with additional data from participants who have been enrolled since the last dataset version release."
        },
        {
          "id": 2,
          "question": "Is this datasheet created for the original version of the dataset? If not, which version of the dataset is this datasheet for?",
          "response": "This datasheet is created for the third version of the dataset."
        },
        {
          "id": 3,
          "question": "Are there any datasheets created for any versions of this dataset?",
          "response": "There was a previous datasheet created for the first version of the dataset, which consisted of data collected during the pilot data collection phase. It is available here: https://docs.aireadi.org/docs/1/dataset/healthsheet. \n\nThere was also a previous datasheet created for the second version of the dataset, which consisted of data collected up to the end of the first full year of the study. It is available here: https://docs.aireadi.org/docs/2/dataset/healthsheet"
        },
        {
          "id": 4,
          "question": "Does the current version/subversion of the dataset come with predefined task(s), labels, and recommended data splits (e.g., for training, development/validation, testing)? If yes, please provide a high-level description of the introduced tasks, data splits, and labeling, and explain the rationale behind them. Please provide the related links and references. If not, is there any resource (website, portal, etc.) to keep track of all defined tasks and/or associated label definitions? (please note that more detailed questions w.r.t labeling is provided in further sections)",
          "response": "See response to question #6 under “Labeling and subjectivity of labeling”."
        },
        {
          "id": 5,
          "question": "If the dataset has multiple versions, and this datasheet represents one of them, answer the following questions:\n\n   a. What are the characteristics that have been changed between different versions of the dataset? ",
          "response": "This version of the dataset includes more patients than the first version of the dataset."
        },
        {
          "id": 6,
          "question": "b. Explain the motivation/rationale for creating the current version of the dataset.",
          "response": "The current version of the dataset includes data from the second year of the study, rather than only the data collected from the pilot data collection phase (which comprised the first version of the dataset) and the first year of the study (which comprised the second version of the dataset)."
        },
        {
          "id": 7,
          "question": "c. Does this version have more subjects/patients represented in the data, or fewer?",
          "response": "This version has more subjects/patients represented in the data (the first version of the dataset contains data from 204 participants, the second version contains additional data from 863 participants for a total of 1067 participants, and this version contains a total of 2280 participants)."
        },
        {
          "id": 8,
          "question": "d. Does this version of the dataset have extended data or new data from the same patients as the older versions? Were any patients, data fields, or data points removed? If so, why?",
          "response": "The data fields/types are largely the same as the prior version of the dataset. However, the Snellen visual acuity variables were dropped in this version of the dataset, such that now only logMAR visual acuity measurements are available."
        },
        {
          "id": 9,
          "question": "e. Do we expect more versions of the dataset to be released?",
          "response": "Yes, enrollment is ongoing, and future versions of the dataset will be released that will include larger numbers of subjects/patients as enrollment increases."
        },
        {
          "id": 10,
          "question": "f. Is this datasheet for a version of the dataset? If yes, does this sub-version of the dataset introduce a new task, labeling, and/or recommended data splits? If the answer to any of these questions is yes, explain the rationale behind it.",
          "response": "This datasheet does not include new tasks or labeling. It does include a new recommended data split for the 2280 participants balancing training, validation, and testing sets for age, sex, races/ethnicities, and study group."
        },
        {
          "id": 11,
          "question": "g. Are you aware of any widespread version(s)/subversion(s) of the dataset? If yes, what is the addressed task, or application that is addressed?",
          "response": "No"
        }
      ],
      "motivation": [
        {
          "id": 1,
          "question": "For what purpose was the dataset created? Was there a specific task in mind? Was there a specific gap that needed to be filled? Please provide a description.",
          "response": "The purpose for creating the dataset was to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. T2DM is a growing public health threat. Yet, the current understanding of T2DM, especially in the context of salutogenesis, is limited. Given the complexity of T2DM, AI-based approaches may help with improving our understanding but a key issue is the lack of data ready for training AI models. The AI-READI dataset is intended to fill this gap."
        },
        {
          "id": 2,
          "question": "What are the applications that the dataset is meant to address? (e.g., administrative applications, software applications, research)",
          "response": "The multi-modal dataset being collected is being gathered to facilitate downstream pseudotime manifolds and various applications in artificial intelligence."
        },
        {
          "id": 3,
          "question": "Are there any types of usage or applications that are discouraged from using this dataset? If so, why?",
          "response": "The AI READI dataset License imposes certain restrictions on the usage of the data. The restrictions are described in the License files available at https://doi.org/10.5281/zenodo.17555036."
        },
        {
          "id": 4,
          "question": "Who created this dataset (e.g., which team, research group), and on behalf of which entity (e.g., company, institution, organization)?",
          "response": "This dataset was created by members of the AI-READI project, hereby referred to as the AI-READI Consortium. Details about each member and their institutions are available on the project website at https://aireadi.org."
        },
        {
          "id": 5,
          "question": "Who funded the creation of the dataset? If there is an associated grant, please provide the name of the grantor and the grant name and number. If the funding institution differs from the research organization creating and managing the dataset, please state how.",
          "response": "The creation of the dataset was funded by the National Institutes of Health (NIH) through their Bridge2AI Program (https://commonfund.nih.gov/bridge2ai). The grant number is OT2ODO32644 and more information about the funding is available at https://reporter.nih.gov/search/T-mv2dbzIEqp9V6UJjHpgw/project-details/10885481. Note that the funding institution is not creating or managing the dataset. The dataset is created and managed by the awardees of the grant (c.f. answer to the previous question)."
        },
        {
          "id": 6,
          "question": "What is the distribution of backgrounds and experience/expertise of the dataset curators/generators?",
          "response": "There is a wide range of experience within the project team, including senior, mid-career, and early career faculty members as well as clinical research coordinators, staff, and interns. They collectively cover many areas of expertise including clinical research, data collection, data management, data standards, bioinformatics, team science, and ethics, among others. Visit https://aireadi.org/team for more information."
        }
      ],
      "composition": [
        {
          "id": 1,
          "question": "What do the instances that comprise the dataset represent (e.g., documents, images, people, countries)? Are there multiple types of instances? Please provide a description.",
          "response": "Each instance represents an individual patient."
        },
        {
          "id": 2,
          "question": "How many instances are there in total (of each type, if appropriate) (breakdown based on schema, provide data stats)?",
          "response": "There are 2280 instances in this current version of the dataset (version 3, released fall 2025)."
        },
        {
          "id": 3,
          "question": "How many patients / subjects does this dataset represent? Answer this for both the preliminary dataset and the current version of the dataset.",
          "response": "This version of the dataset has data from 2280 participants. The first version of the dataset, composed of data from the pilot data collection phase, had 204 instances. The second version of the dataset, composed of data from the first year of the study, had 1067 instances."
        },
        {
          "id": 4,
          "question": "Does the dataset contain all possible instances or is it a sample (not necessarily random) of instances from a larger set? If the dataset is a sample, then what is the larger set? Is the sample representative of the larger set (e.g., geographic coverage)? If so, please describe how this representativeness was validated/verified. If it is not representative of the larger set, please describe why not (e.g., to cover a more diverse range of instances, because instances were withheld or unavailable). Answer this question for the preliminary version and the current version of the dataset in question.",
          "response": "The dataset contains all possible instances. More specifically, the dataset contains data from all participants who have been enrolled during the first year of data collection for AI-READI."
        },
        {
          "id": 5,
          "question": "What data modality does each patient data consist of? If the data is hierarchical, provide the modality details for all levels (e.g: text, image, physiological signal). Break down in all levels and specify the modalities and devices.",
          "response": "Multiple modalities of data are collected for each participant, including survey data, clinical data, retinal imaging data, environmental sensor data, continuous glucose monitor data, and wearable activity monitor data. These encompass tabular data, imaging data, and physiological signal/waveform data. There is no unstructured text data included in this dataset. The exact forms used for data collection in REDCap are available [here](https://docs.aireadi.org/v3/REDCap%20surveys%20and%20forms.pdf). Furthermore, all modalities, file formats, and devices are detailed in the dataset documentation at https://docs.aireadi.org/."
        },
        {
          "id": 6,
          "question": "What data does each instance consist of? “Raw” data (e.g., unprocessed text or images) or features? In either case, please provide a description.",
          "response": "Each instance consists of all of the data available for an individual participating in the study. See answer to question 5 for the data types associated with each instance."
        },
        {
          "id": 7,
          "question": "Is any information missing from individual instances? If so, please provide a description, explaining why this information is missing (e.g., because it was unavailable).?",
          "response": "Yes, not all modalities are available for all participants. Some participants elected not to participate in some study elements. In a few cases, the data collection device did not have any stored results or was returned too late to retrieve the results (e.g. battery died, data was lost). In a few cases, there may have been a data collision at some point in the process and data has been lost."
        },
        {
          "id": 8,
          "question": "Are relationships between individual instances made explicit? (e.g., They are all part of the same clinical trial, or a patient has multiple hospital visits and each visit is one instance)? If so, please describe how these relationships are made explicit.",
          "response": "Yes - all instances are part of the same prospective data generation project (AI-READI). There is currently only one visit per participant."
        },
        {
          "id": 9,
          "question": "Are there any errors, sources of noise, or redundancies in the dataset? If so, please provide a description. (e.g., losing data due to battery failure, or in survey data subjects skip the question, radiological sources of noise).",
          "response": "In cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected."
        },
        {
          "id": 10,
          "question": "Is the dataset self-contained, or does it link to or otherwise rely on external resources (e.g., websites, other datasets)? If it links to or relies on external resources,\n\n a. are there guarantees that they will exist, and remain constant, over time;\n\n b. are there official archival versions of the complete dataset (i.e., including the external resources as they existed at the time the dataset was created);\n\n c. are there any restrictions (e.g., licenses, fees) associated with any of the external resources that might apply to a future user? Please provide descriptions of all external resources and any restrictions associated with them, as well as links or other access points, as appropriate.",
          "response": "The dataset is self-contained but does rely on the dataset documentation for users requiring additional information about the provenance of the dataset. The documentation is available at https://docs.aireadi.org. The documentation is shared under the CC-BY 4.0 license, so there are no restrictions associated with its use."
        },
        {
          "id": 11,
          "question": "Does the dataset contain data that might be considered confidential (e.g., data that is protected by legal privilege or by doctor-patient confidentiality, data that includes the content of individuals' non-public communications that is confidential)? If so, please provide a description.",
          "response": "No, the dataset does not contain data that might be considered confidential. No personally identifiable information is included in the dataset."
        },
        {
          "id": 12,
          "question": "Does the dataset contain data that, if viewed directly, might be offensive, insulting, threatening, or might otherwise pose any safety risk (such as psychological safety and anxiety)? If so, please describe why.",
          "response": "No."
        },
        {
          "id": 13,
          "question": "If the dataset has been de-identified, were any measures taken to avoid the re-identification of individuals? Examples of such measures: removing patients with rare pathologies or shifting time stamps.",
          "response": ""
        },
        {
          "id": 14,
          "question": "Does the dataset contain data that might be considered sensitive in any way (e.g., data that reveals racial or ethnic origins, sexual orientations, religious beliefs, political opinions or union memberships, or locations; financial or health data; biometric or genetic data; forms of government identification, such as social security numbers; criminal history)? If so, please provide a description.",
          "response": "No, the public dataset will not contain data that is considered sensitive. However, the controlled access dataset will contain data regarding racial and ethnic origins, location (5-digit zip code), as well as motor vehicle accident reports."
        }
      ],
      "devices": [
        {
          "id": 1,
          "question": "For data that requires a device or equipment for collection or the context of the experiment, answer the following additional questions or provide relevant information based on the device or context that is used (for example)\n\n   a. If there was an MRI machine used, what is the MRI machine and model used?\n\n   b. If heart rate was measured what is the device for heart rate variation that is used?\n\n   c. If cortisol measurement is reported at multi site, provide details,\n\n   d. If smartphones were used to collect the data, provide the names of models.\n\n   e. And so on,..",
          "response": "The devices included in the study are as follows, and more details can be found at [https://docs.aireadi.org](https://docs.aireadi.org):\n\n   **Environmental sensor device**\n\n   Participants will be sent home with an environmental sensor (a custom-designed sensor unit called the LeeLab Anura), which they will use for 10 continuous days before returning the devices to the clinical research coordinators for data download.\n\n   **Continuous glucose monitor (Dexcom G6)**\n\n   The Dexcom G6 is a real-time, integrated continuous glucose monitoring system (iCGM) that directly monitors blood glucose levels without requiring finger sticks. It must be worn continuously in order to collect data.\n\n   **Wearable accelerometer (Physical activity monitor)**\n\n   The Garmin Vivosmart 5 Fitness Activity tracker will be used to measure data related to physical activity.\n\n   **Heart rate**\n\n   Heart rate can be read from EKG or blood pressure measurement devices.\n\n   **Blood pressure**\n\n   Blood pressure devices used for the study across the various data acquisition sites are: OMRON HEM 907XL Blood Pressure Monitor, Medline MDS4001 Automatic Digital Blood Pressure Monitor, and Welch Allyn 6000 series Vital signs monitor with Welch Allyn FlexiPort Reusable Blood Pressure Cuff.\n\n   **Visual acuity**\n\n   M&S Technologies EVA device to test visual acuity. The test is administered at a distance of 4 meters from a touch-screen monitor that is 12x20 inches. Participants will read letters from the screen. Photopic Conditions: No neutral density filters are used. A general occluder will be used for photopic testing. The participant wears their own prescription spectacles or trial frames. For Mesopic conditions, a neutral density (ND) filter will be used. The ND filter will either be a lens added to trial frames to reduce incoming light on the tested eye, OR a handheld occluder with a neutral density\n   filter (which we will designate as “ND-occluder) over the glasses will be used. The ND-occluder is different from a standard occluder and is used only for vision testing under mesopic conditions.\n\n   **Contrast sensitivity**\n\n   The MARS Letter Contrast Sensitivity test (Perceptrix) was conducted monocularly under both Photopic conditions (with a general occluder) and Mesopic conditions (using a Neutral Density occluder with a low luminance filter lens). The standardized order of MARS cards was as follows: Photopic OD, Photopic OS, Mesopic OD, and Mesopic OS. The background luminance of the charts fell within the range of 60 to 120 cd/m2, with an optimal level of 85 cd/m2. Illuminance was recommended to be between 189 to 377 lux, with an optimal level of 267 lux. While the designed viewing distance was 50 cm, it could vary between 40 to 59 cm. Patients were required to wear their appropriate near correction: reading glasses or trial frames with +2.00D lenses. All testing was carried out under undilated conditions. Patients were instructed to read the letter left to right across each line on the chart. Patients were encouraged to guess, even if they perceived the letters as too faint. Testing was terminated either when the patient made two consecutive errors or reached the end of the chart. The log contrast sensitivity (log CS) values were recorded by multiplying the number of errors prior to the final correct letter by 0.04 and subtracting the result from the log CS value at the final correct letter. If a patient reached the end of the chart without making two consecutive errors, the final correct letter was simply the last one correctly identified.\n\n   **Autorefraction**\n\n   KR 800 Auto Keratometer/Refractor.\n\n   **EKG**\n\n   Philips (manufacturer of Pagewriter TC30 Cardiograph)\n\n   **Lensometer**\n\n   Lensometer devices used at data acquisition sites across the study include: NIDEK LM-600P Auto Lensometer, Topcon-CL-200 computerized Lensometer, and Topcon-CL-300 computerized Lensometer\n\n   **Undilated fundus photography - Optomed Aurora**\n\n   The Optomed Aurora IQ is a handheld fundus camera that can take non-mydriatic images of the ocular fundus. It has a 50° field of view, 5 Mpix sensor, and high-contrast optical design. The camera is non-mydriatic, meaning it doesn't require the pupil to be dilated, so it can be used for detailed viewing of the retina. Images taken during the AI-READI visit, are undilated images taken in a dark room while a patient is sitting on a comfortable chair, laying back. As it becomes challenging to get a good view because of the patients not being dilated and the handheld nature of this imaging modality, the quality of the images vary from patient to patient and within the same patient.\n\n   **Dilated fundus photography - Eidon**\n\n   The iCare EIDON is a widefield TrueColor confocal fundus imaging system that can capture images up to 200°. It comes with multiple imaging modalities, including TrueColor, blue, red, Red-Free, and infrared confocal images. The system offers widefield, ultra-high-resolution imaging and the capability to image through cataract and media opacities. It operates without dilation (minimum pupil 2.5 mm) and provides the flexibility of both fully automated and fully manual modes. Additionally, the iCare EIDON features an all-in-one compact design, eliminating the need for an additional PC. AI READI images using EIDON include two main modalities: 1. Single Field Central IR/FAF 2. Smart Horizontal Mosaic. Imaging is done in fully automated mode in a dark room with the machine moving and positioning according to the patient's head aiming at optimizing the view and minimizing operator's involvement/operator induced noise.\n\n   **Spectralis HRA (Heidelberg Engineering)**\n\n   The Heidelberg Spectralis HRA+OCT is an ophthalmic imaging system that combines optical coherence tomography (OCT) with retinal angiography. It is a modular, upgradable platform that allows clinicians to configure it for their specific diagnostic workflow. It has the confocal scanning laser ophthalmoscope (cSLO) technology that not only offers documentation of clinical findings but also often highlights critical diagnostic details that are not visible on traditional clinical ophthalmoscopy. Since cSLO imaging minimizes the effects of light scatter, it can be used effectively even in patients with cataracts. For AI READI subjects, imaging is done in a dark room using the following modalities: ONH-RC, PPole-H, and OCTA of the macula. As the machine is operated by the imaging team and is not fully automated, quality issues may arise, which may lead to skipping this modality and missing data.\n\n   **Triton DRI OCT (Topcon Healthcare)**\n\n   The DRI OCT Triton is a device from Topcon Healthcare that combines swept-source OCT technology with multimodal fundus imaging. The DRI OCT Triton uses swept-source technology to visualize the deepest layers of the eye, including through cataracts. It also enhances visualization of outer retinal structures and deep pathologies. The DRI OCT Triton has a 1,050 nm wavelength light source and a non-mydriatic color fundus camera. AI READI imaging is done in a dark room with minimal intervention from the imager as the machine positioning is done automatically. This leads to higher quality images with minimal operator induced error. Imaging is done in 12.0X12.0 mm and 6.0X6.0 mm OCTA, and 12.0 mm X9.0 mmX6.0 mm 3D Horizontal and Radial scan modes.\n\n   **Maestro2 3D OCT (Topcon Healthcare)**\n\n   The Maestro2 is a robotic OCT and color fundus camera system from Topcon Healthcare. It can capture a 12 mm x 9 mm wide-field OCT scan that includes the macula and optic disc. The Maestro2 can also capture high-resolution non-mydriatic, true color fundus photography, OCT, and OCTA with a single button press. Imaging is done in a dark room and automatically with minimal involvement of the operator. Protocols include 12.0 mm X9.0 mm widefield, 6.0 mm X 6.0 mm 3D macula scan and 6.0 mm X 6.0 mm OCTA (scan rate: 50 kHz).\n\n   **FLIO (Heidelberg Engineering)**\n\n   Fluorescence Lifetime Imaging Ophthalmoscopy (FLIO) is an advanced imaging technique used in ophthalmology. It is a non-invasive method that provides valuable information about the metabolic and functional status of the retina. FLIO is based on the measurement of fluorescence lifetimes, which is the duration a fluorophore remains in its excited state before emitting a photon and returning to the ground state. FLIO utilizes this fluorescence lifetime information to capture and analyze the metabolic processes occurring in the retina. Different retinal structures and molecules exhibit distinct fluorescence lifetimes, allowing for the visualization of metabolic changes, cellular activity, and the identification of specific biomolecules. The imaging is done by an operator in a dark room analogous to a straightforward heidelberg spectralis OCT. However, as it takes longer than a usual spectralis OCT and exposes patients to uncomfortable levels of light, it is kept to be performed as the last modality of an AI READI visit. Because of this patients may not be at their best possible compliance.\n\n   **Cirrus 5000 Angioplex (Carl Zeiss Meditec)**\n\n   The Zeiss Cirrus 5000 Angioplex is a high-definition optical coherence tomography (OCT) system that offers non-invasive imaging of retinal microvasculature. The imaging is done in a dark room by an operator and it is pretty straightforward and analogous to what is done in the ophthalmology clinics on a day to day basis. Imaging protocols include 512 X 512 and 200 X 200 macula and ONH scans and also OCTA of the macula. Zeiss Cirrus 5000 also provides a 60-degree OCTA widefield view. 8x8mm single scans and 14x14mm automated OCTA montage allow for rapid peripheral assessment of the retina as well.\n\n   **Monofilament testing for peripheral neuropathy**\n\n   Monofilament test is a standard clinical test to monitor peripheral neuropathy in diabetic patients. It is done using a standard 10g monofilament applying pressure to different points on the plantar surface of the feet. If patients sense the monofilament, they confirm by saying “yes”; if patients do not sense the monofilament after it bends, they are considered to be insensate. When the sequence is completed, the insensate area is retested for confirmation. This sequence is further repeated randomly at each of the testing sites on each foot until results are obtained.The results are recorded on an iPad, Laptop, or a paper questionnaire and are directly added to the project's RedCap by the clinical research staff.\n\n   **Montreal Cognitive Assessment (MoCA)**\n\n   The Montreal Cognitive Assessment (MoCA) is a simple, in-office screening tool that helps detect mild cognitive impairment and early onset of dementia. The MoCA evaluates cognitive domains such as: Memory, Executive functioning, Attention, Language, Visuospatial, Orientation, Visuoconstructional skills, Conceptual thinking, Calculations. The MoCA generates a total score and six domain-specific index scores. The maximum score is 30, and anything below 24 is a sign of cognitive impairment. A final total score of 26 and above is considered normal. Some disadvantages of the MoCA include: Professionals require training to score the test, A person's level of education may affect the test, Socioeconomic factors may affect the test, People living with depression or other mental health issues may score similarly to those with mild dementia. AI READI research staff perform this test on an iPad using a pre-installed software (MoCA Duo app downloaded from the app store) that captures all the patients responses in an interactive manner."
        }
      ],
      "challenge": [
        {
          "id": 1,
          "question": "Which factors in the data might limit the generalization of potentially derived models? Is this information available as auxiliary labels for challenge tests? For instance:\n\n   a. Number and diversity of devices included in the dataset.\n\n   b. Data recording specificities, e.g., the view for a chest x-ray image.\n\n   c. Number and diversity of recording sites included in the dataset.\n\n   d. Distribution shifts over time.",
          "response": "While the AI-READI's cross-sectional database ultimately aims to achieve balance across race/ethnicity, biological sex, and diabetes presence and severity, the pilot study is not balanced across these parameters.\n\n   Three recording sites were strategically selected to achieve diverse recruitment: the University of Alabama at Birmingham (UAB), the University of California San Diego (UCSD), and the University of Washington (UW). The sites were chosen for geographic diversity across the United States and to ensure diverse representation across various racial and ethnic groups. Individuals from all demographic backgrounds were recruited at all 3 sites.\n\n   Factors influencing the generalization of derived models include the predominantly urban and hospital-based recruitment, which may not fully capture diverse cultural and socioeconomic backgrounds. The study cohort may not provide a comprehensive representation of the population, as it does not include other races/ethnicities such as Pacific Islanders and Native Americans.\n\n   Information on device make and model, including specific modalities like macula scans or wide scans during OCT, were documented to ensure repeatability. Moreover, the study included multiple devices for one measure to enhance generalizability and represent the diverse range of equipment utilized in clinical settings."
        },
        {
          "id": 2,
          "question": "What confounding factors might be present in the data?\n\n   a. Interactions between demographic or historically marginalized groups and data recordings, e.g., were women patients recorded in one site, and men in another?\n\n   b. Interactions between the labels and data recordings, e.g. were healthy patients recorded on one device and diseased patients on another?",
          "response": "Uniform data collection protocols were implemented for all subjects, irrespective of their race/ethnicity, biological sex, or diabetes severity, across all study sites. The selection of study sites was intended to ensure equitable representation and minimize the potential for sampling bias."
        }
      ],
      "demographic_information": [
        {
          "id": 1,
          "question": "Does the dataset identify any demographic sub-populations (e.g., by age, gender, sex, ethnicity)?",
          "response": "No"
        },
        {
          "id": 2,
          "question": "If no,\n\n   a. Is there any regulation that prevents demographic data collection in your study (for example, the country that the data is collected in)?",
          "response": "No."
        },
        {
          "id": 3,
          "question": "b. Are you employing methods to reduce the disparity of error rate between different demographic subgroups when demographic labels are unavailable? Please describe.",
          "response": "We are suggesting a split for training/validation/testing models that is aimed at reducing disparities in models developed using this dataset."
        }
      ],
      "preprocessing": [
        {
          "id": 1,
          "question": "Was there any pre-processing for the de-identification of the patients? Provide the answer for the preliminary and the current version of the dataset",
          "response": ""
        },
        {
          "id": 2,
          "question": "Was there any pre-processing for cleaning the data? Provide the answer for the preliminary and the current version of the dataset",
          "response": "There were several quality control measures used at the time of data entry/acquisition. For example, clinical data outside of expected min/max ranges were flagged in REDCap, which was visible in reports viewed by clinical research coordinators (CRCs) and Data Managers. Using these REDCap reports as guides, Data Managers and CRCs examined participant records and determined if an error was likely. Data were checked for the following and edited if errors were detected:\n\n   1. Credibility, based on range checks to determine if all responses fall within a prespecified reasonable range\n\n   2. Incorrect flow through prescribed skip patterns\n\n   3. Missing data that can be directly filed from other portions of an individual's record\n\n   4. The omission and/or duplication of records\n\n   Editing was only done under the guidance and approval of the site PI. If corrected data was available from elsewhere in the respondent's answers, the error was corrected. If there was no logical or appropriate way to correct the data, the Data site PI reviewed the values and made decisions about whether those values should be removed from the data.\n\n   Once data were sent from each of the study sites to the central project team, additional processing steps were conducted in preparation for dissemination. For example, all data were mapped to standardized terminologies when possible, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model, a common data model for observational health data, and the Digital Imaging and Communications in Medicine (DICOM), a commonly used standard for medical imaging data. Details about the data processing approaches for each data domain/modality are described in the dataset documentation at https://docs.aireadi.org."
        },
        {
          "id": 3,
          "question": "Was the “raw” data (post de-identification) saved in addition to the preprocessed/cleaned data (e.g., to support unanticipated future uses)? If so, please provide a link or other access point to the “raw” data",
          "response": "The raw data is saved and expected to be preserved by the AI-READI project at least for the duration of the project but is not anticipated to be shared outside the project team right now, because it has not been mapped to standardized terminologies and because the raw data may accidentally include personal health information or personally identifiable information (e.g. in free text fields). There is a possibility that raw data may be included in future releases of the controlled access dataset."
        },
        {
          "id": 4,
          "question": "Were instances excluded from the dataset at the time of preprocessing? If so, why? For example, instances related to patients under 18 might be discarded.",
          "response": "No data were excluded from the dataset at the time of preprocessing. However, regarding to study recruitment (i.e. ability to participate in the study), the following eligibility criteria were used:\n\n   Inclusion Criteria:\n\n   - Able to provide consent\n   - ≥ 40 years old\n   - Persons with or without type 2 diabetes\n   - Must speak and read English\n\n   Exclusion Criteria:\n\n   - Must not be pregnant\n   - Must not have gestational diabetes\n   - Must not have Type 1 diabetes"
        },
        {
          "id": 5,
          "question": "If the dataset is a sample from a larger set, what was the sampling strategy (e.g., deterministic, probabilistic with specific sampling probabilities)? Answer this question for both the preliminary dataset and the current version of the dataset",
          "response": "N/A"
        }
      ],
      "labeling": [
        {
          "id": 1,
          "question": "Is there an explicit label or target associated with each data instance? Please respond for both the preliminary dataset and the current version.\n\n   a. If yes:\n\n   1. What are the labels provided?\n\n   2. Who performed the labeling? For example, was the labeling done by a clinician, ML researcher, university or hospital?",
          "response": "N/A - no labels are provided"
        },
        {
          "id": 2,
          "question": "b. What labeling strategy was used?\n\n   1. Gold standard label available in the data (e.g. cancers validated by biopsies)\n\n   2. Proxy label computed from available data:\n\n      1. Which label definition was used? (e.g. Acute Kidney Injury has multiple definitions)\n\n      2. Which tables and features were considered to compute the label?\n\n   3. Which proportion of the data has gold standard labels?",
          "response": "N/A - no labels are provided"
        },
        {
          "id": 3,
          "question": "c. Human-labeled data\n\n   1. How many labellers were considered?\n\n   2. What is the demographic of the labellers? (countries of residence, of origin, number of years of experience, age, gender, race, ethnicity, …)\n\n   3. What guidelines did they follow?\n\n   4. How many labellers provide a label per instance?\n\n      If multiple labellers per instance:\n\n      1. What is the rater agreement? How was disagreement handled?\n      2. Are all labels provided, or summaries (e.g. maximum vote)?\n\n   5. Is there any subjective source of information that may lead to inconsistencies in the responses? (e.g: multiple people answering a survey having different interpretation of scales, multiple clinicians using scores, or notes)\n\n   6. On average, how much time was required to annotate each instance?\n\n   7. Were the raters compensated for their time? If so, by whom and what amount? What was the compensation strategy (e.g. fixed number of cases, compensated per hour, per cases per hour)?",
          "response": "N/A - no labels are provided. No specific labeling was performed in the dataset, as the dataset is a hypothesis-agnostic dataset aimed at facilitating multiple potential downstream AI/ML applications."
        },
        {
          "id": 4,
          "question": "What are the human level performances in the applications that the dataset is supposed to address?",
          "response": "N/A"
        },
        {
          "id": 5,
          "question": "Is the software used to preprocess/clean/label the instances available? If so, please provide a link or other access point.",
          "response": "N/A - no labeling was performed"
        },
        {
          "id": 6,
          "question": "Is there any guideline that the future researchers are recommended to follow when creating new labels / defining new tasks?",
          "response": "No, we do not have formal guidelines in place."
        },
        {
          "id": 7,
          "question": "Are there recommended data splits (e.g., training, development/validation, testing)? Are there units of data to consider, whatever the task? If so, please provide a description of these splits, explaining the rationale behind them. Please provide the answer for both the preliminary dataset and the current version or any sub-version that is widely used.",
          "response": "The current version of the dataset comes with recommended data splits. Because sex, race, and ethnicity data are not being released with the public version of the dataset, the project team has prepared data splits into proportions (70%/15%/15%) that can be used for subsequent training/validation/testing where the validation and test sets are balanced as well as possible for sex, race/ethnicity and diabetes status (no diabetes, prediabetes/lifestyle controlled, oral medication controlled, and insulin controlled)."
        }
      ],
      "collection": [
        {
          "id": 1,
          "question": "Were any REB/IRB approval (e.g., by an institutional review board or research ethics board) received? If so, please provide a description of these review processes, including the outcomes, as well as a link or other access point to any supporting documentation.",
          "response": "The initial IRB approval at the University of Washington was received on December 20, 2022. The initial approval letter can be found [here](https://docs.aireadi.org/v3/Approval_STUDY00016228_Lee_initial.pdf). An annual renewal application to the IRB about the status and progress of the study is required and due within 90 days of expiration."
        },
        {
          "id": 2,
          "question": "How was the data associated with each instance acquired? Was the data directly observable (e.g., medical images, labs or vitals), reported by subjects (e.g., survey responses, pain levels, itching/burning sensations), or indirectly inferred/derived from other data (e.g., part-of-speech tags, model-based guesses for age or language)? If data was reported by subjects or indirectly inferred/derived from other data, was the data validated/verified? If so, please describe how.",
          "response": "The acquisition of data varied based on the domain; some data were directly observable (such as labs, vitals, and retinal imaging), whereas other data were reported by subjects (e.g. survey responses). Verification of data entry was performed when possible (e.g. cross-referencing entered medications with medications that were physically brought in or photographed by each study participant). Details for each data domain are available in https://docs.aireadi.org."
        },
        {
          "id": 3,
          "question": "What mechanisms or procedures were used to collect the data (e.g., hardware apparatus or sensor, manual human curation, software program, software API)? How were these mechanisms or procedures validated? Provide the answer for all modalities and collected data. Has this information been changed through the process? If so, explain why.",
          "response": "The procedures for data collection and processing is available at https://docs.aireadi.org."
        },
        {
          "id": 4,
          "question": "Who was involved in the data collection process (e.g., patients, clinicians, doctors, ML researchers, hospital staff, vendors, etc.) and how were they compensated (e.g., how much were contributors paid)?",
          "response": "Details about the AI-READI team members involved in the data collection process are available at https://aireadi.org/team. Their effort was supported by the National Institutes of Health award OT2OD032644 based on the percentage of effort contributed, and salaries which aligned with the funding guidelines at each site. Study subjects received a compensation of $200 for the study visit also through the grant funding."
        },
        {
          "id": 5,
          "question": "Over what timeframe was the data collected? Does this timeframe match the creation timeframe of the data associated with the instances (e.g., recent crawl of old news articles)? If not, please describe the timeframe in which the data associated with the instances was created.",
          "response": "The timeline for the overall project spans four years, encompassing one year dedicated to protocol development and training, and years 2-4 allocated for subject recruitment and data collection. Approximately 4% of participants are expected to undergo a follow-up examination in Year 4. The data collection process is specifically tailored to enable downstream pseudotime manifold analysis—an approach used to predict disease trajectories. This involves gathering and learning from complex, multimodal data from participants exhibiting varying disease severity, ranging from normal to insulin-dependent Type 2 Diabetes Mellitus (T2DM). The timeframe also allows for the collection of the highest number of subjects possible to ensure a balanced representation of racial and ethnic groups and mitigate biases in computer vision algorithms.\n\n For this version of the dataset, the timeframe for data collection was July 19, 2023 to May 01, 2025."
        },
        {
          "id": 6,
          "question": "Does the dataset relate to people? If not, you may skip the remaining questions in this section.",
          "response": "Yes"
        },
        {
          "id": 7,
          "question": "Did you collect the data from the individuals in question directly, or obtain it via third parties or other sources (e.g., hospitals, app company)?",
          "response": "The data was collected directly from participants across the three recruiting sites. Recruitment pools were identified by screening Electronic Health Records (EHR) for diabetes and prediabetes ICD-10 codes for all patients who have had an encounter with the sites' health systems within the past 2 years."
        },
        {
          "id": 8,
          "question": "Were the individuals in question notified about the data collection? If so, please describe (or show with screenshots or other information) how notice was provided, and provide a link or other access point to, or otherwise reproduce, the exact language of the notification itself.",
          "response": "Yes, each individual was aware of the data collection, as this was not passive data collection or secondary use of existing data, but rather active data collection directly from participants."
        },
        {
          "id": 9,
          "question": "Did the individuals in question consent to the collection and use of their data? If so, please describe (or show with screenshots or other information) how consent was requested and provided, and provide a link or other access point to, or otherwise reproduce, the exact language to which the individuals consented.",
          "response": "Informed consent to participate was required before participation in any part of the protocol (including questionnaires). Potential participants were given the option to read all consent documentation electronically (e-consent) before their visit and give their consent with an electronic signature without verbal communication with a clinical research coordinator. Participants may access e-consent documentation in REDCap and decide at that point they do not want to participate or would like additional information. The approved consent form for the principal project site University of Washington is available [here](https://docs.aireadi.org/v3/AI-READI%20Consent_Form_Standard_Mod17May2023_useforpilot.pdf). The other clinical sites had IRB reliance and used the same consent form, with minor institution-specific language incorporated depending on individual institutional requirements."
        },
        {
          "id": 10,
          "question": "If consent was obtained, were the consenting individuals provided with a mechanism to revoke their consent in the future or for certain uses? If so, please provide a description, as well as a link or other access point to the mechanism (if appropriate).",
          "response": "Participants were permitted to withdraw consent at any time and cease study participation. However, any data that had been shared or used up to that point would stay in the dataset. This is clearly communicated in the consent document."
        },
        {
          "id": 11,
          "question": "In which countries was the data collected?",
          "response": "USA"
        },
        {
          "id": 12,
          "question": "Has an analysis of the potential impact of the dataset and its use on data subjects (e.g., a data protection impact analysis) been conducted? If so, please provide a description of this analysis, including the outcomes, as well as a link or other access point to any supporting documentation.",
          "response": "No, a data protection impact analysis has not been conducted."
        }
      ],
      "inclusion": [
        {
          "id": 1,
          "question": "Is there any language-based communication with patients (e.g: English, French)? If yes, describe the choices of language(s) for communication. (for example, if there is an app used for communication, what are the language options?)",
          "response": "English language was used for communication with study participants."
        },
        {
          "id": 2,
          "question": "What are the accessibility measurements and what aspects were considered when the study was designed and implemented?",
          "response": "Accessibility measurements were not specifically assessed. However, transportation assistance (rideshare services) was offered to study participants who endorsed barriers to transporting themselves to study visits."
        },
        {
          "id": 3,
          "question": "If data is part of a clinical study, what are the inclusion criteria?",
          "response": "The eligibility criteria for the study were as follows:\n\n Inclusion Criteria:\n\n - Able to provide consent\n - ≥ 40 years old\n - Persons with or without type 2 diabetes\n - Must speak and read English\n\n Exclusion Criteria:\n\n - Must not be pregnant\n - Must not have gestational diabetes\n - Must not have Type 1 diabetes"
        }
      ],
      "uses": [
        {
          "id": 1,
          "question": "Has the dataset been used for any tasks already? If so, please provide a description. ",
          "response": "No"
        },
        {
          "id": 2,
          "question": "Does using the dataset require the citation of the paper or any other forms of acknowledgement? If yes, is it easily accessible through google scholar or other repositories",
          "response": "Yes, use of the dataset requires citation to the resources specified in https://docs.aireadi.org."
        },
        {
          "id": 3,
          "question": "Is there a repository that links to any or all papers or systems that use the dataset? If so, please provide a link or other access point. (besides Google scholar)",
          "response": "No"
        },
        {
          "id": 4,
          "question": "Is there anything about the composition of the dataset or the way it was collected and preprocessed/cleaned/labeled that might impact future uses? For example, is there anything that a future user might need to know to avoid uses that could result in unfair treatment of individuals or groups (e.g., stereotyping, quality of service issues) or other undesirable harms (e.g., financial harms, legal risks) If so, please provide a description. Is there anything a future user could do to mitigate these undesirable harms?",
          "response": "No, to the extent of our knowledge, we do not currently anticipate any uses of the dataset that could result in unfair treatment or harm. However, there is a theoretical risk of future re-identification."
        },
        {
          "id": 5,
          "question": "Are there tasks for which the dataset should not be used? If so, please provide a description. (for example, dataset creators could recommend against using the dataset for considering immigration cases, as part of insurance policies)",
          "response": "This is answered in a prior question (see details regarding license terms)."
        }
      ],
      "distribution": [
        {
          "id": 1,
          "question": "Will the dataset be distributed to third parties outside of the entity (e.g., company, institution, organization) on behalf of which the dataset was created? If so, please provide a description.",
          "response": "The dataset will be distributed and be available for public use."
        },
        {
          "id": 2,
          "question": "How will the dataset be distributed (e.g., tarball on website, API, GitHub)? Does the dataset have a digital object identifier (DOI)?",
          "response": "The dataset will be available through the FAIRhub platform (http://fairhub.io/). The dataset' DOI is https://doi.org/10.60775/fairhub.3"
        },
        {
          "id": 3,
          "question": "When was/will the dataset be distributed?",
          "response": "The first version of the dataset was distributed in May 2024, the second version of the dataset was distributed in November 2024, and the third version of the dataset was distributed in November 2025."
        },
        {
          "id": 4,
          "question": "Assuming the dataset is available, will it be/is the dataset distributed under a copyright or other intellectual property (IP) license, and/or under applicable terms of use (ToU)? If so, please describe this license and/or ToU, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms or ToU, as well as any fees associated with these restrictions.",
          "response": "We provide here the license file containing the terms for reusing the AI-READI dataset (https://doi.org/10.5281/zenodo.17555036). These license terms were specifically tailored to enable reuse of the AI-READI dataset (and other clinical datasets) for commercial or research purpose while putting strong requirements around data usage, security, and secondary sharing to protect study participants, especially when data is reused for artificial intelligence (AI) and machine learning (ML) related applications."
        },
        {
          "id": 5,
          "question": "Have any third parties imposed IP-based or other restrictions on the data associated with the instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms, as well as any fees associated with these restrictions.",
          "response": "Refer to license (https://doi.org/10.5281/zenodo.17555036)"
        },
        {
          "id": 6,
          "question": "Do any export controls or other regulatory restrictions apply to the dataset or to individual instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any supporting documentation.",
          "response": "Refer to license (https://doi.org/10.5281/zenodo.17555036)"
        }
      ],
      "maintenance": [
        {
          "id": 1,
          "question": "Who is supporting/hosting/maintaining the dataset?",
          "response": "The AI-READI team will be supporting and maintaining the dataset. The dataset is hosted on FAIRhub through Microsoft Azure."
        },
        {
          "id": 2,
          "question": "How can the owner/curator/manager of the dataset be contacted (e.g. email address)?",
          "response": "We refer to the README file included with the dataset for contact information."
        },
        {
          "id": 3,
          "question": "Is there an erratum? If so, please provide a link or other access point.",
          "response": ""
        },
        {
          "id": 4,
          "question": "Will the dataset be updated (e.g., to correct labeling errors, add new instances, delete instances)? If so, please describe how often, by whom, and how updates will be communicated to users (e.g., mailing list, GitHub)?",
          "response": "The dataset will not be updated. Rather, new versions of the dataset will be released with additional instances as more study participants complete the study visit."
        },
        {
          "id": 5,
          "question": "If the dataset relates to people, are there applicable limits on the retention of the data associated with the instances (e.g., were individuals in question told that their data would be retained for a fixed period of time and then deleted)? If so, please describe these limits and explain how they will be enforced.",
          "response": "There are no limits on the retention of the data associated with the instances."
        },
        {
          "id": 6,
          "question": "Will older versions of the dataset continue to be supported/hosted/maintained? If so, please describe how and for how long. If not, please describe how its obsolescence will be communicated to users.",
          "response": "N/A - as mentioned in the response to question 4, the dataset will not be updated. Rather, new versions of the dataset will be released with additional instances as more study participants are enrolled."
        },
        {
          "id": 7,
          "question": "If others want to extend/augment/build on/contribute to the dataset, is there a mechanism for them to do so?",
          "response": "No, currently there is no mechanism for others to extend or augment the AI-READI dataset outside of those who are involved in the project."
        }
      ]
    },
    "readme": "## Overview of the study\n\nThe Artificial Intelligence Ready and Exploratory Atlas for Diabetes Insights (AI-READI) project seeks to create a flagship ethically-sourced dataset to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. The ability to understand and affect the course of complex, multi-organ diseases such as T2DM has been limited by a lack of well-designed, high quality, large, and inclusive multimodal datasets. The AI-READI team of investigators will aim to collect a cross-sectional dataset of 4,000 people and longitudinal data from 10% of the study cohort across the US. The study cohort will be balanced for self-reported race/ethnicity, gender, and diabetes disease stage. Data collection will be specifically designed to permit downstream pseudo-time manifold analysis, an approach used to predict disease trajectories by collecting and learning from complex, multimodal data from participants with differing disease severity (normal to insulin-dependent T2DM). The long-term objective for this project is to develop a foundational dataset in T2DM, agnostic to existing classification criteria or biases, which can be used to reconstruct a temporal atlas of T2DM development and reversal towards health (i.e., salutogenesis). Data will be optimized for downstream AI/ML research and made publicly available.\n\n## Description of the dataset\n\nThis dataset contains data from 2280 participants that was collected between July 19, 2023 and May 01, 2025. Data from multiple modalities are included. A full list is provided in the `Data Standards` section below. The data in this dataset contain no protected health information (PHI). Information related to the sex and race/ethnicity of the participants as well as medication used has also been removed.\n\nThe dataset contains 356,343 files and is around 3.82 TB in size.\n\nA detailed description of the dataset is available in the AI-READI documentation for v3.0.0 of the dataset at [docs.aireadi.org](https://docs.aireadi.org/).\n\n## Protocol\n\nThe protocol followed for collecting the data can be found in the AI-READI documentation for v3.0.0 of the dataset at [docs.aireadi.org](https://docs.aireadi.org/).\n\n## Dataset access/restrictions\n\nAccessing the dataset requires several steps, including:\n\n- Login in through a verified ID system\n- Agreeing to use the data only for type 2 diabetes related research.\n- Agreeing to the license terms which set certain restrictions and obligations for data usage (see `License` section below).\n\n## Data standards followed\n\nThis dataset is organized following the [Clinical Dataset Structure (CDS) v0.1.1](https://cds-specification.readthedocs.io/en/v0.1.1/). We refer to the CDS documentation for more details. Briefly, data is organized at the root level into one directory per datatype (c.f. Table below). Within each datatype folder, there is one folder per modality. Within each modality folder, there is one folder per device used to collect that modality. Within each device folder, there is one folder per participant. Each datatype, modality, and device folder is named using a name that best defines it. Each participant folder is named after the participant's ID number used in the study. For each datatype, the data files follow the standards listed in the Table below. More details are available in the dataset_structure_description.json metadata file included in this dataset.\n\n| Datatype directory name   | Description                                                                                                                                                                                      | File format standard followed                                                                                                                                     |\n| ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| cardiac_ecg               | This directory contains electrocardiogram data collected by a 12 lead protocol (the current standard), Holter monitor, or smartwatch. The terms ECG and EKG are often used interchangeably.      | [WaveForm DataBase (WFDB)](https://wfdb.readthedocs.io/en/latest/wfdb.html)                                                                                       |\n| clinical_data             | This directory contains clinical data collected through REDCap. Each CSV file in this directory is a one-to-one mapping to the OMOP CDM tables.                                                  | [Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM)](https://ohdsi.github.io/TheBookOfOhdsi)                                               |\n| environment               | This directory contains data collected through an environmental sensor device custom built for the AI-READI project.                                                                             | [Earth Science Data Systems (ESDS) format](https://www.earthdata.nasa.gov/esdis/esco/standards-and-practices/ascii-file-format-guidelines-for-earth-science-data) |\n| retinal_flio              | This directory contains data collected through fluorescence lifetime imaging ophthalmoscopy (FLIO), an imaging modality for in vivo measurement of lifetimes of endogenous retinal fluorophores. | [Digital Imaging and Communications in Medicine (DICOM)](http://medical.nema.org/)                                                                                |\n| retinal_oct               | This directory contains data collected using optical coherence tomography (OCT), an imaging method using lasers that is used for mapping subsurface structure.                                   | [Digital Imaging and Communications in Medicine (DICOM)](http://medical.nema.org/)                                                                                |\n| retinal_octa              | This directory contains data collected using optical coherence tomography angiography (OCTA), a non-invasive imaging technique that generates volumetric angiography images.                     | [Digital Imaging and Communications in Medicine (DICOM)](http://medical.nema.org/)                                                                                |\n| retinal_photography       | This directory contains retinal photography data, which are 2D images. They are also referred to as fundus photography.                                                                          | [Digital Imaging and Communications in Medicine (DICOM)](http://medical.nema.org/)                                                                                |\n| wearable_activity_monitor | This directory contains data collected through a wearable fitness tracker.                                                                                                                       | [Open mHealth](https://www.openmhealth.org/documentation/#/schema-docs/schema-library)                                                                            |\n| wearable_blood_glucose    | This directory contains data collected through a continuous glucose monitoring (CGM) device.                                                                                                     | [Open mHealth](https://www.openmhealth.org/documentation/#/schema-docs/schema-library)                                                                            |\n\n## Resources\n\nAll of our data files are in formats that are accessible with free software commonly used for such data types so no specific software is required. Some useful resources related to this dataset are listed below:\n\n- Documentation of the dataset: [docs.aireadi.org](https://docs.aireadi.org/) (see 'Dataset v3.0.0' for this version of the dataset)\n- AI-READI project website: [aireadi.org](https://aireadi.org/)\n- Zenodo community of the AI-READI project: [zenodo.org/communities/aireadi](https://zenodo.org/communities/aireadi)\n- GitHub organization of the AI-READI project: [github.com/AI-READI](https://github.com/AI-READI)\n\n### Suggested split\n\nThe suggested split for training, validating, and testing AI/ML models is included in the participants.tsv file that will be included in the dataset. A summary is provided in the table below.\n\n|                         | Train       |           |       |         | Val         |           |       |         | Test        |           |       |         | Total       |           |       |         |\n| ----------------------- |-------------|-----------|-------|---------|-------------|-----------|-------|---------|-------------|-----------|-------|---------|-------------|-----------|-------|---------|\n|                         | Hispanic    | Asian     | Black | White   | Hispanic    | Asian     | Black | White   | Hispanic    | Asian     | Black | White   | Hispanic    | Asian     | Black | White   |\n| Race/ethnicity (count)  | 204         | 369       | 343   | 660     | 88          | 88        | 88    | 88      | 88          | 88        | 88    | 88      | 380         | 545       | 519   | 836     |\n|                         | Male        | Female    |       |         | Male        | Female    |       |         | Male        | Female    |       |         | Male        | Female    |       |         |\n| Sex (count)             | 599         | 977       |       |         | 176         | 176       |       |         | 176         | 176       |       |         | 951         | 1329      |       |         |\n|                         | No DM       | Lifestyle | Oral  | Insulin | No DM       | Lifestyle | Oral  | Insulin | No DM       | Lifestyle | Oral  | Insulin | No DM       | Lifestyle | Oral  | Insulin |\n| Diabetes status (count) | 600         | 384       | 487   | 105     | 88          | 88        | 109   | 67      | 88          | 88        | 90    | 86      | 776         | 560       | 686   | 258     |\n| Mean age (years ± sd)   | 61.1 ± 11.3 |           |       |         | 60.3 ± 10.8 |           |       |         | 60.5 ± 11.4 |           |       |         | 60.8 ± 11.3 |           |       |         |\n| Total                   | 1576        |           |       |         | 352         |           |       |         | 352         |           |       |         | 2280        |           |       |         |\n\n- No DM: Participants who do not have Type 1 or Type 2 Diabetes\n- Lifestyle: Participants with pre-Type 2 Diabetes and those with Type 2 Diabetes whose blood sugar is controlled by lifestyle adjustments\n- Oral: Participants with Type 2 Diabetes whose blood sugar is controlled by oral or injectable medications other than insulin\n- Insulin: Participants with Type 2 Diabetes whose blood sugar is controlled by insulin\n\n### Changes between versions of the dataset\n\nChanges between the current version of the dataset and the previous one are provided in details in the CHANGELOG file included in the dataset (also visible at docs.aireadi.org). A summary of the major changes is provided in the table below.\n\n| Dataset      | v1.0.0 pilot    | year 2 data        | year 3 data       | v3.0.0 main study  |\n| ------------ | --------------- | ------------------ |-------------------|--------------------|\n| Participants | 204             | 863                | 1213              | 2280               |\n| Data types   | 15+ data types  | +1 image device    | +3 image device   | 15+ data types     |\n| Processing   | custom / ad hoc | automated + custom |automated + custom | automated + custom |\n| Release date | 5/3/2024        | included in v2.0.0 |included in v3.0.0 | 11/17/2025         |\n\n## License\n\nThis work is licensed under a custom license specifically tailored to enable the reuse of the AI-READI dataset (and other clinical datasets) for commercial or research purposes while putting strong requirements around data usage, security, and secondary sharing to protect study participants, especially when data is reused for artificial intelligence (AI) and machine learning (ML) related applications. More details are available in the License file included in the dataset and also available at https://doi.org/10.5281/zenodo.17555036.\n\n## How to cite\n\nIf you use this dataset for any purpose, please cite the resources specified in the AI-READI documentation for version 3.0.0 of the dataset at https://docs.aireadi.org.\n\n## Contact\n\nFor any questions, suggestions, or feedback related to this dataset, please go to https://aireadi.org/contact. We refer to the study_description.json and dataset_description.json metadata files included in this dataset for additional information about the contact person/entity, authors, and contributors of the dataset.\n\n## Acknowledgement\n\nThe AI-READI project is supported by NIH grant [1OT2OD032644](https://reporter.nih.gov/search/1ADgncihCk6fdMRJdCnBjg/project-details/10471118) through the NIH Bridge2AI Common Fund program.",
    "studyDescription": {
      "schema": "https://schema.aireadi.org/v0.1.0/study_description.json",
      "identificationModule": {
        "officialTitle": "AI Ready and Exploratory Atlas for Diabetes Insights",
        "acronym": "AI-READI",
        "orgStudyIdInfo": {
          "orgStudyId": "OT2OD032644",
          "orgStudyIdType": "U.S. National Institutes of Health (NIH) Grant/Contract Award Number",
          "orgStudyIdLink": "https://reporter.nih.gov/search/yatARMM-qUyKAhnQgsCTAQ/project-details/10885481"
        },
        "secondaryIdInfoList": [
          {
            "secondaryId": "NCT06002048",
            "secondaryIdType": "ClinicalTrials.gov",
            "secondaryIdLink": "https://classic.clinicaltrials.gov/ct2/show/NCT06002048"
          }
        ]
      },
      "statusModule": {
        "overallStatus": "Enrolling by invitation",
        "startDateStruct": {
          "startDate": "2023-07-19",
          "startDateType": "Actual"
        },
        "completionDateStruct": {
          "completionDate": "2027-01-01",
          "completionDateType": "Anticipated"
        }
      },
      "sponsorCollaboratorsModule": {
        "leadSponsor": {
          "leadSponsorName": "Washington University in St. Louis",
          "leadSponsorIdentifier": {
            "leadSponsorIdentifierValue": "https://ror.org/01yc7t268",
            "leadSponsorIdentifierScheme": "ROR",
            "schemeURI": "https://ror.org/"
          }
        },
        "responsibleParty": {
          "responsiblePartyType": "Principal Investigator",
          "responsiblePartyInvestigatorFirstName": "Aaron",
          "responsiblePartyInvestigatorLastName": "Lee",
          "responsiblePartyInvestigatorTitle": "Associate Professor",
          "responsiblePartyInvestigatorIdentifier": [
            {
              "responsiblePartyInvestigatorIdentifierValue": "https://orcid.org/0000-0002-7452-1648",
              "responsiblePartyInvestigatorIdentifierScheme": "ORCID",
              "schemeURI": "https://orcid.org/"
            }
          ],
          "responsiblePartyInvestigatorAffiliation": {
            "responsiblePartyInvestigatorAffiliationName": "Washington University in St. Louis",
            "responsiblePartyInvestigatorAffiliationIdentifier": {
              "responsiblePartyInvestigatorAffiliationIdentifierValue": "https://ror.org/01yc7t268",
              "responsiblePartyInvestigatorAffiliationIdentifierScheme": "ROR",
              "schemeURI": "https://ror.org/"
            }
          }
        },
        "collaboratorList": [
          {
            "collaboratorName": "National Institutes of Health",
            "collaboratorNameIdentifier": {
              "collaboratorNameIdentifierValue": "https://ror.org/01cwqze88",
              "collaboratorNameIdentifierScheme": "ROR",
              "schemeURI": "https://ror.org/"
            }
          },
          {
            "collaboratorName": "California Medical Innovations Institute",
            "collaboratorNameIdentifier": {
              "collaboratorNameIdentifierValue": "https://ror.org/0156zyn36",
              "collaboratorNameIdentifierScheme": "ROR",
              "schemeURI": "https://ror.org/"
            }
          },
          {
            "collaboratorName": "Johns Hopkins University",
            "collaboratorNameIdentifier": {
              "collaboratorNameIdentifierValue": "https://ror.org/00za53h95",
              "collaboratorNameIdentifierScheme": "ROR",
              "schemeURI": "https://ror.org/"
            }
          },
          {
            "collaboratorName": "Oregon Health & Science University",
            "collaboratorNameIdentifier": {
              "collaboratorNameIdentifierValue": "https://ror.org/009avj582",
              "collaboratorNameIdentifierScheme": "ROR",
              "schemeURI": "https://ror.org/"
            }
          },
          {
            "collaboratorName": "Stanford University",
            "collaboratorNameIdentifier": {
              "collaboratorNameIdentifierValue": "https://ror.org/00f54p054",
              "collaboratorNameIdentifierScheme": "ROR",
              "schemeURI": "https://ror.org/"
            }
          },
          {
            "collaboratorName": "University of Alabama at Birmingham",
            "collaboratorNameIdentifier": {
              "collaboratorNameIdentifierValue": "https://ror.org/008s83205",
              "collaboratorNameIdentifierScheme": "ROR",
              "schemeURI": "https://ror.org/"
            }
          },
          {
            "collaboratorName": "University of California, San Diego",
            "collaboratorNameIdentifier": {
              "collaboratorNameIdentifierValue": "https://ror.org/0168r3w48",
              "collaboratorNameIdentifierScheme": "ROR",
              "schemeURI": "https://ror.org/"
            }
          }
        ]
      },
      "oversightModule": {
        "isFDARegulatedDrug": "No",
        "isFDARegulatedDevice": "No",
        "humanSubjectReviewStatus": "Submitted, approved",
        "oversightHasDMC": "No"
      },
      "descriptionModule": {
        "briefSummary": "The study will collect a cross-sectional dataset of 4000 people across the US from diverse racial/ethnic groups who are either 1) healthy, or 2) belong in one of the three stages of diabetes severity (pre-diabetes/lifestyle controlled, oral medication and/or non-insulin-injectable medication controlled, or insulin dependent), forming a total of four groups of patients. Clinical data (social determinants of health surveys, continuous glucose monitoring data, biomarkers, genetic data, retinal imaging, cognitive testing, etc.) will be collected. The purpose of this project is data generation to allow future creation of artificial intelligence/machine learning (AI/ML) algorithms aimed at defining disease trajectories and underlying genetic links in different racial/ethnic cohorts. A smaller subgroup of participants will be invited to come for a follow-up visit in year 4 of the project (longitudinal arm of the study). Data will be placed in an open-source repository and samples will be sent to the study sample repository and used for future research.",
        "detailedDescription": "The Artificial Intelligence Ready and Exploratory Atlas for Diabetes Insights (AI-READI) project seeks to create a flagship ethically-sourced dataset to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. The ability to understand and affect the course of complex, multi-organ diseases such as T2DM has been limited by a lack of well-designed, high quality, large, and inclusive multimodal datasets. The AI-READI team of investigators will aim to collect a cross-sectional dataset of 4,000 people and longitudinal data from 10% of the study cohort across the US. The study cohort will be balanced for self-reported race/ethnicity, gender, and diabetes disease stage. Data collection will be specifically designed to permit downstream pseudo-time manifold analysis, an approach used to predict disease trajectories by collecting and learning from complex, multimodal data from participants with differing disease severity (normal to insulin-dependent T2DM). The long-term objective for this project is to develop a foundational dataset in T2DM, agnostic to existing classification criteria or biases, which can be used to reconstruct a temporal atlas of T2DM development and reversal towards health (i.e., salutogenesis). Data will be optimized for downstream AI/ML research and made publicly available."
      },
      "conditionsModule": {
        "conditionList": [
          {
            "conditionName": "Type 2 Diabetes",
            "conditionIdentifier": {
              "conditionClassificationCode": "D003924",
              "conditionScheme": "Medical Subject Headings (MeSH)",
              "schemeURI": "https://meshb.nlm.nih.gov/",
              "conditionURI": "https://meshb.nlm.nih.gov/record/ui?ui=D003924"
            }
          }
        ],
        "keywordList": [
          {
            "keywordValue": "Retinal Imaging"
          },
          {
            "keywordValue": "Data Sharing",
            "keywordIdentifier": {
              "keywordClassificationCode": "D033181",
              "keywordScheme": "Medical Subject Headings (MeSH)",
              "schemeURI": "https://meshb.nlm.nih.gov/",
              "keywordURI": "https://meshb.nlm.nih.gov/record/ui?ui=D033181"
            }
          },
          {
            "keywordValue": "Exploratory Data Collection"
          },
          {
            "keywordValue": "Machine Learning",
            "keywordIdentifier": {
              "keywordClassificationCode": "D000069550",
              "keywordScheme": "Medical Subject Headings (MeSH)",
              "schemeURI": "https://meshb.nlm.nih.gov/",
              "keywordURI": "https://meshb.nlm.nih.gov/record/ui?ui=D000069550"
            }
          },
          {
            "keywordValue": "Artificial Intelligence",
            "keywordIdentifier": {
              "keywordClassificationCode": "D001185",
              "keywordScheme": "Medical Subject Headings (MeSH)",
              "schemeURI": "https://meshb.nlm.nih.gov/",
              "keywordURI": "https://meshb.nlm.nih.gov/record/ui?ui=D001185"
            }
          },
          {
            "keywordValue": "Electrocardiography",
            "keywordIdentifier": {
              "keywordClassificationCode": "D004562",
              "keywordScheme": "Medical Subject Headings (MeSH)",
              "schemeURI": "https://meshb.nlm.nih.gov/",
              "keywordURI": "https://meshb.nlm.nih.gov/record/ui?ui=D004562"
            }
          },
          {
            "keywordValue": "Continuous Glucose Monitoring",
            "keywordIdentifier": {
              "keywordClassificationCode": "D000095583",
              "keywordScheme": "Medical Subject Headings (MeSH)",
              "schemeURI": "https://meshb.nlm.nih.gov/",
              "keywordURI": "https://meshb.nlm.nih.gov/record/ui?ui=D000095583"
            }
          },
          {
            "keywordValue": "Retinal imaging"
          },
          {
            "keywordValue": "Eye exam"
          }
        ]
      },
      "designModule": {
        "studyType": "Observational",
        "isPatientRegistry": "No",
        "designInfo": {
          "designObservationalModelList": [
            "Cohort"
          ],
          "designTimePerspectiveList": [
            "Cross-sectional"
          ]
        },
        "bioSpec": {
          "bioSpecRetention": "Samples With DNA",
          "bioSpecDescription": "Participants who consent are asked to provide both blood and urine for research use.  There are three distinct elements that follow from these collections; first, a small amount of the collected blood is sent to the local collection site hospital lab for a complete blood cell count on the day of the encounter. Second, an additional portion of the blood and the urine sample are processed and stored at the collection site for batch shipment to the University of Washington Nutrition and Obesity Research Center for specialized lab tests.  Third, the remaining majority of the collected blood is processed into a variety of derivatives that are being used to establish a large repository of biospecimens at the University of Alabama at Birmingham to be used in research to further our understanding of diabetes, diabetic associated eye disease and other applications related to health and related diseases. These derivatives include isolated plasma, serum, DNA, buffy coats (white blood cells), blood stabilized for future isolation of RNA and peripheral blood mononuclear cells (PBMC). These derivatives are separated into small aliquots to facilitate future approved research test and are stored at either -80oC or at cryogenic temperatures (PBMC)."
        },
        "enrollmentInfo": {
          "enrollmentCount": "4000",
          "enrollmentType": "Anticipated"
        }
      },
      "armsInterventionsModule": {
        "armGroupList": [
          {
            "armGroupLabel": "Healthy",
            "armGroupDescription": "Participants who do not have Type 1 or Type 2 Diabetes",
            "armGroupInterventionList": [
              "AI-READI protocol"
            ]
          },
          {
            "armGroupLabel": "Pre-diabetes/Lifestyle Controlled",
            "armGroupDescription": "Participants with pre-Type 2 Diabetes and those with Type 2 Diabetes whose blood sugar is controlled by lifesyle adjustments",
            "armGroupInterventionList": [
              "AI-READI protocol"
            ]
          },
          {
            "armGroupLabel": "Oral Medication and/or Non-insulin-injectable Medication Controlled",
            "armGroupDescription": "Participants with Type 2 Diabetes whose blood sugar is controlled by oral or injectable medications other than insulin",
            "armGroupInterventionList": [
              "AI-READI protocol"
            ]
          },
          {
            "armGroupLabel": "Insulin Dependent",
            "armGroupDescription": "Participants with Type 2 Diabetes whose blood sugar is controlled by insulin",
            "armGroupInterventionList": [
              "AI-READI protocol"
            ]
          }
        ],
        "interventionList": [
          {
            "interventionType": "Other",
            "interventionName": "AI-READI protocol",
            "interventionDescription": "Custom protocol followed for the AI-READI study. Details are available at in the documentation of v2.0.0 of the dataset at https://docs.aireadi.org/"
          }
        ]
      },
      "eligibilityModule": {
        "sex": "All",
        "genderBased": "No",
        "minimumAge": "40 Years",
        "maximumAge": "85 Years",
        "healthyVolunteers": "No",
        "eligibilityCriteria": {
          "eligibilityCriteriaInclusion": [
            "Adults (≥ 40 years old)",
            "Patients with and without type 2 diabetes",
            "Able to provide consent",
            "Must be able to read and speak English"
          ],
          "eligibilityCriteriaExclusion": [
            "Adults older than 85 years of age",
            "Pregnancy",
            "Gestational diabetes",
            "Type 1 diabetes"
          ]
        },
        "studyPopulation": "Adult patients will be recruited into one of four groups: 1) healthy/no diabetes, 2) pre-diabetes/borderline diabetes/lifestyle-controlled diabetes, 3) oral medication and/or non-insulin injectable medication controlled type 2 diabetes, or 4) insulin dependent type 2 diabetes. The investigators aim to recruit approximately 1000 patients into each of the four groups. Patients will be recruited from University of Washington (UW), University of California at San Diego (UCSD), and University of Alabama at Birmingham (UAB). The study aims to recruit 1,000 subjects from each of the following racial and ethnic groups: White, Asian, Hispanic, and Black. Subjects will be age- and sex-matched within and between groups.",
        "samplingMethod": "Non-Probability Sample"
      },
      "contactsLocationsModule": {
        "centralContactList": [
          {
            "centralContactFirstName": "Aaron",
            "centralContactLastName": "Lee",
            "centralContactDegree": "MD",
            "centralContactIdentifier": [
              {
                "centralContactIdentifierValue": "https://orcid.org/0000-0002-7452-1648",
                "centralContactIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "centralContactAffiliation": {
              "centralContactAffiliationName": "Washington University in St. Louis",
              "centralContactAffiliationIdentifier": {
                "centralContactAffiliationIdentifierValue": "https://ror.org/01yc7t268",
                "centralContactAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "centralContactEMail": "contact@aireadi.org"
          }
        ],
        "overallOfficialList": [
          {
            "overallOfficialFirstName": "Aaron",
            "overallOfficialLastName": "Lee",
            "overallOfficialDegree": "MD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0002-7452-1648",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "Washington University in St. Louis",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/01yc7t268",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Cecilia",
            "overallOfficialLastName": "Lee",
            "overallOfficialDegree": "MD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0003-1994-7213",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "Washington University in St. Louis",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/01yc7t268",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Amir",
            "overallOfficialLastName": "Bahmani",
            "overallOfficialDegree": "PhD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0003-4533-9334",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "Stanford University",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/00f54p054",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Sally L.",
            "overallOfficialLastName": "Baxter",
            "overallOfficialDegree": "MD, MSc",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0002-5271-7690",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "University of California, San Diego",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/0168r3w48",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Christopher G.",
            "overallOfficialLastName": "Chute",
            "overallOfficialDegree": "MD, DrPH",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0001-5437-2545",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "Johns Hopkins University",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/00za53h95",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Jorge",
            "overallOfficialLastName": "Contreras",
            "overallOfficialDegree": "J.D.",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0002-7899-3060",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "University of Utah",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/03r0ha626",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Nicholas",
            "overallOfficialLastName": "Evans",
            "overallOfficialDegree": "PhD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0002-3330-0224",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "University of Massachusetts Lowell",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/03hamhx47",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Samantha",
            "overallOfficialLastName": "Hurst",
            "overallOfficialDegree": "PhD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0001-9843-2845",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "University of California, San Diego",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/0168r3w48",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "T. Y. Alvin",
            "overallOfficialLastName": "Liu",
            "overallOfficialDegree": "MD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0003-2957-0755",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "Johns Hopkins University",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/00za53h95",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Gerald",
            "overallOfficialLastName": "McGwin",
            "overallOfficialDegree": "PhD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0001-9592-1133",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "University of Alabama at Birmingham",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/008s83205",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Shannon",
            "overallOfficialLastName": "McWeeney",
            "overallOfficialDegree": "PhD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0001-8333-6607",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "Oregon Health & Science University",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/009avj582",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Cynthia",
            "overallOfficialLastName": "Owsley",
            "overallOfficialDegree": "PhD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0003-3424-011X",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "University of Alabama at Birmingham",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/008s83205",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Bhavesh",
            "overallOfficialLastName": "Patel",
            "overallOfficialDegree": "PhD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0002-0307-262X",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "California Medical Innovations Institute",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/0156zyn36",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Michael",
            "overallOfficialLastName": "Snyder",
            "overallOfficialDegree": "PhD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0003-0784-7987",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "Stanford University",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/00f54p054",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Sara J.",
            "overallOfficialLastName": "Singer",
            "overallOfficialDegree": "MBA, PhD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "http://orcid.org/0000-0002-3374-1177",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "Stanford University",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/00f54p054",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          },
          {
            "overallOfficialFirstName": "Linda M.",
            "overallOfficialLastName": "Zangwill",
            "overallOfficialDegree": "PhD",
            "overallOfficialIdentifier": [
              {
                "overallOfficialIdentifierValue": "https://orcid.org/0000-0002-1143-5224",
                "overallOfficialIdentifierScheme": "ORCID",
                "schemeURI": "https://orcid.org"
              }
            ],
            "overallOfficialAffiliation": {
              "overallOfficialAffiliationName": "University of California, San Diego",
              "overallOfficialAffiliationIdentifier": {
                "overallOfficialAffiliationIdentifierValue": "https://ror.org/0168r3w48",
                "overallOfficialAffiliationIdentifierScheme": "ROR",
                "schemeURI": "https://ror.org"
              }
            },
            "overallOfficialRole": "Study Principal Investigator"
          }
        ],
        "locationList": [
          {
            "locationFacility": "University of Alabama at Birmingham",
            "locationStatus": "Enrolling by invitation",
            "locationCity": "Birmingham",
            "locationZip": "35233",
            "locationState": "Alabama",
            "locationCountry": "United States",
            "locationIdentifier": {
              "locationIdentifierValue": "https://ror.org/008s83205",
              "locationIdentifierScheme": "ROR",
              "schemeURI": "https://ror.org/"
            }
          },
          {
            "locationFacility": "University of California, San Diego",
            "locationStatus": "Enrolling by invitation",
            "locationCity": "San Diego",
            "locationZip": "92093",
            "locationState": "California",
            "locationCountry": "United States",
            "locationIdentifier": {
              "locationIdentifierValue": "https://ror.org/0168r3w48",
              "locationIdentifierScheme": "ROR",
              "schemeURI": "https://ror.org/"
            }
          },
          {
            "locationFacility": "University of Washington",
            "locationStatus": "Enrolling by invitation",
            "locationCity": "Seattle",
            "locationZip": "98109",
            "locationState": "Washington",
            "locationCountry": "United States",
            "locationIdentifier": {
              "locationIdentifierValue": "https://ror.org/00cvxb145",
              "locationIdentifierScheme": "ROR",
              "schemeURI": "https://ror.org/"
            }
          }
        ]
      }
    }
  },
  "study_id": "f5d4172f-b0fb-48c8-b545-0192d812b329",
  "study_title": "AI Ready and Exploratory Atlas for Diabetes Insights",
  "version_id": "504ba283-30fc-4a78-b2f9-857d3108a4ad",
  "version_title": "3.0.0",
  "versions": [
    {
      "id": "3",
      "title": "3.0.0",
      "createdAt": 1763366400,
      "doi": "10.60775/fairhub.3"
    },
    {
      "id": "2",
      "title": "2.0.0",
      "createdAt": 1731060000,
      "doi": "10.60775/fairhub.2"
    },
    {
      "id": "1",
      "title": "1.0.0",
      "createdAt": 1714762800,
      "doi": "10.60775/fairhub.1"
    }
  ]
}


================================================================================

FILE: aireadi_ro_crate_metadata_2026-08-12.txt
PATH: data/preprocessed/individual/AI_READI/aireadi_ro_crate_metadata_2026-08-12.txt
SIZE: 111757 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: AI_READI
Source ID: ro_crate_metadata
Source type: RO-Crate
Source URL: https://drive.google.com/uc?export=download&id=1appdz89WfHXkkXvJIv2Q_ziegNxdnhAF
Raw file: data/raw/AI_READI/aireadi_ro_crate_metadata_2026-08-12.json
--------------------------------------------------------------------------------
{
  "@context": {
    "EVI": "https://w3id.org/EVI#",
    "@vocab": "https://schema.org/"
  },
  "@graph": [
    {
      "@id": "ro-crate-metadata.json",
      "@type": "CreativeWork",
      "conformsTo": {
        "@id": "https://w3id.org/ro/crate/1.2-DRAFT"
      },
      "about": {
        "@id": "ark:59853/rocrate-b2ai-aireadi-release-3-0-0"
      },
      "fairscapeVersion": "1.2.1"
    },
    {
      "@id": "ark:59853/rocrate-b2ai-aireadi-release-3-0-0",
      "@type": [
        "https://w3id.org/EVI#Dataset",
        "https://w3id.org/EVI#ROCrate"
      ],
      "conformsTo": {
        "@id": "https://w3id.org/fairscape/profile/0.1"
      },
      "name": "Flagship Dataset of Type 2 Diabetes from the AI-READI Project",
      "description": "\nThe Artificial Intelligence Ready and Exploratory Atlas for Diabetes Insights (AI-READI) project seeks to create a flagship ethically-sourced dataset to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. The ability to understand and affect the course of complex, multi-organ diseases such as T2DM has been limited by a lack of well-designed, high quality, large, and inclusive multimodal datasets. The AI-READI team of investigators will aim to collect a cross-sectional dataset of 4,000 people and longitudinal data from 10% of the study cohort across the US. The study cohort will be balanced for self-reported race/ethnicity, gender, and diabetes disease stage. Data collection will be specifically designed to permit downstream pseudo-time manifold analysis, an approach used to predict disease trajectories by collecting and learning from complex, multimodal data from participants with differing disease severity (normal to insulin-dependent T2DM). The long-term objective for this project is to develop a foundational dataset in T2DM, agnostic to existing classification criteria or biases, which can be used to reconstruct a temporal atlas of T2DM development and reversal towards health (i.e., salutogenesis). Data will be optimized for downstream AI/ML research and made publicly available\n\nThis dataset contains data from 2280 participants that was collected between July 19, 2023 and May 01, 2025. Data from multiple modalities are included. A full list is provided in the Data Standards section below. The data in this dataset contain no protected health information (PHI). Information related to the sex and race/ethnicity of the participants as well as medication used has also been removed.\n\nThe dataset contains 356,343 files and is around 3.82 TB in size.\n\nA detailed description of the dataset is available in the AI-READI documentation for v3.0.0 of the dataset at docs.aireadi.org.\n",
      "keywords": [
        "diabetes mellitus",
        "Machine Learning",
        "Artificial Intelligence",
        "Electrocardiography",
        "Continuous Glucose Monitoring",
        "Retinal Imaging",
        "Eye Exam"
      ],
      "version": "3.0.0",
      "datePublished": "11/17/25",
      "isPartOf": [],
      "hasPart": [
        {
          "@id": "ark:59853/rocrate-b2ai-ai-readi-environmental-sensor"
        },
        {
          "@id": "ark:59853/rocrate-b2ai-ai-readi-flio"
        },
        {
          "@id": "ark:59853/rocrate-b2ai-ai-readi-ecg"
        },
        {
          "@id": "ark:59853/rocrate-b2ai-ai-readi-omop"
        },
        {
          "@id": "ark:59853/rocrate-b2ai-ai-readi-retinal-oct"
        },
        {
          "@id": "ark:59853/rocrate-b2ai-ai-readi-retinal-octa"
        },
        {
          "@id": "ark:59853/rocrate-b2ai-ai-readi-retinal-photography"
        },
        {
          "@id": "ark:59853/rocrate-b2ai-ai-readi-wearable-blood-glucose"
        },
        {
          "@id": "ark:59853/rocrate-b2ai-ai-readi-wearable-activity-monitor"
        }
      ],
      "author": [
        "AI-READI Consortium"
      ],
      "publisher": "AI-READI Consortium",
      "principalInvestigator": "Aaron Lee, Department of Ophthalmology, University of Washington",
      "funder": "NIH grant 1OT2OD032644 to the Bridge2AI: Salutogenesis Data Generation Project through the NIH Bridge2AI Common Fund program",
      "citation": "https://docs.aireadi.org",
      "associatedPublication": [
        "AI-READI Consortium. (2024). \"AI-READI: rethinking data collection, preparation and\nsharing for propelling AI-based discoveries in diabetes research and beyond.\"\nNature metabolism. https://doi.org/10.1038/s42255-024-01165-x",
        "AI-READI Consortium. (2025). Flagship Dataset of Type 2 Diabetes from the\nAI-READI Project (3.0.0) [Data set]. FAIRhub.\nhttps://doi.org/10.60775/fairhub.3"
      ],
      "identifier": "https://doi.org/10.60775/fairhub.3",
      "license": "https://doi.org/10.5281/zenodo.17555036",
      "conditionsOfAccess": "https://fairhub.io/datasets/3/access",
      "copyrightNotice": "Copyright © 2026 AI-READI",
      "contentSize": "3.82 TB",
      "ethicalReview": "Camille Nebeker, Debra Mathews, Kadija Ferryman, Nicholas Evans",
      "confidentialityLevel": "HL7:2N (normal)",
      "irb": {
        "@type": "Organization",
        "name": "Washington University IRB",
        "contactPoint": {
          "@type": "ContactPoint",
          "contactType": "IRB Reliance Team",
          "email": "hsdrely@uw.edu",
          "telephone": ""
        },
        "address": {
          "@type": "PostalAddress",
          "streetAddress": "Human Subjects Division University of Washington 4333 Brooklyn Ave NE Box 359470",
          "addressLocality": "Seattle",
          "addressRegion": "WA",
          "postalCode": "98195-9470",
          "addressCountry": "US"
        }
      },
      "irbProtocolId": "STUDY00016228",
      "humanSubjectExemption": "",
      "fdaRegulated": false,
      "deidentified": true,
      "humanSubjectResearch": "Yes",
      "dataGovernanceCommittee": "AI-READI Consortium",
      "rai:dataLimitations": "\nWhile the AI-READI's cross-sectional database ultimately aims to achieve balance across race/ethnicity, biological sex, and diabetes presence and severity, the pilot study is not balanced across these parameters.\n\nThree recording sites were strategically selected to achieve broad recruitment: the University of Alabama at Birmingham (UAB), the University of California San Diego (UCSD), and the University of Washington (UW). The sites were chosen for geographic variability across the United States and to ensure representation across various racial and ethnic groups. Individuals from all demographic backgrounds were recruited at all 3 sites. Factors influencing the generalization of derived models include the predominantly urban and hospital-based recruitment, which may not fully capture all possible cultural and socioeconomic backgrounds. The study cohort may not provide a comprehensive representation of the population, as it does not include other races/ethnicities such as Pacific Islanders and Native Americans. Information on device make and model, including specific modalities like macula scans or wide scans during OCT, were documented to ensure repeatability. Moreover, the study included multiple devices for one measure to enhance generalizability and represent the broad range of equipment utilized in clinical settings.\n\nIn cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.\n",
      "rai:dataBiases": "\nUniform data collection protocols were implemented for all subjects, irrespective of their race/ethnicity, biological sex, or diabetes severity, across all study sites. The selection of study sites was intended to ensure varied representation and minimize the potential for sampling bias.\n",
      "rai:dataUseCases": "The purpose for creating the dataset was to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. T2DM is a growing public health threat. Yet, the current understanding of T2DM, especially in the context of salutogenesis, is limited. Given the complexity of T2DM, AI-based approaches may help with improving our understanding but a key issue is the lack of data ready for training AI models. The AI-READI dataset is intended to fill this gap.",
      "rai:dataReleaseMaintenancePlan": "The dataset gets released as static versions. This is the third version of the dataset and consists of data collected up through the end of the second year of the study, i.e. between July 19, 2023 and May 1st, 2025. There are plans to release new versions of the dataset approximately once a year with additional data from participants who have been enrolled since the last dataset version release.",
      "rai:dataCollection": "Multiple modalities of data are collected for each participant, including survey data, clinical data, retinal imaging data, environmental sensor data, continuous glucose monitor data, and wearable activity monitor data. These encompass tabular data, imaging data, and physiological signal/waveform data. There is no unstructured text data included in this dataset. The exact forms used for data collection in REDCap are available here. Furthermore, all modalities, file formats, and devices are detailed in the dataset documentation at https://docs.aireadi.org/.",
      "rai:dataCollectionType": [
        "Manual Human Curation"
      ],
      "rai:dataCollectionMissingData": "Yes, not all modalities are available for all participants. Some participants elected not to participate in some study elements. In a few cases, the data collection device did not have any stored results or was returned too late to retrieve the results (e.g. battery died, data was lost). In a few cases, there may have been a data collision at some point in the process and data has been lost.",
      "rai:dataCollectionRawData": "Each instance consists of all of the data available for an individual participating in the study.",
      "rai:dataPreprocessingProtocol": [
        "\nThere were several quality control measures used at the time of data entry/acquisition. For example, clinical data outside of expected min/max ranges were flagged in REDCap, which was visible in reports viewed by clinical research coordinators (CRCs) and Data Managers. Using these REDCap reports as guides, Data Managers and CRCs examined participant records and determined if an error was likely. Data were checked for the following and edited if errors were detected:\n\n\ti. Credibility, based on range checks to determine if all responses fall within a prespecified reasonable range\n ii. Incorrect flow through prescribed skip patterns\n\tiii. Missing data that can be directly filed from other portions of an individual’ s record\n\tiv. The omission and/or duplication of records\n\nEditing was only done under the guidance and approval of the site PI. If corrected data was available from elsewhere in the respondent’s answers, the error was corrected. If there was no logical or appropriate way to correct the data, the Data site PI reviewed the values and made decisions about whether those values should be removed from the data.\n\nOnce data +were sent from each of the study sites to the central project team, additional processing steps were conducted in preparation for dissemination. For example, all data were mapped to standardized terminologies when possible, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model, a common data model for observational health data, and the Digital Imaging and Communications in Medicine (DICOM), a commonly used standard for medical imaging data. Details about the data processing approaches for each data domain/modality are described in the dataset documentation at https://docs.aireadi.org.\n"
      ],
      "rai:dataAnnotationProtocol": "N/A - no labels are provided",
      "rai:personalSensitiveInformation": [
        "EHR",
        "Wearable Monitoring",
        "ECG",
        "Environmental Sensor",
        "Continuous glucose monitor",
        "Wearable accelerometer"
      ],
      "completeness": "In cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.",
      "project_name": "AI READI",
      "ro-crate-metadata": "ro-crate-metadata.json",
      "contentUrl": "https://doi.org/10.60775/fairhub.3",
      "contact": "https://docs.aireadi.org/docs/3/contact"
    },
    {
      "@id": "ark:59853/rocrate-b2ai-ai-readi-ecg",
      "@type": [
        "https://w3id.org/EVI#Dataset",
        "https://w3id.org/EVI#ROCrate"
      ],
      "conformsTo": {
        "@id": "https://w3id.org/fairscape/profile/0.1"
      },
      "name": "AI-READI ECG Subcrate",
      "description": "The ECG Modality of the ROCrate",
      "keywords": [
        "diabetes mellitus",
        "Machine Learning",
        "Artificial Intelligence",
        "Electrocardiography",
        "Continuous Glucose Monitoring",
        "Retinal Imaging",
        "Eye Exam"
      ],
      "version": "3.0.0",
      "datePublished": "11/17/25",
      "isPartOf": [
        {
          "@id": "ark:59853/rocrate-b2ai-aireadi-release-3-0-0"
        }
      ],
      "hasPart": [],
      "author": [
        "AI-READI Consortium"
      ],
      "publisher": "AI-READI Consortium",
      "principalInvestigator": "Aaron Lee, Department of Ophthalmology, University of Washington",
      "funder": "NIH grant 1OT2OD032644 to the Bridge2AI: Salutogenesis Data Generation Project through the NIH Bridge2AI Common Fund program",
      "citation": "https://docs.aireadi.org",
      "associatedPublication": [
        "AI-READI Consortium. (2024). \"AI-READI: rethinking data collection, preparation and\nsharing for propelling AI-based discoveries in diabetes research and beyond.\"\nNature metabolism. https://doi.org/10.1038/s42255-024-01165-x",
        "AI-READI Consortium. (2025). Flagship Dataset of Type 2 Diabetes from the\nAI-READI Project (3.0.0) [Data set]. FAIRhub.\nhttps://doi.org/10.60775/fairhub.3"
      ],
      "identifier": "https://doi.org/10.60775/fairhub.3",
      "license": "https://doi.org/10.5281/zenodo.17555036",
      "conditionsOfAccess": "https://fairhub.io/datasets/3/access",
      "copyrightNotice": "Copyright © 2026 AI-READI",
      "ethicalReview": "Camille Nebeker, Debra Mathews, Kadija Ferryman, Nicholas Evans",
      "confidentialityLevel": "HL7:2N (normal)",
      "irb": {
        "@type": "Organization",
        "name": "Washington University IRB",
        "contactPoint": {
          "@type": "ContactPoint",
          "contactType": "IRB Reliance Team",
          "email": "hsdrely@uw.edu",
          "telephone": ""
        },
        "address": {
          "@type": "PostalAddress",
          "streetAddress": "Human Subjects Division University of Washington 4333 Brooklyn Ave NE Box 359470",
          "addressLocality": "Seattle",
          "addressRegion": "WA",
          "postalCode": "98195-9470",
          "addressCountry": "US"
        }
      },
      "irbProtocolId": "STUDY00016228",
      "humanSubjectExemption": "",
      "fdaRegulated": false,
      "deidentified": true,
      "humanSubjectResearch": "Yes",
      "dataGovernanceCommittee": "AI-READI Consortium",
      "rai:dataLimitations": "\nWhile the AI-READI's cross-sectional database ultimately aims to achieve balance across race/ethnicity, biological sex, and diabetes presence and severity, the pilot study is not balanced across these parameters.\n\nThree recording sites were strategically selected to achieve broad recruitment: the University of Alabama at Birmingham (UAB), the University of California San Diego (UCSD), and the University of Washington (UW). The sites were chosen for geographic variability across the United States and to ensure representation across various racial and ethnic groups. Individuals from all demographic backgrounds were recruited at all 3 sites. Factors influencing the generalization of derived models include the predominantly urban and hospital-based recruitment, which may not fully capture all possible cultural and socioeconomic backgrounds. The study cohort may not provide a comprehensive representation of the population, as it does not include other races/ethnicities such as Pacific Islanders and Native Americans. Information on device make and model, including specific modalities like macula scans or wide scans during OCT, were documented to ensure repeatability. Moreover, the study included multiple devices for one measure to enhance generalizability and represent the broad range of equipment utilized in clinical settings.\n\nIn cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.\n",
      "rai:dataBiases": "\nUniform data collection protocols were implemented for all subjects, irrespective of their race/ethnicity, biological sex, or diabetes severity, across all study sites. The selection of study sites was intended to ensure varied representation and minimize the potential for sampling bias.\n",
      "rai:dataUseCases": "The purpose for creating the dataset was to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. T2DM is a growing public health threat. Yet, the current understanding of T2DM, especially in the context of salutogenesis, is limited. Given the complexity of T2DM, AI-based approaches may help with improving our understanding but a key issue is the lack of data ready for training AI models. The AI-READI dataset is intended to fill this gap.",
      "rai:dataReleaseMaintenancePlan": "The dataset gets released as static versions. This is the third version of the dataset and consists of data collected up through the end of the second year of the study, i.e. between July 19, 2023 and May 1st, 2025. There are plans to release new versions of the dataset approximately once a year with additional data from participants who have been enrolled since the last dataset version release.",
      "rai:dataCollection": "Multiple modalities of data are collected for each participant, including survey data, clinical data, retinal imaging data, environmental sensor data, continuous glucose monitor data, and wearable activity monitor data. These encompass tabular data, imaging data, and physiological signal/waveform data. There is no unstructured text data included in this dataset. The exact forms used for data collection in REDCap are available here. Furthermore, all modalities, file formats, and devices are detailed in the dataset documentation at https://docs.aireadi.org/.",
      "rai:dataCollectionType": [
        "Manual Human Curation"
      ],
      "rai:dataCollectionMissingData": "Yes, not all modalities are available for all participants. Some participants elected not to participate in some study elements. In a few cases, the data collection device did not have any stored results or was returned too late to retrieve the results (e.g. battery died, data was lost). In a few cases, there may have been a data collision at some point in the process and data has been lost.",
      "rai:dataCollectionRawData": "Each instance consists of all of the data available for an individual participating in the study.",
      "rai:dataPreprocessingProtocol": [
        "\nThere were several quality control measures used at the time of data entry/acquisition. For example, clinical data outside of expected min/max ranges were flagged in REDCap, which was visible in reports viewed by clinical research coordinators (CRCs) and Data Managers. Using these REDCap reports as guides, Data Managers and CRCs examined participant records and determined if an error was likely. Data were checked for the following and edited if errors were detected:\n\n\ti. Credibility, based on range checks to determine if all responses fall within a prespecified reasonable range\n ii. Incorrect flow through prescribed skip patterns\n\tiii. Missing data that can be directly filed from other portions of an individual’ s record\n\tiv. The omission and/or duplication of records\n\nEditing was only done under the guidance and approval of the site PI. If corrected data was available from elsewhere in the respondent’s answers, the error was corrected. If there was no logical or appropriate way to correct the data, the Data site PI reviewed the values and made decisions about whether those values should be removed from the data.\n\nOnce data +were sent from each of the study sites to the central project team, additional processing steps were conducted in preparation for dissemination. For example, all data were mapped to standardized terminologies when possible, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model, a common data model for observational health data, and the Digital Imaging and Communications in Medicine (DICOM), a commonly used standard for medical imaging data. Details about the data processing approaches for each data domain/modality are described in the dataset documentation at https://docs.aireadi.org.\n"
      ],
      "rai:dataAnnotationProtocol": "N/A - no labels are provided",
      "rai:personalSensitiveInformation": [
        "EHR",
        "Wearable Monitoring",
        "ECG",
        "Environmental Sensor",
        "Continuous glucose monitor",
        "Wearable accelerometer"
      ],
      "completeness": "In cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.",
      "ro-crate-metadata": "cardiac_ecg/ro-crate-metadata.json",
      "contentUrl": "https://doi.org/10.60775/fairhub.3",
      "contact": "https://docs.aireadi.org/docs/3/contact"
    },
    {
      "@id": "ark:59853/rocrate-b2ai-ai-readi-omop",
      "@type": [
        "https://w3id.org/EVI#Dataset",
        "https://w3id.org/EVI#ROCrate"
      ],
      "conformsTo": {
        "@id": "https://w3id.org/fairscape/profile/0.1"
      },
      "name": "AI-READI EHR Subcrate",
      "description": "Clinical data in AI-READI consists of OMOP Common Data Model–formatted health-related measurements and assessments, including lab tests, MoCA cognitive testing, monofilament testing, physical and vision assessments, and survey questionnaire results.",
      "keywords": [
        "diabetes mellitus",
        "Machine Learning",
        "Artificial Intelligence",
        "Electrocardiography",
        "Continuous Glucose Monitoring",
        "Retinal Imaging",
        "Eye Exam"
      ],
      "version": "3.0.0",
      "datePublished": "11/17/25",
      "isPartOf": [
        {
          "@id": "ark:59853/rocrate-b2ai-aireadi-release-3-0-0"
        }
      ],
      "hasPart": [],
      "author": [
        "AI-READI Consortium"
      ],
      "publisher": "AI-READI Consortium",
      "principalInvestigator": "Aaron Lee, Department of Ophthalmology, University of Washington",
      "funder": "NIH grant 1OT2OD032644 to the Bridge2AI: Salutogenesis Data Generation Project through the NIH Bridge2AI Common Fund program",
      "citation": "https://docs.aireadi.org",
      "associatedPublication": [
        "AI-READI Consortium. (2024). \"AI-READI: rethinking data collection, preparation and\nsharing for propelling AI-based discoveries in diabetes research and beyond.\"\nNature metabolism. https://doi.org/10.1038/s42255-024-01165-x",
        "AI-READI Consortium. (2025). Flagship Dataset of Type 2 Diabetes from the\nAI-READI Project (3.0.0) [Data set]. FAIRhub.\nhttps://doi.org/10.60775/fairhub.3"
      ],
      "identifier": "https://doi.org/10.60775/fairhub.3",
      "license": "https://doi.org/10.5281/zenodo.17555036",
      "conditionsOfAccess": "https://fairhub.io/datasets/3/access",
      "copyrightNotice": "Copyright © 2026 AI-READI",
      "ethicalReview": "Camille Nebeker, Debra Mathews, Kadija Ferryman, Nicholas Evans",
      "confidentialityLevel": "HL7:2N (normal)",
      "irb": {
        "@type": "Organization",
        "name": "Washington University IRB",
        "contactPoint": {
          "@type": "ContactPoint",
          "contactType": "IRB Reliance Team",
          "email": "hsdrely@uw.edu",
          "telephone": ""
        },
        "address": {
          "@type": "PostalAddress",
          "streetAddress": "Human Subjects Division University of Washington 4333 Brooklyn Ave NE Box 359470",
          "addressLocality": "Seattle",
          "addressRegion": "WA",
          "postalCode": "98195-9470",
          "addressCountry": "US"
        }
      },
      "irbProtocolId": "STUDY00016228",
      "humanSubjectExemption": "",
      "fdaRegulated": false,
      "deidentified": true,
      "humanSubjectResearch": "Yes",
      "dataGovernanceCommittee": "AI-READI Consortium",
      "rai:dataLimitations": "\nWhile the AI-READI's cross-sectional database ultimately aims to achieve balance across race/ethnicity, biological sex, and diabetes presence and severity, the pilot study is not balanced across these parameters.\n\nThree recording sites were strategically selected to achieve broad recruitment: the University of Alabama at Birmingham (UAB), the University of California San Diego (UCSD), and the University of Washington (UW). The sites were chosen for geographic variability across the United States and to ensure representation across various racial and ethnic groups. Individuals from all demographic backgrounds were recruited at all 3 sites. Factors influencing the generalization of derived models include the predominantly urban and hospital-based recruitment, which may not fully capture all possible cultural and socioeconomic backgrounds. The study cohort may not provide a comprehensive representation of the population, as it does not include other races/ethnicities such as Pacific Islanders and Native Americans. Information on device make and model, including specific modalities like macula scans or wide scans during OCT, were documented to ensure repeatability. Moreover, the study included multiple devices for one measure to enhance generalizability and represent the broad range of equipment utilized in clinical settings.\n\nIn cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.\n",
      "rai:dataBiases": "\nUniform data collection protocols were implemented for all subjects, irrespective of their race/ethnicity, biological sex, or diabetes severity, across all study sites. The selection of study sites was intended to ensure varied representation and minimize the potential for sampling bias.\n",
      "rai:dataUseCases": "The purpose for creating the dataset was to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. T2DM is a growing public health threat. Yet, the current understanding of T2DM, especially in the context of salutogenesis, is limited. Given the complexity of T2DM, AI-based approaches may help with improving our understanding but a key issue is the lack of data ready for training AI models. The AI-READI dataset is intended to fill this gap.",
      "rai:dataReleaseMaintenancePlan": "The dataset gets released as static versions. This is the third version of the dataset and consists of data collected up through the end of the second year of the study, i.e. between July 19, 2023 and May 1st, 2025. There are plans to release new versions of the dataset approximately once a year with additional data from participants who have been enrolled since the last dataset version release.",
      "rai:dataCollection": "Multiple modalities of data are collected for each participant, including survey data, clinical data, retinal imaging data, environmental sensor data, continuous glucose monitor data, and wearable activity monitor data. These encompass tabular data, imaging data, and physiological signal/waveform data. There is no unstructured text data included in this dataset. The exact forms used for data collection in REDCap are available here. Furthermore, all modalities, file formats, and devices are detailed in the dataset documentation at https://docs.aireadi.org/.",
      "rai:dataCollectionType": [
        "Manual Human Curation"
      ],
      "rai:dataCollectionMissingData": "Yes, not all modalities are available for all participants. Some participants elected not to participate in some study elements. In a few cases, the data collection device did not have any stored results or was returned too late to retrieve the results (e.g. battery died, data was lost). In a few cases, there may have been a data collision at some point in the process and data has been lost.",
      "rai:dataCollectionRawData": "Each instance consists of all of the data available for an individual participating in the study.",
      "rai:dataPreprocessingProtocol": [
        "\nThere were several quality control measures used at the time of data entry/acquisition. For example, clinical data outside of expected min/max ranges were flagged in REDCap, which was visible in reports viewed by clinical research coordinators (CRCs) and Data Managers. Using these REDCap reports as guides, Data Managers and CRCs examined participant records and determined if an error was likely. Data were checked for the following and edited if errors were detected:\n\n\ti. Credibility, based on range checks to determine if all responses fall within a prespecified reasonable range\n ii. Incorrect flow through prescribed skip patterns\n\tiii. Missing data that can be directly filed from other portions of an individual’ s record\n\tiv. The omission and/or duplication of records\n\nEditing was only done under the guidance and approval of the site PI. If corrected data was available from elsewhere in the respondent’s answers, the error was corrected. If there was no logical or appropriate way to correct the data, the Data site PI reviewed the values and made decisions about whether those values should be removed from the data.\n\nOnce data +were sent from each of the study sites to the central project team, additional processing steps were conducted in preparation for dissemination. For example, all data were mapped to standardized terminologies when possible, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model, a common data model for observational health data, and the Digital Imaging and Communications in Medicine (DICOM), a commonly used standard for medical imaging data. Details about the data processing approaches for each data domain/modality are described in the dataset documentation at https://docs.aireadi.org.\n"
      ],
      "rai:dataAnnotationProtocol": "N/A - no labels are provided",
      "rai:personalSensitiveInformation": [
        "EHR",
        "Wearable Monitoring",
        "ECG",
        "Environmental Sensor",
        "Continuous glucose monitor",
        "Wearable accelerometer"
      ],
      "completeness": "In cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.",
      "ro-crate-metadata": "clinical_data/ro-crate-metadata.json",
      "contentUrl": "https://doi.org/10.60775/fairhub.3",
      "contact": "https://docs.aireadi.org/docs/3/contact"
    },
    {
      "@id": "ark:59853/rocrate-b2ai-ai-readi-flio",
      "@type": [
        "https://w3id.org/EVI#Dataset",
        "https://w3id.org/EVI#ROCrate"
      ],
      "conformsTo": {
        "@id": "https://w3id.org/fairscape/profile/0.1"
      },
      "name": "AI-READI FLIO Subcrate",
      "description": "\nFLIO stands for Fluorescence Lifetime Imaging Ophthalmoscopy, a technology used to capture and analyze fluorescence behavior in the eye over time. It is based on a concept called fluorescence decay, which refers to the process by which fluorescent molecules lose their excitation energy and return to their ground state, emitting light in the form of fluorescence. When a fluorescent molecule absorbs light energy, it becomes excited and emits light of a longer wavelength as it returns to its ground state. The time it takes for this emission to occur, along with the intensity of the emitted light, provides valuable information about the molecular environment and dynamics. Understanding fluorescence decay is crucial for interpreting FLIO data, as it informs about the lifetime and behavior of fluorescent signals within the eye.\n\nDuring the imaging, the FLIO device exposes retinal tissue to intermittent laser beams and records the timing and number of photons for each point after each laser pulse. This information forms a picture of how fluorescence behaves over time in different parts of the eye. When analyzing FLIO data, it is important to understand that each pixel in the image contains a curve showing how many photons arrived at different times. This photon data follows Poisson distribution. The signal-to-noise ratio varies for different aspects of the data. Additionally, unwanted signals, like background light or fluorescence from the lens, can not simply be subtracted because they are part of the same distribution as the useful data. Our FLIO scan records images using two wavelengths, short (498-560nm) and long (560-720nm), which are provided in two different .dcm files for each eye in the dataset. These wavelengths provide different penetrance into the retinal tissue and usage of short- or long-wavelength .dcm file can be based on the region and layer of interest. A common approach is analyzing each of these files separately for each subject.\n\nTo get the best FLIO data, it is crucial to maximize the number of photons recorded by ensuring proper focus and avoiding unwanted signals like background light or lens fluorescence. Keeping a consistent distance between the scanner and the patient during data collection is also important for accurate results. These are particularly important when working with the data because the device is very sensitive, and several factors can induce noise in the images. The device can not handle more than 1 centimeter of movement during image acquisition.\n",
      "keywords": [
        "diabetes mellitus",
        "Machine Learning",
        "Artificial Intelligence",
        "Electrocardiography",
        "Continuous Glucose Monitoring",
        "Retinal Imaging",
        "Eye Exam"
      ],
      "version": "3.0.0",
      "datePublished": "11/17/25",
      "isPartOf": [
        {
          "@id": "ark:59853/rocrate-b2ai-aireadi-release-3-0-0"
        }
      ],
      "hasPart": [],
      "author": [
        "AI-READI Consortium"
      ],
      "publisher": "AI-READI Consortium",
      "principalInvestigator": "Aaron Lee, Department of Ophthalmology, University of Washington",
      "funder": "NIH grant 1OT2OD032644 to the Bridge2AI: Salutogenesis Data Generation Project through the NIH Bridge2AI Common Fund program",
      "citation": "https://docs.aireadi.org",
      "associatedPublication": [
        "AI-READI Consortium. (2024). \"AI-READI: rethinking data collection, preparation and\nsharing for propelling AI-based discoveries in diabetes research and beyond.\"\nNature metabolism. https://doi.org/10.1038/s42255-024-01165-x",
        "AI-READI Consortium. (2025). Flagship Dataset of Type 2 Diabetes from the\nAI-READI Project (3.0.0) [Data set]. FAIRhub.\nhttps://doi.org/10.60775/fairhub.3"
      ],
      "identifier": "https://doi.org/10.60775/fairhub.3",
      "license": "https://doi.org/10.5281/zenodo.17555036",
      "conditionsOfAccess": "https://fairhub.io/datasets/3/access",
      "copyrightNotice": "Copyright © 2026 AI-READI",
      "ethicalReview": "Camille Nebeker, Debra Mathews, Kadija Ferryman, Nicholas Evans",
      "confidentialityLevel": "HL7:2N (normal)",
      "irb": {
        "@type": "Organization",
        "name": "Washington University IRB",
        "contactPoint": {
          "@type": "ContactPoint",
          "contactType": "IRB Reliance Team",
          "email": "hsdrely@uw.edu",
          "telephone": ""
        },
        "address": {
          "@type": "PostalAddress",
          "streetAddress": "Human Subjects Division University of Washington 4333 Brooklyn Ave NE Box 359470",
          "addressLocality": "Seattle",
          "addressRegion": "WA",
          "postalCode": "98195-9470",
          "addressCountry": "US"
        }
      },
      "irbProtocolId": "STUDY00016228",
      "humanSubjectExemption": "",
      "fdaRegulated": false,
      "deidentified": true,
      "humanSubjectResearch": "Yes",
      "dataGovernanceCommittee": "AI-READI Consortium",
      "rai:dataLimitations": "\nWhile the AI-READI's cross-sectional database ultimately aims to achieve balance across race/ethnicity, biological sex, and diabetes presence and severity, the pilot study is not balanced across these parameters.\n\nThree recording sites were strategically selected to achieve broad recruitment: the University of Alabama at Birmingham (UAB), the University of California San Diego (UCSD), and the University of Washington (UW). The sites were chosen for geographic variability across the United States and to ensure representation across various racial and ethnic groups. Individuals from all demographic backgrounds were recruited at all 3 sites. Factors influencing the generalization of derived models include the predominantly urban and hospital-based recruitment, which may not fully capture all possible cultural and socioeconomic backgrounds. The study cohort may not provide a comprehensive representation of the population, as it does not include other races/ethnicities such as Pacific Islanders and Native Americans. Information on device make and model, including specific modalities like macula scans or wide scans during OCT, were documented to ensure repeatability. Moreover, the study included multiple devices for one measure to enhance generalizability and represent the broad range of equipment utilized in clinical settings.\n\nIn cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.\n",
      "rai:dataBiases": "\nUniform data collection protocols were implemented for all subjects, irrespective of their race/ethnicity, biological sex, or diabetes severity, across all study sites. The selection of study sites was intended to ensure varied representation and minimize the potential for sampling bias.\n",
      "rai:dataUseCases": "The purpose for creating the dataset was to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. T2DM is a growing public health threat. Yet, the current understanding of T2DM, especially in the context of salutogenesis, is limited. Given the complexity of T2DM, AI-based approaches may help with improving our understanding but a key issue is the lack of data ready for training AI models. The AI-READI dataset is intended to fill this gap.",
      "rai:dataReleaseMaintenancePlan": "The dataset gets released as static versions. This is the third version of the dataset and consists of data collected up through the end of the second year of the study, i.e. between July 19, 2023 and May 1st, 2025. There are plans to release new versions of the dataset approximately once a year with additional data from participants who have been enrolled since the last dataset version release.",
      "rai:dataCollection": "Multiple modalities of data are collected for each participant, including survey data, clinical data, retinal imaging data, environmental sensor data, continuous glucose monitor data, and wearable activity monitor data. These encompass tabular data, imaging data, and physiological signal/waveform data. There is no unstructured text data included in this dataset. The exact forms used for data collection in REDCap are available here. Furthermore, all modalities, file formats, and devices are detailed in the dataset documentation at https://docs.aireadi.org/.",
      "rai:dataCollectionType": [
        "Manual Human Curation"
      ],
      "rai:dataCollectionMissingData": "Yes, not all modalities are available for all participants. Some participants elected not to participate in some study elements. In a few cases, the data collection device did not have any stored results or was returned too late to retrieve the results (e.g. battery died, data was lost). In a few cases, there may have been a data collision at some point in the process and data has been lost.",
      "rai:dataCollectionRawData": "Each instance consists of all of the data available for an individual participating in the study.",
      "rai:dataPreprocessingProtocol": [
        "\nThere were several quality control measures used at the time of data entry/acquisition. For example, clinical data outside of expected min/max ranges were flagged in REDCap, which was visible in reports viewed by clinical research coordinators (CRCs) and Data Managers. Using these REDCap reports as guides, Data Managers and CRCs examined participant records and determined if an error was likely. Data were checked for the following and edited if errors were detected:\n\n\ti. Credibility, based on range checks to determine if all responses fall within a prespecified reasonable range\n ii. Incorrect flow through prescribed skip patterns\n\tiii. Missing data that can be directly filed from other portions of an individual’ s record\n\tiv. The omission and/or duplication of records\n\nEditing was only done under the guidance and approval of the site PI. If corrected data was available from elsewhere in the respondent’s answers, the error was corrected. If there was no logical or appropriate way to correct the data, the Data site PI reviewed the values and made decisions about whether those values should be removed from the data.\n\nOnce data +were sent from each of the study sites to the central project team, additional processing steps were conducted in preparation for dissemination. For example, all data were mapped to standardized terminologies when possible, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model, a common data model for observational health data, and the Digital Imaging and Communications in Medicine (DICOM), a commonly used standard for medical imaging data. Details about the data processing approaches for each data domain/modality are described in the dataset documentation at https://docs.aireadi.org.\n"
      ],
      "rai:dataAnnotationProtocol": "N/A - no labels are provided",
      "rai:personalSensitiveInformation": [
        "EHR",
        "Wearable Monitoring",
        "ECG",
        "Environmental Sensor",
        "Continuous glucose monitor",
        "Wearable accelerometer"
      ],
      "completeness": "In cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.",
      "ro-crate-metadata": "retinal_flio/ro-crate-metadata.json",
      "contentUrl": "https://doi.org/10.60775/fairhub.3",
      "contact": "https://docs.aireadi.org/docs/3/contact"
    },
    {
      "@id": "ark:59853/rocrate-b2ai-ai-readi-retinal-oct",
      "@type": [
        "https://w3id.org/EVI#Dataset",
        "https://w3id.org/EVI#ROCrate"
      ],
      "conformsTo": {
        "@id": "https://w3id.org/fairscape/profile/0.1"
      },
      "name": "AI-READI Retinal OCT Subcrate",
      "description": "Optical Coherence Tomography (OCT) is a non-invasive diagnostic technique that renders an in vivo cross-sectional view of the retina. OCT utilizes a concept known as interferometry to create a cross-sectional map of the retina that is accurate to within at least 10-15 microns. This technique provides detailed, cross-sectional images of the various layers of the retina with high resolution, enabling clinicians to diagnose and monitor a wide range of retinal diseases and conditions.",
      "keywords": [
        "diabetes mellitus",
        "Machine Learning",
        "Artificial Intelligence",
        "Electrocardiography",
        "Continuous Glucose Monitoring",
        "Retinal Imaging",
        "Eye Exam"
      ],
      "version": "3.0.0",
      "datePublished": "11/17/25",
      "isPartOf": [
        {
          "@id": "ark:59853/rocrate-b2ai-aireadi-release-3-0-0"
        }
      ],
      "hasPart": [],
      "author": [
        "AI-READI Consortium"
      ],
      "publisher": "AI-READI Consortium",
      "principalInvestigator": "Aaron Lee, Department of Ophthalmology, University of Washington",
      "funder": "NIH grant 1OT2OD032644 to the Bridge2AI: Salutogenesis Data Generation Project through the NIH Bridge2AI Common Fund program",
      "citation": "https://docs.aireadi.org",
      "associatedPublication": [
        "AI-READI Consortium. (2024). \"AI-READI: rethinking data collection, preparation and\nsharing for propelling AI-based discoveries in diabetes research and beyond.\"\nNature metabolism. https://doi.org/10.1038/s42255-024-01165-x",
        "AI-READI Consortium. (2025). Flagship Dataset of Type 2 Diabetes from the\nAI-READI Project (3.0.0) [Data set]. FAIRhub.\nhttps://doi.org/10.60775/fairhub.3"
      ],
      "identifier": "https://doi.org/10.60775/fairhub.3",
      "license": "https://doi.org/10.5281/zenodo.17555036",
      "conditionsOfAccess": "https://fairhub.io/datasets/3/access",
      "copyrightNotice": "Copyright © 2026 AI-READI",
      "ethicalReview": "Camille Nebeker, Debra Mathews, Kadija Ferryman, Nicholas Evans",
      "confidentialityLevel": "HL7:2N (normal)",
      "irb": {
        "@type": "Organization",
        "name": "Washington University IRB",
        "contactPoint": {
          "@type": "ContactPoint",
          "contactType": "IRB Reliance Team",
          "email": "hsdrely@uw.edu",
          "telephone": ""
        },
        "address": {
          "@type": "PostalAddress",
          "streetAddress": "Human Subjects Division University of Washington 4333 Brooklyn Ave NE Box 359470",
          "addressLocality": "Seattle",
          "addressRegion": "WA",
          "postalCode": "98195-9470",
          "addressCountry": "US"
        }
      },
      "irbProtocolId": "STUDY00016228",
      "humanSubjectExemption": "",
      "fdaRegulated": false,
      "deidentified": true,
      "humanSubjectResearch": "Yes",
      "dataGovernanceCommittee": "AI-READI Consortium",
      "rai:dataLimitations": "\nWhile the AI-READI's cross-sectional database ultimately aims to achieve balance across race/ethnicity, biological sex, and diabetes presence and severity, the pilot study is not balanced across these parameters.\n\nThree recording sites were strategically selected to achieve broad recruitment: the University of Alabama at Birmingham (UAB), the University of California San Diego (UCSD), and the University of Washington (UW). The sites were chosen for geographic variability across the United States and to ensure representation across various racial and ethnic groups. Individuals from all demographic backgrounds were recruited at all 3 sites. Factors influencing the generalization of derived models include the predominantly urban and hospital-based recruitment, which may not fully capture all possible cultural and socioeconomic backgrounds. The study cohort may not provide a comprehensive representation of the population, as it does not include other races/ethnicities such as Pacific Islanders and Native Americans. Information on device make and model, including specific modalities like macula scans or wide scans during OCT, were documented to ensure repeatability. Moreover, the study included multiple devices for one measure to enhance generalizability and represent the broad range of equipment utilized in clinical settings.\n\nIn cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.\n",
      "rai:dataBiases": "\nUniform data collection protocols were implemented for all subjects, irrespective of their race/ethnicity, biological sex, or diabetes severity, across all study sites. The selection of study sites was intended to ensure varied representation and minimize the potential for sampling bias.\n",
      "rai:dataUseCases": "The purpose for creating the dataset was to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. T2DM is a growing public health threat. Yet, the current understanding of T2DM, especially in the context of salutogenesis, is limited. Given the complexity of T2DM, AI-based approaches may help with improving our understanding but a key issue is the lack of data ready for training AI models. The AI-READI dataset is intended to fill this gap.",
      "rai:dataReleaseMaintenancePlan": "The dataset gets released as static versions. This is the third version of the dataset and consists of data collected up through the end of the second year of the study, i.e. between July 19, 2023 and May 1st, 2025. There are plans to release new versions of the dataset approximately once a year with additional data from participants who have been enrolled since the last dataset version release.",
      "rai:dataCollection": "Multiple modalities of data are collected for each participant, including survey data, clinical data, retinal imaging data, environmental sensor data, continuous glucose monitor data, and wearable activity monitor data. These encompass tabular data, imaging data, and physiological signal/waveform data. There is no unstructured text data included in this dataset. The exact forms used for data collection in REDCap are available here. Furthermore, all modalities, file formats, and devices are detailed in the dataset documentation at https://docs.aireadi.org/.",
      "rai:dataCollectionType": [
        "Manual Human Curation"
      ],
      "rai:dataCollectionMissingData": "Yes, not all modalities are available for all participants. Some participants elected not to participate in some study elements. In a few cases, the data collection device did not have any stored results or was returned too late to retrieve the results (e.g. battery died, data was lost). In a few cases, there may have been a data collision at some point in the process and data has been lost.",
      "rai:dataCollectionRawData": "Each instance consists of all of the data available for an individual participating in the study.",
      "rai:dataPreprocessingProtocol": [
        "\nThere were several quality control measures used at the time of data entry/acquisition. For example, clinical data outside of expected min/max ranges were flagged in REDCap, which was visible in reports viewed by clinical research coordinators (CRCs) and Data Managers. Using these REDCap reports as guides, Data Managers and CRCs examined participant records and determined if an error was likely. Data were checked for the following and edited if errors were detected:\n\n\ti. Credibility, based on range checks to determine if all responses fall within a prespecified reasonable range\n ii. Incorrect flow through prescribed skip patterns\n\tiii. Missing data that can be directly filed from other portions of an individual’ s record\n\tiv. The omission and/or duplication of records\n\nEditing was only done under the guidance and approval of the site PI. If corrected data was available from elsewhere in the respondent’s answers, the error was corrected. If there was no logical or appropriate way to correct the data, the Data site PI reviewed the values and made decisions about whether those values should be removed from the data.\n\nOnce data +were sent from each of the study sites to the central project team, additional processing steps were conducted in preparation for dissemination. For example, all data were mapped to standardized terminologies when possible, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model, a common data model for observational health data, and the Digital Imaging and Communications in Medicine (DICOM), a commonly used standard for medical imaging data. Details about the data processing approaches for each data domain/modality are described in the dataset documentation at https://docs.aireadi.org.\n"
      ],
      "rai:dataAnnotationProtocol": "N/A - no labels are provided",
      "rai:personalSensitiveInformation": [
        "EHR",
        "Wearable Monitoring",
        "ECG",
        "Environmental Sensor",
        "Continuous glucose monitor",
        "Wearable accelerometer"
      ],
      "completeness": "In cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.",
      "ro-crate-metadata": "retinal_oct/ro-crate-metadata.json",
      "contentUrl": "https://doi.org/10.60775/fairhub.3",
      "contact": "https://docs.aireadi.org/docs/3/contact"
    },
    {
      "@id": "ark:59853/rocrate-b2ai-ai-readi-retinal-octa",
      "@type": [
        "https://w3id.org/EVI#Dataset",
        "https://w3id.org/EVI#ROCrate"
      ],
      "conformsTo": {
        "@id": "https://w3id.org/fairscape/profile/0.1"
      },
      "name": "AI-READI Retinal OCTA Subcrate",
      "description": "\nOptical coherence tomography angiography (OCTA) is a non-invasive imaging modality that permits visualization of retinal and inner choroidal circulation without the need for the dye injection, based on the principle of mapping red blood cell movement over time by comparing sequential OCT sagittal scans at a given cross-section.\n\nOCT-A technology uses laser light reflectance of the surface of moving red blood cells to accurately depict vessels through different segmented areas of the eye, thus eliminating the need for intravascular dyes. The OCT scan of a patient's retina consists of multiple individual A-scans, which when compiled into a B-scan provides cross-sectional structural information. With OCT-A technology, the same tissue area is repeatedly imaged, and differences are analyzed between scans (over time), thus allowing one to detect zones containing high flow rates (i.e. with marked changes between scans) and zones with slower, or no flow at all, which will be similar among scans. The main advantages are the shorter acquisition time and that it is a non-invasive process. Fluorescein and indocyanine-green angiography require an injectable dye (which takes time to reach retinal vessels, and may be associated with systemic adverse effects and even anaphylactic reactions. One asset of this OCT-based approach is that it provides a quantitative analysis of the retinal vessels (in addition to the qualitative analysis done on standard angiography). Moreover, and contrary to the \"2-D\" conventional angiograms, OCT-A technology provides \"3-D\" imaging information of the macula and visualizes peripapillary capillaries that supply the retinal nerve fiber layer.\n",
      "keywords": [
        "diabetes mellitus",
        "Machine Learning",
        "Artificial Intelligence",
        "Electrocardiography",
        "Continuous Glucose Monitoring",
        "Retinal Imaging",
        "Eye Exam"
      ],
      "version": "3.0.0",
      "datePublished": "11/17/25",
      "isPartOf": [
        {
          "@id": "ark:59853/rocrate-b2ai-aireadi-release-3-0-0"
        }
      ],
      "hasPart": [],
      "author": [
        "AI-READI Consortium"
      ],
      "publisher": "AI-READI Consortium",
      "principalInvestigator": "Aaron Lee, Department of Ophthalmology, University of Washington",
      "funder": "NIH grant 1OT2OD032644 to the Bridge2AI: Salutogenesis Data Generation Project through the NIH Bridge2AI Common Fund program",
      "citation": "https://docs.aireadi.org",
      "associatedPublication": [
        "AI-READI Consortium. (2024). \"AI-READI: rethinking data collection, preparation and\nsharing for propelling AI-based discoveries in diabetes research and beyond.\"\nNature metabolism. https://doi.org/10.1038/s42255-024-01165-x",
        "AI-READI Consortium. (2025). Flagship Dataset of Type 2 Diabetes from the\nAI-READI Project (3.0.0) [Data set]. FAIRhub.\nhttps://doi.org/10.60775/fairhub.3"
      ],
      "identifier": "https://doi.org/10.60775/fairhub.3",
      "license": "https://doi.org/10.5281/zenodo.17555036",
      "conditionsOfAccess": "https://fairhub.io/datasets/3/access",
      "copyrightNotice": "Copyright © 2026 AI-READI",
      "ethicalReview": "Camille Nebeker, Debra Mathews, Kadija Ferryman, Nicholas Evans",
      "confidentialityLevel": "HL7:2N (normal)",
      "irb": {
        "@type": "Organization",
        "name": "Washington University IRB",
        "contactPoint": {
          "@type": "ContactPoint",
          "contactType": "IRB Reliance Team",
          "email": "hsdrely@uw.edu",
          "telephone": ""
        },
        "address": {
          "@type": "PostalAddress",
          "streetAddress": "Human Subjects Division University of Washington 4333 Brooklyn Ave NE Box 359470",
          "addressLocality": "Seattle",
          "addressRegion": "WA",
          "postalCode": "98195-9470",
          "addressCountry": "US"
        }
      },
      "irbProtocolId": "STUDY00016228",
      "humanSubjectExemption": "",
      "fdaRegulated": false,
      "deidentified": true,
      "humanSubjectResearch": "Yes",
      "dataGovernanceCommittee": "AI-READI Consortium",
      "rai:dataLimitations": "\nWhile the AI-READI's cross-sectional database ultimately aims to achieve balance across race/ethnicity, biological sex, and diabetes presence and severity, the pilot study is not balanced across these parameters.\n\nThree recording sites were strategically selected to achieve broad recruitment: the University of Alabama at Birmingham (UAB), the University of California San Diego (UCSD), and the University of Washington (UW). The sites were chosen for geographic variability across the United States and to ensure representation across various racial and ethnic groups. Individuals from all demographic backgrounds were recruited at all 3 sites. Factors influencing the generalization of derived models include the predominantly urban and hospital-based recruitment, which may not fully capture all possible cultural and socioeconomic backgrounds. The study cohort may not provide a comprehensive representation of the population, as it does not include other races/ethnicities such as Pacific Islanders and Native Americans. Information on device make and model, including specific modalities like macula scans or wide scans during OCT, were documented to ensure repeatability. Moreover, the study included multiple devices for one measure to enhance generalizability and represent the broad range of equipment utilized in clinical settings.\n\nIn cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.\n",
      "rai:dataBiases": "\nUniform data collection protocols were implemented for all subjects, irrespective of their race/ethnicity, biological sex, or diabetes severity, across all study sites. The selection of study sites was intended to ensure varied representation and minimize the potential for sampling bias.\n",
      "rai:dataUseCases": "The purpose for creating the dataset was to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. T2DM is a growing public health threat. Yet, the current understanding of T2DM, especially in the context of salutogenesis, is limited. Given the complexity of T2DM, AI-based approaches may help with improving our understanding but a key issue is the lack of data ready for training AI models. The AI-READI dataset is intended to fill this gap.",
      "rai:dataReleaseMaintenancePlan": "The dataset gets released as static versions. This is the third version of the dataset and consists of data collected up through the end of the second year of the study, i.e. between July 19, 2023 and May 1st, 2025. There are plans to release new versions of the dataset approximately once a year with additional data from participants who have been enrolled since the last dataset version release.",
      "rai:dataCollection": "Multiple modalities of data are collected for each participant, including survey data, clinical data, retinal imaging data, environmental sensor data, continuous glucose monitor data, and wearable activity monitor data. These encompass tabular data, imaging data, and physiological signal/waveform data. There is no unstructured text data included in this dataset. The exact forms used for data collection in REDCap are available here. Furthermore, all modalities, file formats, and devices are detailed in the dataset documentation at https://docs.aireadi.org/.",
      "rai:dataCollectionType": [
        "Manual Human Curation"
      ],
      "rai:dataCollectionMissingData": "Yes, not all modalities are available for all participants. Some participants elected not to participate in some study elements. In a few cases, the data collection device did not have any stored results or was returned too late to retrieve the results (e.g. battery died, data was lost). In a few cases, there may have been a data collision at some point in the process and data has been lost.",
      "rai:dataCollectionRawData": "Each instance consists of all of the data available for an individual participating in the study.",
      "rai:dataPreprocessingProtocol": [
        "\nThere were several quality control measures used at the time of data entry/acquisition. For example, clinical data outside of expected min/max ranges were flagged in REDCap, which was visible in reports viewed by clinical research coordinators (CRCs) and Data Managers. Using these REDCap reports as guides, Data Managers and CRCs examined participant records and determined if an error was likely. Data were checked for the following and edited if errors were detected:\n\n\ti. Credibility, based on range checks to determine if all responses fall within a prespecified reasonable range\n ii. Incorrect flow through prescribed skip patterns\n\tiii. Missing data that can be directly filed from other portions of an individual’ s record\n\tiv. The omission and/or duplication of records\n\nEditing was only done under the guidance and approval of the site PI. If corrected data was available from elsewhere in the respondent’s answers, the error was corrected. If there was no logical or appropriate way to correct the data, the Data site PI reviewed the values and made decisions about whether those values should be removed from the data.\n\nOnce data +were sent from each of the study sites to the central project team, additional processing steps were conducted in preparation for dissemination. For example, all data were mapped to standardized terminologies when possible, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model, a common data model for observational health data, and the Digital Imaging and Communications in Medicine (DICOM), a commonly used standard for medical imaging data. Details about the data processing approaches for each data domain/modality are described in the dataset documentation at https://docs.aireadi.org.\n"
      ],
      "rai:dataAnnotationProtocol": "N/A - no labels are provided",
      "rai:personalSensitiveInformation": [
        "EHR",
        "Wearable Monitoring",
        "ECG",
        "Environmental Sensor",
        "Continuous glucose monitor",
        "Wearable accelerometer"
      ],
      "completeness": "In cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.",
      "ro-crate-metadata": "retinal_octa\\ro-crate-metadata.json",
      "contentUrl": "https://doi.org/10.60775/fairhub.3",
      "contact": "https://docs.aireadi.org/docs/3/contact"
    },
    {
      "@id": "ark:59853/rocrate-b2ai-ai-readi-retinal-photography",
      "@type": [
        "https://w3id.org/EVI#Dataset",
        "https://w3id.org/EVI#ROCrate"
      ],
      "conformsTo": {
        "@id": "https://w3id.org/fairscape/profile/0.1"
      },
      "name": "AI-READI Retinal Photography Subcrate",
      "description": "Retinal photography is a non-invasive procedure that photographs the posterior segment of an eye, also known as the fundus. It provides two-dimensional images of the fundus and can be performed with different filters of different wavelengths when an eye is undilated or dilated. The main structures that can be visualized inon a fundus photo are the optic nerve, macula, retinal vasculature, and central and peripheral retina. Fundus photography is used to record the condition of these structures to document the presence of abnormalities and monitor the changes over time.",
      "keywords": [
        "diabetes mellitus",
        "Machine Learning",
        "Artificial Intelligence",
        "Electrocardiography",
        "Continuous Glucose Monitoring",
        "Retinal Imaging",
        "Eye Exam"
      ],
      "version": "3.0.0",
      "datePublished": "11/17/25",
      "isPartOf": [
        {
          "@id": "ark:59853/rocrate-b2ai-aireadi-release-3-0-0"
        }
      ],
      "hasPart": [],
      "author": [
        "AI-READI Consortium"
      ],
      "publisher": "AI-READI Consortium",
      "principalInvestigator": "Aaron Lee, Department of Ophthalmology, University of Washington",
      "funder": "NIH grant 1OT2OD032644 to the Bridge2AI: Salutogenesis Data Generation Project through the NIH Bridge2AI Common Fund program",
      "citation": "https://docs.aireadi.org",
      "associatedPublication": [
        "AI-READI Consortium. (2024). \"AI-READI: rethinking data collection, preparation and\nsharing for propelling AI-based discoveries in diabetes research and beyond.\"\nNature metabolism. https://doi.org/10.1038/s42255-024-01165-x",
        "AI-READI Consortium. (2025). Flagship Dataset of Type 2 Diabetes from the\nAI-READI Project (3.0.0) [Data set]. FAIRhub.\nhttps://doi.org/10.60775/fairhub.3"
      ],
      "identifier": "https://doi.org/10.60775/fairhub.3",
      "license": "https://doi.org/10.5281/zenodo.17555036",
      "conditionsOfAccess": "https://fairhub.io/datasets/3/access",
      "copyrightNotice": "Copyright © 2026 AI-READI",
      "ethicalReview": "Camille Nebeker, Debra Mathews, Kadija Ferryman, Nicholas Evans",
      "confidentialityLevel": "HL7:2N (normal)",
      "irb": {
        "@type": "Organization",
        "name": "Washington University IRB",
        "contactPoint": {
          "@type": "ContactPoint",
          "contactType": "IRB Reliance Team",
          "email": "hsdrely@uw.edu",
          "telephone": ""
        },
        "address": {
          "@type": "PostalAddress",
          "streetAddress": "Human Subjects Division University of Washington 4333 Brooklyn Ave NE Box 359470",
          "addressLocality": "Seattle",
          "addressRegion": "WA",
          "postalCode": "98195-9470",
          "addressCountry": "US"
        }
      },
      "irbProtocolId": "STUDY00016228",
      "humanSubjectExemption": "",
      "fdaRegulated": false,
      "deidentified": true,
      "humanSubjectResearch": "Yes",
      "dataGovernanceCommittee": "AI-READI Consortium",
      "rai:dataLimitations": "\nWhile the AI-READI's cross-sectional database ultimately aims to achieve balance across race/ethnicity, biological sex, and diabetes presence and severity, the pilot study is not balanced across these parameters.\n\nThree recording sites were strategically selected to achieve broad recruitment: the University of Alabama at Birmingham (UAB), the University of California San Diego (UCSD), and the University of Washington (UW). The sites were chosen for geographic variability across the United States and to ensure representation across various racial and ethnic groups. Individuals from all demographic backgrounds were recruited at all 3 sites. Factors influencing the generalization of derived models include the predominantly urban and hospital-based recruitment, which may not fully capture all possible cultural and socioeconomic backgrounds. The study cohort may not provide a comprehensive representation of the population, as it does not include other races/ethnicities such as Pacific Islanders and Native Americans. Information on device make and model, including specific modalities like macula scans or wide scans during OCT, were documented to ensure repeatability. Moreover, the study included multiple devices for one measure to enhance generalizability and represent the broad range of equipment utilized in clinical settings.\n\nIn cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.\n",
      "rai:dataBiases": "\nUniform data collection protocols were implemented for all subjects, irrespective of their race/ethnicity, biological sex, or diabetes severity, across all study sites. The selection of study sites was intended to ensure varied representation and minimize the potential for sampling bias.\n",
      "rai:dataUseCases": "The purpose for creating the dataset was to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. T2DM is a growing public health threat. Yet, the current understanding of T2DM, especially in the context of salutogenesis, is limited. Given the complexity of T2DM, AI-based approaches may help with improving our understanding but a key issue is the lack of data ready for training AI models. The AI-READI dataset is intended to fill this gap.",
      "rai:dataReleaseMaintenancePlan": "The dataset gets released as static versions. This is the third version of the dataset and consists of data collected up through the end of the second year of the study, i.e. between July 19, 2023 and May 1st, 2025. There are plans to release new versions of the dataset approximately once a year with additional data from participants who have been enrolled since the last dataset version release.",
      "rai:dataCollection": "Multiple modalities of data are collected for each participant, including survey data, clinical data, retinal imaging data, environmental sensor data, continuous glucose monitor data, and wearable activity monitor data. These encompass tabular data, imaging data, and physiological signal/waveform data. There is no unstructured text data included in this dataset. The exact forms used for data collection in REDCap are available here. Furthermore, all modalities, file formats, and devices are detailed in the dataset documentation at https://docs.aireadi.org/.",
      "rai:dataCollectionType": [
        "Manual Human Curation"
      ],
      "rai:dataCollectionMissingData": "Yes, not all modalities are available for all participants. Some participants elected not to participate in some study elements. In a few cases, the data collection device did not have any stored results or was returned too late to retrieve the results (e.g. battery died, data was lost). In a few cases, there may have been a data collision at some point in the process and data has been lost.",
      "rai:dataCollectionRawData": "Each instance consists of all of the data available for an individual participating in the study.",
      "rai:dataPreprocessingProtocol": [
        "\nThere were several quality control measures used at the time of data entry/acquisition. For example, clinical data outside of expected min/max ranges were flagged in REDCap, which was visible in reports viewed by clinical research coordinators (CRCs) and Data Managers. Using these REDCap reports as guides, Data Managers and CRCs examined participant records and determined if an error was likely. Data were checked for the following and edited if errors were detected:\n\n\ti. Credibility, based on range checks to determine if all responses fall within a prespecified reasonable range\n ii. Incorrect flow through prescribed skip patterns\n\tiii. Missing data that can be directly filed from other portions of an individual’ s record\n\tiv. The omission and/or duplication of records\n\nEditing was only done under the guidance and approval of the site PI. If corrected data was available from elsewhere in the respondent’s answers, the error was corrected. If there was no logical or appropriate way to correct the data, the Data site PI reviewed the values and made decisions about whether those values should be removed from the data.\n\nOnce data +were sent from each of the study sites to the central project team, additional processing steps were conducted in preparation for dissemination. For example, all data were mapped to standardized terminologies when possible, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model, a common data model for observational health data, and the Digital Imaging and Communications in Medicine (DICOM), a commonly used standard for medical imaging data. Details about the data processing approaches for each data domain/modality are described in the dataset documentation at https://docs.aireadi.org.\n"
      ],
      "rai:dataAnnotationProtocol": "N/A - no labels are provided",
      "rai:personalSensitiveInformation": [
        "EHR",
        "Wearable Monitoring",
        "ECG",
        "Environmental Sensor",
        "Continuous glucose monitor",
        "Wearable accelerometer"
      ],
      "completeness": "In cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.",
      "ro-crate-metadata": "retinal_photography/ro-crate-metadata.json",
      "contentUrl": "https://doi.org/10.60775/fairhub.3",
      "contact": "https://docs.aireadi.org/docs/3/contact"
    },
    {
      "@id": "ark:59853/rocrate-b2ai-ai-readi-wearable-blood-glucose",
      "@type": [
        "https://w3id.org/EVI#Dataset",
        "https://w3id.org/EVI#ROCrate"
      ],
      "conformsTo": {
        "@id": "https://w3id.org/fairscape/profile/0.1"
      },
      "name": "AI-READI Continuous Glucose Monitoring Subcrate",
      "description": "The Dexcom G6 is a real-time, integrated continuous glucose monitoring system (iCGM) that directly monitors blood glucose levels without requiring finger pricks. The device must be worn continuously in order to collect data day and night. A tiny filament called a glucose sensor is inserted under the skin to measure glucose levels in tissue fluid. This filament remains under the skin while it is worn. The internal sensor is connected to the transmitter that sits on top of the skin. The battery life for the transmitter is sufficient to power the system for three months. It is approximately the size of a quarter and adheres to the skin with medical tape. The Dexcom G6 Continuous Glucose Monitor (CGM) captures blood glucose readings every five minutes using this sensor. A single sensor is designed to last for a maximum of ten days, after which time the Dexcom G6 will require the insertion of a new sensor. The G6 transmitter will only save data for thirty days, therefore the data must be downloaded within thirty days from activation or all data will be lost.\n\nIn the AI-READI program, we have asked the research participants to wear the Dexcom CGM for ten days concurrently with wearing the Garmin Activity Monitor and using the home environmental sensor.\n",
      "keywords": [
        "diabetes mellitus",
        "Machine Learning",
        "Artificial Intelligence",
        "Electrocardiography",
        "Continuous Glucose Monitoring",
        "Retinal Imaging",
        "Eye Exam"
      ],
      "version": "3.0.0",
      "datePublished": "11/17/25",
      "isPartOf": [
        {
          "@id": "ark:59853/rocrate-b2ai-aireadi-release-3-0-0"
        }
      ],
      "hasPart": [],
      "author": [
        "AI-READI Consortium"
      ],
      "publisher": "AI-READI Consortium",
      "principalInvestigator": "Aaron Lee, Department of Ophthalmology, University of Washington",
      "funder": "NIH grant 1OT2OD032644 to the Bridge2AI: Salutogenesis Data Generation Project through the NIH Bridge2AI Common Fund program",
      "citation": "https://docs.aireadi.org",
      "associatedPublication": [
        "AI-READI Consortium. (2024). \"AI-READI: rethinking data collection, preparation and\nsharing for propelling AI-based discoveries in diabetes research and beyond.\"\nNature metabolism. https://doi.org/10.1038/s42255-024-01165-x",
        "AI-READI Consortium. (2025). Flagship Dataset of Type 2 Diabetes from the\nAI-READI Project (3.0.0) [Data set]. FAIRhub.\nhttps://doi.org/10.60775/fairhub.3"
      ],
      "identifier": "https://doi.org/10.60775/fairhub.3",
      "license": "https://doi.org/10.5281/zenodo.17555036",
      "conditionsOfAccess": "https://fairhub.io/datasets/3/access",
      "copyrightNotice": "Copyright © 2026 AI-READI",
      "ethicalReview": "Camille Nebeker, Debra Mathews, Kadija Ferryman, Nicholas Evans",
      "confidentialityLevel": "HL7:2N (normal)",
      "irb": {
        "@type": "Organization",
        "name": "Washington University IRB",
        "contactPoint": {
          "@type": "ContactPoint",
          "contactType": "IRB Reliance Team",
          "email": "hsdrely@uw.edu",
          "telephone": ""
        },
        "address": {
          "@type": "PostalAddress",
          "streetAddress": "Human Subjects Division University of Washington 4333 Brooklyn Ave NE Box 359470",
          "addressLocality": "Seattle",
          "addressRegion": "WA",
          "postalCode": "98195-9470",
          "addressCountry": "US"
        }
      },
      "irbProtocolId": "STUDY00016228",
      "humanSubjectExemption": "",
      "fdaRegulated": false,
      "deidentified": true,
      "humanSubjectResearch": "Yes",
      "dataGovernanceCommittee": "AI-READI Consortium",
      "rai:dataLimitations": "\nWhile the AI-READI's cross-sectional database ultimately aims to achieve balance across race/ethnicity, biological sex, and diabetes presence and severity, the pilot study is not balanced across these parameters.\n\nThree recording sites were strategically selected to achieve broad recruitment: the University of Alabama at Birmingham (UAB), the University of California San Diego (UCSD), and the University of Washington (UW). The sites were chosen for geographic variability across the United States and to ensure representation across various racial and ethnic groups. Individuals from all demographic backgrounds were recruited at all 3 sites. Factors influencing the generalization of derived models include the predominantly urban and hospital-based recruitment, which may not fully capture all possible cultural and socioeconomic backgrounds. The study cohort may not provide a comprehensive representation of the population, as it does not include other races/ethnicities such as Pacific Islanders and Native Americans. Information on device make and model, including specific modalities like macula scans or wide scans during OCT, were documented to ensure repeatability. Moreover, the study included multiple devices for one measure to enhance generalizability and represent the broad range of equipment utilized in clinical settings.\n\nIn cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.\n",
      "rai:dataBiases": "\nUniform data collection protocols were implemented for all subjects, irrespective of their race/ethnicity, biological sex, or diabetes severity, across all study sites. The selection of study sites was intended to ensure varied representation and minimize the potential for sampling bias.\n",
      "rai:dataUseCases": "The purpose for creating the dataset was to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. T2DM is a growing public health threat. Yet, the current understanding of T2DM, especially in the context of salutogenesis, is limited. Given the complexity of T2DM, AI-based approaches may help with improving our understanding but a key issue is the lack of data ready for training AI models. The AI-READI dataset is intended to fill this gap.",
      "rai:dataReleaseMaintenancePlan": "The dataset gets released as static versions. This is the third version of the dataset and consists of data collected up through the end of the second year of the study, i.e. between July 19, 2023 and May 1st, 2025. There are plans to release new versions of the dataset approximately once a year with additional data from participants who have been enrolled since the last dataset version release.",
      "rai:dataCollection": "Multiple modalities of data are collected for each participant, including survey data, clinical data, retinal imaging data, environmental sensor data, continuous glucose monitor data, and wearable activity monitor data. These encompass tabular data, imaging data, and physiological signal/waveform data. There is no unstructured text data included in this dataset. The exact forms used for data collection in REDCap are available here. Furthermore, all modalities, file formats, and devices are detailed in the dataset documentation at https://docs.aireadi.org/.",
      "rai:dataCollectionType": [
        "Manual Human Curation"
      ],
      "rai:dataCollectionMissingData": "Yes, not all modalities are available for all participants. Some participants elected not to participate in some study elements. In a few cases, the data collection device did not have any stored results or was returned too late to retrieve the results (e.g. battery died, data was lost). In a few cases, there may have been a data collision at some point in the process and data has been lost.",
      "rai:dataCollectionRawData": "Each instance consists of all of the data available for an individual participating in the study.",
      "rai:dataPreprocessingProtocol": [
        "\nThere were several quality control measures used at the time of data entry/acquisition. For example, clinical data outside of expected min/max ranges were flagged in REDCap, which was visible in reports viewed by clinical research coordinators (CRCs) and Data Managers. Using these REDCap reports as guides, Data Managers and CRCs examined participant records and determined if an error was likely. Data were checked for the following and edited if errors were detected:\n\n\ti. Credibility, based on range checks to determine if all responses fall within a prespecified reasonable range\n ii. Incorrect flow through prescribed skip patterns\n\tiii. Missing data that can be directly filed from other portions of an individual’ s record\n\tiv. The omission and/or duplication of records\n\nEditing was only done under the guidance and approval of the site PI. If corrected data was available from elsewhere in the respondent’s answers, the error was corrected. If there was no logical or appropriate way to correct the data, the Data site PI reviewed the values and made decisions about whether those values should be removed from the data.\n\nOnce data +were sent from each of the study sites to the central project team, additional processing steps were conducted in preparation for dissemination. For example, all data were mapped to standardized terminologies when possible, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model, a common data model for observational health data, and the Digital Imaging and Communications in Medicine (DICOM), a commonly used standard for medical imaging data. Details about the data processing approaches for each data domain/modality are described in the dataset documentation at https://docs.aireadi.org.\n"
      ],
      "rai:dataAnnotationProtocol": "N/A - no labels are provided",
      "rai:personalSensitiveInformation": [
        "EHR",
        "Wearable Monitoring",
        "ECG",
        "Environmental Sensor",
        "Continuous glucose monitor",
        "Wearable accelerometer"
      ],
      "completeness": "In cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.",
      "ro-crate-metadata": "wearable_blood_glucose//ro-crate-metadata.json",
      "contentUrl": "https://doi.org/10.60775/fairhub.3",
      "contact": "https://docs.aireadi.org/docs/3/contact"
    },
    {
      "@id": "ark:59853/rocrate-b2ai-ai-readi-wearable-activity-monitor",
      "@type": [
        "https://w3id.org/EVI#Dataset",
        "https://w3id.org/EVI#ROCrate"
      ],
      "conformsTo": {
        "@id": "https://w3id.org/fairscape/profile/0.1"
      },
      "name": "AI READI Wearable Activity Monitoring Subcrate",
      "description": "Activity monitoring involves the use of wearable trackers, which track metrics like heart rate, steps, calories, and active minutes, etc, creating a detailed record of their daily activity.\n\nFor the AI-READI research program, the Garmin Vivosmart 5 fitness tracker was utilized for activity monitoring. Participants were advised to wear the wristwatch on their non-dominant wrist for 10 consecutive days, but they were free to choose their preferred wrist for wearing the device. Upon completion of the 10-day period, participants returned the wristwatch along with a form specifying 1) the wrist on which they wore the tracker and 2) their dominant hand. The Garmin Vivosmart 5 was worn concurrently with the continuous glucose monitoring device and the use of the home environmental sensor. The Garmin Vivosmart 5 continuously recorded data related to physical activities and sleep. There are gaps in data collection as the battery had to be charged every 2-3 days. The sampling frequency of the Fitness tracker is 5 seconds.\n",
      "keywords": [
        "diabetes mellitus",
        "Machine Learning",
        "Artificial Intelligence",
        "Electrocardiography",
        "Continuous Glucose Monitoring",
        "Retinal Imaging",
        "Eye Exam"
      ],
      "version": "3.0.0",
      "datePublished": "11/17/25",
      "isPartOf": [
        {
          "@id": "ark:59853/rocrate-b2ai-aireadi-release-3-0-0"
        }
      ],
      "hasPart": [],
      "author": [
        "AI-READI Consortium"
      ],
      "publisher": "AI-READI Consortium",
      "principalInvestigator": "Aaron Lee, Department of Ophthalmology, University of Washington",
      "funder": "NIH grant 1OT2OD032644 to the Bridge2AI: Salutogenesis Data Generation Project through the NIH Bridge2AI Common Fund program",
      "citation": "https://docs.aireadi.org",
      "associatedPublication": [
        "AI-READI Consortium. (2024). \"AI-READI: rethinking data collection, preparation and\nsharing for propelling AI-based discoveries in diabetes research and beyond.\"\nNature metabolism. https://doi.org/10.1038/s42255-024-01165-x",
        "AI-READI Consortium. (2025). Flagship Dataset of Type 2 Diabetes from the\nAI-READI Project (3.0.0) [Data set]. FAIRhub.\nhttps://doi.org/10.60775/fairhub.3"
      ],
      "identifier": "https://doi.org/10.60775/fairhub.3",
      "license": "https://doi.org/10.5281/zenodo.17555036",
      "conditionsOfAccess": "https://fairhub.io/datasets/3/access",
      "copyrightNotice": "Copyright © 2026 AI-READI",
      "ethicalReview": "Camille Nebeker, Debra Mathews, Kadija Ferryman, Nicholas Evans",
      "confidentialityLevel": "HL7:2N (normal)",
      "irb": {
        "@type": "Organization",
        "name": "Washington University IRB",
        "contactPoint": {
          "@type": "ContactPoint",
          "contactType": "IRB Reliance Team",
          "email": "hsdrely@uw.edu",
          "telephone": ""
        },
        "address": {
          "@type": "PostalAddress",
          "streetAddress": "Human Subjects Division University of Washington 4333 Brooklyn Ave NE Box 359470",
          "addressLocality": "Seattle",
          "addressRegion": "WA",
          "postalCode": "98195-9470",
          "addressCountry": "US"
        }
      },
      "irbProtocolId": "STUDY00016228",
      "humanSubjectExemption": "",
      "fdaRegulated": false,
      "deidentified": true,
      "humanSubjectResearch": "Yes",
      "dataGovernanceCommittee": "AI-READI Consortium",
      "rai:dataLimitations": "\nWhile the AI-READI's cross-sectional database ultimately aims to achieve balance across race/ethnicity, biological sex, and diabetes presence and severity, the pilot study is not balanced across these parameters.\n\nThree recording sites were strategically selected to achieve broad recruitment: the University of Alabama at Birmingham (UAB), the University of California San Diego (UCSD), and the University of Washington (UW). The sites were chosen for geographic variability across the United States and to ensure representation across various racial and ethnic groups. Individuals from all demographic backgrounds were recruited at all 3 sites. Factors influencing the generalization of derived models include the predominantly urban and hospital-based recruitment, which may not fully capture all possible cultural and socioeconomic backgrounds. The study cohort may not provide a comprehensive representation of the population, as it does not include other races/ethnicities such as Pacific Islanders and Native Americans. Information on device make and model, including specific modalities like macula scans or wide scans during OCT, were documented to ensure repeatability. Moreover, the study included multiple devices for one measure to enhance generalizability and represent the broad range of equipment utilized in clinical settings.\n\nIn cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.\n",
      "rai:dataBiases": "\nUniform data collection protocols were implemented for all subjects, irrespective of their race/ethnicity, biological sex, or diabetes severity, across all study sites. The selection of study sites was intended to ensure varied representation and minimize the potential for sampling bias.\n",
      "rai:dataUseCases": "The purpose for creating the dataset was to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. T2DM is a growing public health threat. Yet, the current understanding of T2DM, especially in the context of salutogenesis, is limited. Given the complexity of T2DM, AI-based approaches may help with improving our understanding but a key issue is the lack of data ready for training AI models. The AI-READI dataset is intended to fill this gap.",
      "rai:dataReleaseMaintenancePlan": "The dataset gets released as static versions. This is the third version of the dataset and consists of data collected up through the end of the second year of the study, i.e. between July 19, 2023 and May 1st, 2025. There are plans to release new versions of the dataset approximately once a year with additional data from participants who have been enrolled since the last dataset version release.",
      "rai:dataCollection": "Multiple modalities of data are collected for each participant, including survey data, clinical data, retinal imaging data, environmental sensor data, continuous glucose monitor data, and wearable activity monitor data. These encompass tabular data, imaging data, and physiological signal/waveform data. There is no unstructured text data included in this dataset. The exact forms used for data collection in REDCap are available here. Furthermore, all modalities, file formats, and devices are detailed in the dataset documentation at https://docs.aireadi.org/.",
      "rai:dataCollectionType": [
        "Manual Human Curation"
      ],
      "rai:dataCollectionMissingData": "Yes, not all modalities are available for all participants. Some participants elected not to participate in some study elements. In a few cases, the data collection device did not have any stored results or was returned too late to retrieve the results (e.g. battery died, data was lost). In a few cases, there may have been a data collision at some point in the process and data has been lost.",
      "rai:dataCollectionRawData": "Each instance consists of all of the data available for an individual participating in the study.",
      "rai:dataPreprocessingProtocol": [
        "\nThere were several quality control measures used at the time of data entry/acquisition. For example, clinical data outside of expected min/max ranges were flagged in REDCap, which was visible in reports viewed by clinical research coordinators (CRCs) and Data Managers. Using these REDCap reports as guides, Data Managers and CRCs examined participant records and determined if an error was likely. Data were checked for the following and edited if errors were detected:\n\n\ti. Credibility, based on range checks to determine if all responses fall within a prespecified reasonable range\n ii. Incorrect flow through prescribed skip patterns\n\tiii. Missing data that can be directly filed from other portions of an individual’ s record\n\tiv. The omission and/or duplication of records\n\nEditing was only done under the guidance and approval of the site PI. If corrected data was available from elsewhere in the respondent’s answers, the error was corrected. If there was no logical or appropriate way to correct the data, the Data site PI reviewed the values and made decisions about whether those values should be removed from the data.\n\nOnce data +were sent from each of the study sites to the central project team, additional processing steps were conducted in preparation for dissemination. For example, all data were mapped to standardized terminologies when possible, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model, a common data model for observational health data, and the Digital Imaging and Communications in Medicine (DICOM), a commonly used standard for medical imaging data. Details about the data processing approaches for each data domain/modality are described in the dataset documentation at https://docs.aireadi.org.\n"
      ],
      "rai:dataAnnotationProtocol": "N/A - no labels are provided",
      "rai:personalSensitiveInformation": [
        "EHR",
        "Wearable Monitoring",
        "ECG",
        "Environmental Sensor",
        "Continuous glucose monitor",
        "Wearable accelerometer"
      ],
      "completeness": "In cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.",
      "ro-crate-metadata": "wearable_activity_monitor/ro-crate-metadata.json",
      "contentUrl": "https://doi.org/10.60775/fairhub.3",
      "contact": "https://docs.aireadi.org/docs/3/contact"
    },
    {
      "@id": "ark:59853/rocrate-b2ai-ai-readi-environmental-sensor",
      "@type": [
        "https://w3id.org/EVI#Dataset",
        "https://w3id.org/EVI#ROCrate"
      ],
      "conformsTo": {
        "@id": "https://w3id.org/fairscape/profile/0.1"
      },
      "name": "AI-READI Environmental Sensor Subcrate",
      "description": "Environmental sensors are devices designed to detect and measure various environmental parameters such as temperature, humidity, air quality, and light intensity. Research indicates that environmental factors play a significant role in health outcomes, yet most of these studies have focused on outdoor conditions at a broad city or regional scale. They often overlook the critical aspect of the individual's home environment.\n\nIn the AI-READI study, a custom-designed sensor unit (LeeLab Anura) was utilized to obtain environmental sensor data from each subject's home for a period of 10 days. Clinical research coordinators provided the subjects with the device along with take-home instructions to place the device in an area frequently used. Upon return, the subject was asked to make a note of the location of the environmental sensor. The sensor then recorded particulate matter counts (PM 1.0, 2.5, 4, and 10), temperature, relative humidity, volatile organic compounds (VOCs), nitrogen oxides (NO and NO2), and 11 multi-spectral light intensity measurements.\n",
      "keywords": [
        "diabetes mellitus",
        "Machine Learning",
        "Artificial Intelligence",
        "Electrocardiography",
        "Continuous Glucose Monitoring",
        "Retinal Imaging",
        "Eye Exam"
      ],
      "version": "3.0.0",
      "datePublished": "11/17/25",
      "isPartOf": [
        {
          "@id": "ark:59853/rocrate-b2ai-aireadi-release-3-0-0"
        }
      ],
      "hasPart": [],
      "author": [
        "AI-READI Consortium"
      ],
      "publisher": "AI-READI Consortium",
      "principalInvestigator": "Aaron Lee, Department of Ophthalmology, University of Washington",
      "funder": "NIH grant 1OT2OD032644 to the Bridge2AI: Salutogenesis Data Generation Project through the NIH Bridge2AI Common Fund program",
      "citation": "https://docs.aireadi.org",
      "associatedPublication": [
        "AI-READI Consortium. (2024). \"AI-READI: rethinking data collection, preparation and\nsharing for propelling AI-based discoveries in diabetes research and beyond.\"\nNature metabolism. https://doi.org/10.1038/s42255-024-01165-x",
        "AI-READI Consortium. (2025). Flagship Dataset of Type 2 Diabetes from the\nAI-READI Project (3.0.0) [Data set]. FAIRhub.\nhttps://doi.org/10.60775/fairhub.3"
      ],
      "identifier": "https://doi.org/10.60775/fairhub.3",
      "license": "https://doi.org/10.5281/zenodo.17555036",
      "conditionsOfAccess": "https://fairhub.io/datasets/3/access",
      "copyrightNotice": "Copyright © 2026 AI-READI",
      "ethicalReview": "Camille Nebeker, Debra Mathews, Kadija Ferryman, Nicholas Evans",
      "confidentialityLevel": "HL7:2N (normal)",
      "irb": {
        "@type": "Organization",
        "name": "Washington University IRB",
        "contactPoint": {
          "@type": "ContactPoint",
          "contactType": "IRB Reliance Team",
          "email": "hsdrely@uw.edu",
          "telephone": ""
        },
        "address": {
          "@type": "PostalAddress",
          "streetAddress": "Human Subjects Division University of Washington 4333 Brooklyn Ave NE Box 359470",
          "addressLocality": "Seattle",
          "addressRegion": "WA",
          "postalCode": "98195-9470",
          "addressCountry": "US"
        }
      },
      "irbProtocolId": "STUDY00016228",
      "humanSubjectExemption": "",
      "fdaRegulated": false,
      "deidentified": true,
      "humanSubjectResearch": "Yes",
      "dataGovernanceCommittee": "AI-READI Consortium",
      "rai:dataLimitations": "\nWhile the AI-READI's cross-sectional database ultimately aims to achieve balance across race/ethnicity, biological sex, and diabetes presence and severity, the pilot study is not balanced across these parameters.\n\nThree recording sites were strategically selected to achieve broad recruitment: the University of Alabama at Birmingham (UAB), the University of California San Diego (UCSD), and the University of Washington (UW). The sites were chosen for geographic variability across the United States and to ensure representation across various racial and ethnic groups. Individuals from all demographic backgrounds were recruited at all 3 sites. Factors influencing the generalization of derived models include the predominantly urban and hospital-based recruitment, which may not fully capture all possible cultural and socioeconomic backgrounds. The study cohort may not provide a comprehensive representation of the population, as it does not include other races/ethnicities such as Pacific Islanders and Native Americans. Information on device make and model, including specific modalities like macula scans or wide scans during OCT, were documented to ensure repeatability. Moreover, the study included multiple devices for one measure to enhance generalizability and represent the broad range of equipment utilized in clinical settings.\n\nIn cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.\n",
      "rai:dataBiases": "\nUniform data collection protocols were implemented for all subjects, irrespective of their race/ethnicity, biological sex, or diabetes severity, across all study sites. The selection of study sites was intended to ensure varied representation and minimize the potential for sampling bias.\n",
      "rai:dataUseCases": "The purpose for creating the dataset was to enable future generations of artificial intelligence/machine learning (AI/ML) research to provide critical insights into type 2 diabetes mellitus (T2DM), including salutogenic pathways to return to health. T2DM is a growing public health threat. Yet, the current understanding of T2DM, especially in the context of salutogenesis, is limited. Given the complexity of T2DM, AI-based approaches may help with improving our understanding but a key issue is the lack of data ready for training AI models. The AI-READI dataset is intended to fill this gap.",
      "rai:dataReleaseMaintenancePlan": "The dataset gets released as static versions. This is the third version of the dataset and consists of data collected up through the end of the second year of the study, i.e. between July 19, 2023 and May 1st, 2025. There are plans to release new versions of the dataset approximately once a year with additional data from participants who have been enrolled since the last dataset version release.",
      "rai:dataCollection": "Multiple modalities of data are collected for each participant, including survey data, clinical data, retinal imaging data, environmental sensor data, continuous glucose monitor data, and wearable activity monitor data. These encompass tabular data, imaging data, and physiological signal/waveform data. There is no unstructured text data included in this dataset. The exact forms used for data collection in REDCap are available here. Furthermore, all modalities, file formats, and devices are detailed in the dataset documentation at https://docs.aireadi.org/.",
      "rai:dataCollectionType": [
        "Manual Human Curation"
      ],
      "rai:dataCollectionMissingData": "Yes, not all modalities are available for all participants. Some participants elected not to participate in some study elements. In a few cases, the data collection device did not have any stored results or was returned too late to retrieve the results (e.g. battery died, data was lost). In a few cases, there may have been a data collision at some point in the process and data has been lost.",
      "rai:dataCollectionRawData": "Each instance consists of all of the data available for an individual participating in the study.",
      "rai:dataPreprocessingProtocol": [
        "\nThere were several quality control measures used at the time of data entry/acquisition. For example, clinical data outside of expected min/max ranges were flagged in REDCap, which was visible in reports viewed by clinical research coordinators (CRCs) and Data Managers. Using these REDCap reports as guides, Data Managers and CRCs examined participant records and determined if an error was likely. Data were checked for the following and edited if errors were detected:\n\n\ti. Credibility, based on range checks to determine if all responses fall within a prespecified reasonable range\n ii. Incorrect flow through prescribed skip patterns\n\tiii. Missing data that can be directly filed from other portions of an individual’ s record\n\tiv. The omission and/or duplication of records\n\nEditing was only done under the guidance and approval of the site PI. If corrected data was available from elsewhere in the respondent’s answers, the error was corrected. If there was no logical or appropriate way to correct the data, the Data site PI reviewed the values and made decisions about whether those values should be removed from the data.\n\nOnce data +were sent from each of the study sites to the central project team, additional processing steps were conducted in preparation for dissemination. For example, all data were mapped to standardized terminologies when possible, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model, a common data model for observational health data, and the Digital Imaging and Communications in Medicine (DICOM), a commonly used standard for medical imaging data. Details about the data processing approaches for each data domain/modality are described in the dataset documentation at https://docs.aireadi.org.\n"
      ],
      "rai:dataAnnotationProtocol": "N/A - no labels are provided",
      "rai:personalSensitiveInformation": [
        "EHR",
        "Wearable Monitoring",
        "ECG",
        "Environmental Sensor",
        "Continuous glucose monitor",
        "Wearable accelerometer"
      ],
      "completeness": "In cases of survey data, skipped questions or incomplete responses are expected. In cases of using wearables, improper use, technical failure such as battery failure or system malfunction are expected. In cases of imaging data, patient uncooperation, noise that may obscure the images and technical failure such as system malfunction, and data transfer failures are expected.",
      "ro-crate-metadata": "environment/ro-crate-metadata.json",
      "contentUrl": "https://doi.org/10.60775/fairhub.3",
      "contact": "https://docs.aireadi.org/docs/3/contact"
    }
  ]
}


================================================================================

FILE: gdrive_1rJsa5kySlBRRNhsO_WY7N3bfSKtqDi-Q_row13.txt
PATH: data/preprocessed/individual/AI_READI/gdrive_1rJsa5kySlBRRNhsO_WY7N3bfSKtqDi-Q_row13.txt
SIZE: 151236 bytes
--------------------------------------------------------------------------------

SOURCE METADATA
Project: AI_READI
Source ID: irb_protocol
Source type: IRB
Source URL: https://docs.google.com/document/d/1rJsa5kySlBRRNhsO_WY7N3bfSKtqDi-Q/edit
Raw file: data/raw/AI_READI/gdrive_1rJsa5kySlBRRNhsO_WY7N3bfSKtqDi-Q_row13.docx
--------------------------------------------------------------------------------
The Human Subjects Division (HSD) strives to ensure that people with disabilities have access to all services and content. If you experience any accessibility-related issues with this form or any aspect of the application process, email hsdinfo@uw.edu for assistance.
INSTRUCTIONS
This form is only for studies that will be reviewed by the UW IRB. Before completing this form, check HSD’s website to confirm that this should not be reviewed by an external (non-UW) IRB.
If you are requesting a determination about whether the planned activity is human subjects research or qualifies for exempt status, you may skip all questions except those marked with [DETERMINATION] For example 1.1. [DETERMINATION] must be answered. Do not upload consent materials for determinations in Zipline as HSD does not review or approve them.
Answer all questions. If a question is not applicable to the research or if you believe you have already answered a question elsewhere in the application, state “NA” (and if applicable, refer to the question where you provided the information). If you do not answer a question, the IRB does not know whether the question was overlooked or whether it is not applicable. This may result in unnecessary “back and forth” for clarification. Use non-technical language as much as possible.
For collaborative or multi-site research, describe only the UW activities unless you are requesting that the UW IRB provide the review and oversight for non-UW collaborators or co-investigators as well.
You may reference other documents (such as a grant application) if they provide the requested information in non-technical language. Be sure to provide the document name, page(s), and specific sections, and upload it to Zipline. Also, describe any changes that may have occurred since the document was written (for example, changes that you’ve made during or after the grant review process). In some cases, you may need to provide additional details in the answer space as well as referencing a document.
NOTE: Do not convert this Word document to PDF. The ability to use “tracked changes” is required in order to modify your study and respond to screening requests
INDEX
1. Overview
2. Participants
3. Non-UW Research Setting
4. Recruiting and Screening Participants
5. Procedures
6. Children (Minors) and Parental Permission
7. Assent of Children (Minors)
8. Consent of Adults
9. Privacy and Confidentiality
10. Risk / Benefit Assessment
11. Economic Burden to Participants
12. Resources
13. Other Approvals, Permissions, and Regulatory Issues
1. OVERVIEW
Study Title:
1.1. [DETERMINATION]  Home institution. Identify the institution through which the lead researcher listed on the IRB application will conduct the research. Provide any helpful explanatory information.
In general, the home institution is the institution (1) that provides the researcher’s paycheck and that considers them to be a paid employee, or (2) at which the researcher is a matriculated student. Scholars, faculty, fellows, and students who are visiting the UW and who are the lead researcher: identify your home institution and describe the purpose and duration of your UW visit, as well as the UW department/center with which you are affiliated while at the UW.
Note that many UW clinical faculty members are paid employees of non-UW institutions.
The UW IRB provides IRB review and oversight for only those researchers who meet the criteria described in the SOP Use of the UW IRB.
1.2. [DETERMINATION]  Consultation history. Has there been any consultation with someone at HSD about this study?
It is not necessary to obtain advance consultation. However, if advance consultation was obtained, answering this question will help ensure that the IRB is aware of and considers the advice and guidance provided in that consultation.
 No
 Yes → Briefly describe the consultation: approximate date, with whom, and method (e.g., by email, phone call, in-person meeting).
1.3. [DETERMINATION]  Similar and/or related studies. Are there any related IRB applications that provide context for the proposed activities?
Examples of studies for which there is likely to be a related IRB application: Using samples or data collected by another study; recruiting subjects from a registry established by a colleague’s research activity; conducting Phase 2 of a multi-part project, or conducting a continuation of another study; serving as the data coordinating center for a multi-site study that includes a UW site.
Providing this information (if relevant) may significantly improve the efficiency and consistency of the IRB’s review.
 No
 Yes → Briefly describe the other studies or applications and how they relate to the proposed activities. If the other applications were reviewed by the UW IRB, please also provide: the UW IRB number, the study title, and the lead researcher’s name.
1.4. [DETERMINATION]  Externally-imposed urgency or time deadlines. Are there any externally-imposed deadlines or urgency that affect the proposed activity?
HSD recognizes that everyone would like their IRB applications to be reviewed as quickly as possible. To ensure fairness, it is HSD policy to review applications in the order in which they are received. However, HSD will assign a higher priority to research with externally-imposed urgency that is beyond the control of the researcher. Researchers are encouraged to communicate as soon as possible with their HSD staff contact person when there is an urgent situation (in other words, before submitting the IRB application). Examples: a researcher plans to test an experimental vaccine that has just been developed for a newly emerging epidemic; a researcher has an unexpected opportunity to collect data from students when the end of the school year is only four weeks away.
HSD may ask for documentation of the externally-imposed urgency. A higher priority should not be requested to compensate for a researcher’s failure to prepare an IRB application in a timely manner. Note that IRB review requires a certain minimum amount of time; without sufficient time, the IRB may not be able to review and approve an application by a deadline.
 No
 Yes → Briefly describe the urgency or deadline as well as the reason for it.
1.5. [DETERMINATION]  Objectives. Using lay language, describe the purpose, specific aims, or objectives that will be met by this specific project. If hypotheses are being tested, describe them. You will be asked to describe the specific procedures in a later section.
If this application involves the use of a HUD “humanitarian” device: describe whether the use is for “on-label” clinical patient care, “off-label” clinical patient care, and/or research (collecting safety and/or effectiveness data).
1.6. [DETERMINATION]  Study design. Provide a one-sentence description of the general study design and/or type of methodology.
Your answer will help HSD in assigning applications to reviewers and in managing workload. Examples: a longitudinal observational study; a double-blind, placebo-controlled randomized study; ethnographic interviews; web scraping from a convenience sample of blogs; medical record review; coordinating center for a multi-site study.
1.7. [DETERMINATION]  Intent. Check all the descriptors that apply to your study. You must check at least one box.
This question is essential for ensuring that your application is correctly reviewed. Please read each option carefully.
1.8. Background, experience, and preliminary work. Answer this question only if the proposed activity has one or more of the following characteristics. The purpose of this question is to provide the IRB with information that is relevant to its risk/benefit analysis.
Involves more than minimal risk (physical or non-physical)
Is a clinical trial, or
Involves having the subjects use a drug, biological, botanical, nutritional supplement, or medical device.
“Minimal risk” means that the probability and magnitude of harm or discomfort anticipated in the research are not greater than those ordinarily encountered in daily life or during the performance of routine physical or psychological examinations or tests.
1.8.a. Background. Provide the rationale and the scientific or scholarly background for the proposed activity, based on existing literature (or clinical knowledge). Describe the gaps in current knowledge that the project is intended to address.
This should be a plain language description. Do not provide scholarly citations. Limit your answer to less than one page, or refer to an attached document with background information that is no more than three pages long.
1.8.b. Experience and preliminary work. Briefly describe experience or preliminary work or data (if any) that you, your team, or your collaborators/co-investigators have that supports the feasibility and/or safety of this study.
It is not necessary to summarize all discussion that has led to the development of the study protocol. The IRB is interested only in short summaries about experiences or preliminary work that suggest the study is feasible and that risks are reasonable relative to the benefits. Examples: Your team has already conducted a Phase 1 study of an experimental drug which supports the Phase 2 study being proposed in this application; your team has already done a small pilot study showing that the reading skills intervention described in this application is feasible in an after-school program with classroom aides; your team has experience with the type of surgery that is required to implant the study device; the study coordinator is experienced in working with subjects who have significant cognitive impairment.
1.9. Supplements. Check all boxes that apply, to identify relevant SUPPLEMENTS that should be completed and uploaded to Zipline.
This section is here instead of at the end of the form to reduce the risk of duplicating information in this IRB Protocol form that you will need to provide in these Supplements.
1.10. [DETERMINATION]  Confirm by checking the box below that you will comply with the COVID requirements described on HSD’s COVID webpage, which are based on the location of the in-person study procedures and the vaccination status of study team members and study participants.
Review the HSD website for current guidelines about which in-person research activities are allowable.
 Confirmed
2. PARTICIPANTS
2.1. [DETERMINATION]  Participants. Describe the general characteristics of the subject populations or groups, including age range, gender, health status, and any other relevant characteristics.
2.2. [DETERMINATION]  Inclusion and exclusion criteria.
2.2.a. Inclusion criteria. Describe the specific criteria that will be used to decide who will be included in the research from among interested or potential subjects. Define any technical terms in lay language.
2.2.b. Exclusion criteria. Describe the specific criteria that will be used to decide which of the subjects who meet the inclusion criteria listed above will be excluded from the research. Define any technical terms in lay language.
2.3. [DETERMINATION]  Prisoners. IRB approval is required in order to include prisoners in research, even when prisoners are not an intended target population.
Is the research likely to have subjects who become prisoners while participating in the study?
For example, a longitudinal study of youth with drug problems is likely to have subjects who will be prisoners at some point during the study.
 No
 Yes → If a subject becomes a prisoner while participating in the study, will any study procedures and/or data collection related to the subject be continued while the subject is a prisoner?
 No
 Yes → Describe the procedures and/or data collection that will continue with prisoner subjects.
2.4. [DETERMINATION]  Will the proposed research recruit or obtain data from individuals that are known to be prisoners?
For records reviews: if the records do not indicate prisoner status and prisoners are not a target population, select “No”. See the GUIDANCE Prisoners for the definition of “prisoner”, which is not necessarily tied to the type of facility in which a person is residing.
 No
 Yes → Answer the following questions (2.4.a. – 2.4.d.)
2.4.a. Describe the type of prisoners, and their locations(s).
2.4.b. One concern about prisoner research is whether the effect of participation on prisoners’ general living conditions, medical care, quality of food, amenities, and/or opportunity for earnings in prison will be so great that it will make it difficult for prisoners to adequately consider the research risks. How will the chances of this be reduced?
2.4.c. Describe what will be done to make sure that (a) recruitment and subject selection procedures will be fair to all eligible prisoners and (b) prison authorities or other prisoners will not be able to arbitrarily prevent or require particular prisoners from participating.
2.4.d. If the research is funded by one of these federal departments and agencies (Health & Human Services; Energy; Defense; Homeland Security; CIA; Social Security Administration), and/or will involve prisoners in federal facilities or in state/local facilities outside of Washington State: check the box below to provide assurance that study team members will (a) not encourage or facilitate the use of a prisoner’s participation in the research to influence parole or pardon decisions, and (b) clearly inform each prisoner in advance (for example, in a consent form) that participation in the research will have no effect on his or her parole or pardon.
 Confirmed
2.5. [DETERMINATION]  Protected populations. IRB approval is required for the use of the subject populations listed here. Check the boxes for any of these populations that will be purposefully included. (In other words, being a part of the populations is an inclusion criterion for the study.)
The WORKSHEETS describe the criteria for approval but do not need to be completed and should not be submitted.
2.5.a. If you check any of the boxes above, use this space to provide any information that may be relevant for the IRB to consider.
2.6. [DETERMINATION]  Native Americans or non-U.S. indigenous populations. Will Native American or non-U.S. indigenous populations be actively recruited through a tribe, tribe-focused organization, or similar community-based organization?
Indigenous people are defined in international or national legislation as having a set of specific rights based on their historical ties to a particular territory and their cultural or historical distinctiveness from other populations that are often politically dominant.
Examples: a reservation school or health clinic; recruiting during a tribal community gathering.
 No
 Yes → Name the tribe, tribal-focused organization, or similar community-based organization. The UW IRB expects that tribal/indigenous approval will be obtained before beginning the research. This may or may not involve approval from a tribal IRB. The study team and any collaborators/investigators are also responsible for identifying any tribal laws that may affect the research.
2.7.	[DETERMINATION]  UW Medicine and UW Dentistry residents and fellows. Will the research involve UW Medicine or UW Dentistry residents or fellows as study subjects?
 No
 Yes → (1) Describe in the Recruiting section (4.1) and Risks section (10.1) how you will ensure that residents feel free to truly make a voluntary decision about participation (i.e., no negative consequences from supervisors for saying “No”) and how you will ensure that any research data will not be used in the residents’ supervisor or program evaluation of them; AND (2) You must inform the UW HR Labor Relations representative who negotiates with the resident’s union about the study before beginning it. This is currently Jennifer Mallahan mallaj@uw.edu .
2.8.	 [DETERMINATION]  Third party subjects. Will the research collect private identifiable information about individuals other than the study subjects? Common examples include: collecting medical history information or contact information about family members, friends, co-workers.
“Identifiable” means any direct or indirect identifier that, alone or in combination, would allow you or another member of the research team to readily identify the person. For example, suppose that the research is about immigration history. If subjects are asked questions about their grandparents but are not asked for names or other information that would allow easy identification of the grandparents, then private identifiable information is not being collected about the grandparents and the grandparents are not subjects.
 No
 Yes → These individuals are considered human subjects in the study. Describe them and what data will be collected about them.
2.9. Number of subjects. Is it possible to predict or describe the maximum number of subjects (or subject units) needed to complete the study, for each subject group?
Subject units mean units within a group. For most research studies, a group will consist of individuals. However, the unit of interest in some research is not the individual. Examples:
Dyads such as caregiver-and-Alzheimer’s patient, or parent and child
Families
Other units, such as student-parent-teacher
Subject group means categories of subjects that are meaningful for the specific study. Some research has only one subject group – for example, all UW students taking Introductory Psychology. Some common ways in which subjects are grouped include:
By intervention – for example, an intervention group and a control group.
By subject population or setting – for example, urban versus rural families
By age – for example, children who are 6, 10, or 14 years old.
The IRB reviews the number of subjects in the context of risks and benefits. Unless otherwise specified, if the IRB determines that the research involves no more than minimal risk: there are no restrictions on the total number of subjects that may be enrolled. If the research involves more than minimal risk: The number of enrolled subjects must be limited to the number described in this application. If it is necessary later to increase the number of subjects, submit a Modification. Exceeding the IRB-approved number (over-enrollment) will be considered non-compliance.
 No → Provide the rationale in the box below. Also, provide any other available information about the scope/size of the research. You do not need to complete the table.
Example: It may not be possible to predict the number of subjects who will complete an online survey advertised through Craigslist, but you can state that the survey will be posted for two weeks and the number who respond is the number who will be in the study.
 Yes → For each subject group, use the table below to provide the estimate of the maximum desired number of individuals (or other subject unit, such as families) who will complete the research.
3. NON-UW RESEARCH SETTINGS
Complete this section only if UW investigators and people named in the SUPPLEMENT Non-UW Individual Investigators will conduct research procedures outside of UW and Harborview
3.1. [DETERMINATION]  Research locations and rationale. Identify the locations where the research will be conducted and include a description of the reason(s) for choosing the locations. If the research will be conducted internationally, be sure to list all the countries where the research will take place.
This is especially important when the research will occur in locations or with populations that may be vulnerable to exploitation. One of the three ethical principles the IRB must consider is Justice: ensuring that reasonable, non-exploitative, and well-considered procedures are administered fairly, with a fair distribution of costs and potential benefits.
3.2. [DETERMINATION]  Local context. Culturally appropriate procedures and an understanding of local context are an important part of protecting subjects. Describe any site-specific cultural issues, customs, beliefs, or values that may affect the research, how it is conducted, or how consent is obtained or documented.
Examples: It would be culturally inappropriate in some international settings for a woman to be directly contacted by a male researcher; instead, the researcher may need to ask a male family member for permission before the woman can be approached. It may be appropriate to obtain permission from community leaders prior to obtaining consent from individual members of a group. In some distinct cultural groups, signing forms may not be the norm.
This federal site maintains an international list of human research standards and requirements: http://www.hhs.gov/ohrp/international/index.html
3.3. [DETERMINATION]   Location-specific laws. Describe any local laws that may affect the research (especially the research design and consent procedures). The most common examples are laws about:
Specimens – for example, some countries will not allow biospecimens to be taken out of the country.
Age of consent – laws about when an individual is considered old enough to be able to provide consent vary across states, and countries.
Legally authorized representative – laws about who can serve as a legally authorized representative (and who has priority when more than one person is available) vary across states and countries.
Use of healthcare records – many states have laws that are similar to the federal HIPAA law but that have additional requirements.
3.4. [DETERMINATION]  Location specific administrative or ethical requirements. Describe local administrative or ethical requirements that affect the research.
Example: A school district may require researchers to obtain permission from the head district office as well as school principals before approaching teachers or students; a factory in China may allow researchers to interview factory workers but not allow the workers to be paid for their participation.
3.5. [DETERMINATION]  If the PI is a student: Does the research involve traveling outside of the U.S.?
 No
 Yes → Confirm by checking the box that (1) you will register with the UW Office of Global Affairs before traveling; (2) you will notify your advisor when the registration is complete; and (3) you will request a UW Travel Waiver is the research involves travel to the list of countries requiring a UW Travel Waiver.
 Confirmed
4. RECRUITING AND SCREENING PARTICIPANTS
4.1. [DETERMINATION]  Recruiting and screening. Describe how subjects will be identified, recruited, and screened. Include information about: how, when, where, and in what setting. Identify who (by position or role, not name) will approach and recruit subjects, and who will screen them for eligibility.
Note: Per UW Medicine policy, the UW Medicine eCare/MyChart system may not be used for research recruitment purposes. Additionally, researchers may not use UW Medicine’s Epic Care Everywhere data for research purposes unless the clinical data is necessary for patient/participant safety activities. This means Care Everywhere data cannot be used for recruitment, data abstraction, or any research activities other than those necessary for patient/participant safety.
4.2. Recruitment materials.
4.2.a. What materials (if any) will be used to recruit and screen subjects?
Examples: talking points for phone or in-person conversations; video or audio presentations; websites; social media messages; written materials such as letters, flyers for posting, brochures, or printed advertisements; questionnaires filled out by potential subjects.
4.2.b. Upload descriptions of each type of material (or the materials themselves) to Zipline. If letters or emails will be sent to any subjects, these should include a statement about how the subject’s name and contact information were obtained. No sensitive information about the person (such as a diagnosis of a medical condition) should be included in the letter. The text of these letters and emails must be uploaded to Zipline (i.e., a description will not suffice).
HSD encourages researchers to consider uploading descriptions of most recruitment and screening materials instead of the materials themselves. The goal is to provide the researchers with the flexibility to change some information on the materials without submitting a Modification for IRB approval of the changes. Examples:
Provide a list of talking points that will be used for phone or in-person conversations instead of a script.
For the description of a flyer, include the information that it will provide the study phone number and the name of a study contact person (without providing the actual phone number or name). This means that a Modification would not be necessary if/when the study phone number or contact person changes. Also, instead of listing the inclusion/exclusion criteria, the description below might state that the flyer will list one or a few of the major inclusion/exclusion criteria.
For the description of a video or a website, include a description of the possible visual elements and a list of the content (e.g., study phone number; study contact person; top three inclusion/exclusion criteria; payment of $
50; study name; UW researcher).
4.3. [DETERMINATION]  Relationship with participant population. Do any members of the study team have an existing relationship with the study population(s)?
Example: a study team member may have a dual role with the study population such as being their clinical care provider, teacher, laboratory directory or tribal leader in addition to recruiting them for their research.
 No
 Yes → Describe the nature of the relationship.
4.4. Payment to participants. The IRB must evaluate subject payment for the possibility that it will unduly influence subjects to participate. Refer to GUIDANCE Subject Payment when designing subject payment plans. Provide the following information about your plans for paying research subjects in the text box below or note that the information can be found in the consent form.
The total amount/value of the payment
Schedule/timing of the payment [i.e., when will subjects receive the payment(s)]
Purpose of the payment [e.g., reimbursement, compensation, incentive]
Whether payment will be “pro-rated” so that participants who are unable to complete the research may still receive some part of the payment
The IRB expects the consent process or study information provided to the subjects to include all of the above-listed information about payment, including the number and amount of payments, and especially when subjects can expect to receive payment. One of the most frequent complaints received by HSD is from subjects who expected to receive cash or a check on the day that they completed a study and who were angry or disappointed when payment took 6-8 weeks to reach them.
Researchers should review current UW Financial Management requirements about when Social Security Numbers must be collected, and when research payment must be reported to the UW Tax Office and the IRS: https://finance.uw.edu/ps/how-pay/research-subjects.
If your study involves the use of Amazon’s Mechanical Turk (MTurk), you must comply with the UW Procurement Services policy that no UW employee, family member, or student directly involved in the research will participate as a subject. The policy requires adding a qualifying question that asks whether the subject is a UW employee or family member, or UW student who is directly involved in the research. If they answer yes, they must be disqualified from MTurk activities.
4.5. [DETERMINATION]  Non-monetary compensation. Describe any non-monetary compensation that will be provided. Example; extra credit for students; a toy for a child.
4.5.a. If class credit will be offered to students, there must be an alternate way for the students to earn the extra credit without participating in the research. If class credit will be offered, describe the alternative non-research method by which students can earn that same course credit, including who will provide the alternative (e.g., a student subject pool; the course instructor).
4.6. [DETERMINATION]  Will data or specimens be accessed or obtained for recruiting and screening procedures prior to enrollment?
Examples: names and contact information; the information gathered from records that were screened; results of screening questionnaires or screening blood tests; Protected Health Information (PHI) from screening medical records to identify possible subjects.
 No → Skip the rest of this section; go to question 5.1.
 Yes → Describe the data and/or specimens (including PHI) and whether it will be retained as part of the study data.
4.7. Consent for recruiting and screening. Will consent be obtained for any of the recruiting and screening procedures? (Section 8: Consent of Adults asks about consent for the main study procedures).
“Consent” includes: consent from individuals for their own participation; parental permission; assent from children; consent from a legally authorized representative for adult individuals who are unable to provide consent.
Examples:
For a study in which names and contact information will be obtained from a registry: the registry should have consent from the registry participants to release their names and contact information to researchers.
For a study in which possible subjects are identified by screening records: there will be no consent process.
For a study in which individuals respond to an announcement and call into a study phone line: the study team person talking to the individual may obtain non-written consent to ask eligibility questions over the phone.
 No → Skip the rest of this section; go to question 5.1.
 Yes → Describe the consent process.
4.7.a. Documentation of consent. Will a written or verifiable electronic signature from the subject on a consent form be used to document consent for the recruiting and screening procedures?
 No → Describe the information that will be provided during the consent process and for which procedures.
 Yes, written → If yes, and a written signature will be used to document consent:
Upload the consent form to Zipline.
 Yes, electronic → If yes, and an electronic signature will be used to document consent:
Upload the consent form to Zipline.
If the eSignature process or method for recruiting and screening is different than for the main study procedures, use the questions about electronic consent in Sections 8.3. and 8.4. to differentiate between recruiting/screening and main study electronic consent. If electronic consent will be used for recruiting/screening but not main study consent, use 8.3. and 8.4. to describe e-consent and note that it is only for recruiting/screening.
5. PROCEDURES
5.1. [DETERMINATION]  Study procedures. Using lay language, provide a complete description of the study procedures, including the sequence, intervention or manipulation (if any), drug dosing information (if any), blood volumes and frequency of draws (if any), use of records, time required, and setting/location. If it is available: Upload a study flow sheet or table to Zipline.
For studies comparing standards of care: It is important to accurately identify the research procedures. See UW IRB GUIDANCE Risks of Harm from Standard Care and the draft guidance from the federal Office of Human Research Protections, “Guidance on Disclosing Reasonably Foreseeable Risks in Research Evaluating Standards of Care”; October 20, 2014.
Information about pediatric blood volume and frequency of draws that would qualify for expedited review can be found in this reference table on the Seattle Children’s IRB website.
5.2. [DETERMINATION]  Recordings. Does the research involve creating audio or video recordings?
 No → Go to question 5.3.
 Yes → Verify that you have described what will be recorded in the answer to question 5.1., and answer question 5.2.a.
5.2.a. Before recording, will consent for being recorded be obtained from subjects and any other individuals who may be recorded?
 No → Email hsdinfo@uw.edu before submitting this application in Zipline. In the email, include a brief description of the research and a note that individuals will be recorded without their advance consent.
 Yes
5.3. [DETERMINATION]  MRI scans. Will any subjects have a Magnetic Resonance Imaging (MRI) scan as part of the study procedures?
This means scans that are performed solely for research purposes or clinical scans that are modified for research purposes (for example, using a gadolinium-based contrast agent when it is not required for clinical reasons).
 No → Go to question 5.4.
 Yes → Answer questions 5.3.a through 5.3.c.
5.3.a. Describe the MRI scan(s). Specifically:
What is the purpose of the scan(s)? Examples: obtain research data; safety assessment associated with a research procedure.
Which subjects will receive an MRI scan?
Describe the minimum and maximum number of scans per subject, and over what time period the scans will occur. For example: all subjects will undergo two MRI scans, six months apart.
5.3.b. MRI facility. At which facility(ies) will the MRI scans occur? Check all that apply.
 UWMC Radiology/Imaging Services (the UWMC clinical facility)
 DISC Diagnostic Imaging Sciences Center (UWMC research facility)
 CHN Center for Human Neuroscience MRI Center (Arts & Sciences research facility)
 BMIC Biomolecular Imaging Center (South Lake Union research facility)
 Harborview Radiology/Imaging Services (the Harborview clinical facility)
 SCCA Imaging Services
 Northwest Diagnostic Imaging
 Other: identify in the text box below:
5.3.c. Personnel. For MRI scans that will be conducted at the DISC, CHN or BMIC research facilities: Indicate who will be responsible for operating the MRI scanner by checking all that apply.
 MRI technician who is formally qualified
 Researcher who has completed scanner operator training provided by a qualified MRI operator
5.4. [DETERMINATION]  Data variables. Describe the specific data that will be obtained (including a description of the most sensitive items). Alternatively, a list of the data variables may be uploaded to Zipline.
5.5. [DETERMINATION]  Data sources. For all types of data that will be accessed or collected for this research: Identify whether the data are being obtained from the subjects (or subjects’ specimens) or whether they are being obtained from some other source (and identify the source).
If you have already provided this information in Question 5.1, you do not need to repeat the information here.
5.6. [DETERMINATION]  Identifiability of data and specimens. Answer these questions carefully and completely. This will allow HSD to accurately determine the type of review that is required and the relevant compliance requirements. Review the following definitions before answering the questions:
Access means to view or perceive data, but not to possess or record it. See, in contrast, the definition of “obtain”.
Identifiable means that the identity of an individual is or may be readily (1) ascertained by the researcher or any other member of the study team from specific data variables or from a combination of data variables, or (2) associated with the information.
Direct identifiers are direct links between a subject and data/specimens. Examples include (but are not limited to): name, date of birth, medical record number, email or IP address, pathology or surgery accession number, student number, or a collection of data that is (when taken together) identifiable.
Indirect identifiers are information that links between direct identifiers and data/specimens. Examples: a subject code or pseudonym.
Key refers to a single place where direct identifiers and indirect identifiers are linked together so that, for example, coded data can be identified as relating to a specific person. Example: a master list that contains the data code and the identifiers linked to the codes.
Obtain means to possess or record in any fashion (writing, electronic document, video, email, voice recording, etc.) for research purposes and to retain for any length of time. This is different from accessing, which means to view or perceive data.
5.6.a. Will you or any members of you team have access to any direct or indirect identifiers?
 Yes → Describe which identifiers and for which data/specimens.
 No → Select the reason(s) why you (and all members of your team) will not have access to direct or indirect identifiers.
 There will be no identifiers
 Identifiers or the key have been (or will have been) destroyed before access.
 There is an agreement with the holder of the identifiers (or key) that prohibits the release of the identifiers (or key) to study team members under any circumstances.
This agreement should be available upon request from the IRB. Examples: a Data Use Agreement, Repository Gatekeeping form, or documented email.
 There are written policies and procedures for the repository/database/data management center that prohibit the release of the identifiers (or identifying link). This includes situations involving an Honest Broker.
 There are other legal requirements prohibiting the release of the identifiers or key. Describe them below.
5.6.b. Will you or any study team members obtain any direct or indirect identifiers?
 Yes → Describe which identifiers and for which data/specimens.
 No → Select the reason(s) why you (and all members of your team) will not obtain direct or indirect identifiers.
 There will be no identifiers.
 Identifiers or the key have been (or will have been) destroyed before access.
 There will be an agreement with the holder of the identifiers (or key) that prohibits the release of the identifiers (or key under any circumstances.
This agreement should be available upon request from the IRB. Examples: a Data Use Agreement, Repository Gatekeeping form, or documented email.
 There are written policies and procedures for the repository/database/data management center that prohibit the release of the identifiers (or identifying link). This includes situations involving an Honest Broker.
 There are other legal requirements prohibiting the release of the identifiers or key. Describe them below.
5.6.c. If any identifiers will be obtained, indicate how the identifiers will be stored (and for which data). NOT: Do not describe the data security plan here, that information is requested in question 9.6.
 Identifiers will be stored with the data. Describe the data to which this applies:
 Identifiers and study data will be stored separately but a link will be maintained between the identifiers and the study data (for example, through the use of a code). Describe the data to which this applies:
 Identifiers and study data will be stored separately, with no link between the identifiers and the study data. Describe the data to which this applies:
5.6.d. Research collaboration. Will individuals who provide coded information or specimens for the research also collaborate on other activities for this research? If yes, identify the activities and provide the name of the collaborator’s institution/organization.
Examples include but are not limited to: (1) study, interpretation, or analysis of the data that results from the coded information or specimens; and (2) authorship on presentations or manuscripts related to this work.
5.7. [DETERMINATION]  Protected Health Information (PHI). Will participants’ identifiable PHI be accessed, obtained, used, or disclosed for any reason (for example, to identify or screen potential subjects, to obtain study data or specimens, for study follow-up) that does not involve the creation or obtaining of a Limited Data Set?
PHI is individually identifiable healthcare record information or clinical specimens from an organization considered a “covered entity” by federal HIPAA regulations, in any form or media, whether electronic, paper, or oral. You must answer yes to this question if the research involves identifiable health care records (e.g., medical, dental, pharmacy, nursing, billing, etc.), identifiable healthcare information from a clinical department repository, or observations or recordings of clinical interactions.
For information about what constitutes the UW Covered Entity, see UW Medicine Compliance Patient Information Privacy Policy 101 and diagram of the healthcare components.
 No → Skip the rest of this question; go to question 5.8.
 Yes → Answer all of the questions below (5.7.a. through 5.7.f.)
5.7.a. Describe the PHI and the reason for using it. Be specific. For example, will any “free text” fields (such as physician notes) be accessed, obtained or used?
5.7.b. Is any of the PHI located in Washington State?
 No
 Yes
5.7.c. Describe the pathway of how the PHI will be accessed or obtained, starting with the source/location and then describing the system/path/mechanism by which it will be identified, accessed, and copied for the research. Be specific. For example: directly view records; search through a department’s clinical database; submit a request to Leaf.
5.7.d. For which PHI will subjects provide HIPAA authorization before the PHI is accessed, obtained and/or used?
Confirm by checking the box that UW Medicine HIPAA Authorization form maintained on the HSD website will be used to access obtain, use, or disclose any UW Medicine PHI.
 Confirmed
5.7.e. Will you obtain any HIPAA authorizations electronically (i.e., e-signature)?
 No
 Yes → Confirm by checking the box that you have read and understand the Electronic Documentation of Consent section of the WORKSHEET Consent Requirements and Waivers the GUIDANCE Consent Documentation of Consent for information regarding the use of electronic signatures and HIPAA authorizations.
 Confirmed
5.7.f. For which PHI will HIPAA authorization NOT be obtained from the subjects?
Provide the following assurances by checking the boxes.
 The minimum necessary amount of PHI to accomplish the purposes described in this application will be accessed, obtained and/or used.
 The PHI will not be reused or disclosed to any other person or entity, except as required by law, for authorized oversight of the research study, or for other research for which the use or disclosure of PHI would be permitted.
 The HIPAA “accounting for disclosures” requirement will be fulfilled, if applicable. See UW Medicine Compliance Policy #104.
 There will be reasonable safeguards to protect against identifying, directly or indirectly, any patient in any report of the research.
5.8. [DETERMINATION]  Genomic data sharing. Will the research obtain or generate genomic data?
 No
 Yes → Answer the question below.
5.8.a. Will genomic data from this research be sent to a national database (for example, NIH’s dbGaP database)?
 No
 Yes → Complete the SUPPLEMENT Genomic Data Sharing and upload it to Zipline.
5.9. Whole genome sequencing. For research involving biospecimens: Will the research include whole genome sequencing?
Whole genome sequencing is sequencing of a human germline or somatic specimen with the intent to generate the genome or exome sequence of that specimen.
 No
 Yes
5.10. [DETERMINATION]  Cannabis (marijuana), hemp, and related compounds. These questions are about: cannabis (any part of the plant in any form), hemp, cannabidiol (CBD), delta-8-THC, any product derived from cannabis or hemp, and related synthesized compounds. All UW research must comply with federal laws about cannabis because of conditions associated with the federal money that UW receives. Answer the questions below so that HSD can determine whether the federal laws apply to your specific situation. See the UW Guidance on Research Involving Marijuana for additional information.
5.10.a. Does your research involve any of the following? Check all that apply.
 Study staff will obtain or handle any of the above items
 Study will provide money to the participants to obtain any of the above items
 Study participants will use or consume any of the above items on campus or in any UW-owned or leased facility
 None of the above
5.10.b. If you checked any box except “None of the above”, provide the following information about each cannabis and related item your research will involve: Name of the item, how you will obtain it, the source, and whether it contains ≥0.3% THC (tetrahydrocannabinol).
5.11. Possible secondary use or sharing of information, specimens, or subject contact information. Is it likely that the obtained or collected information, specimens, or subject contact information will be used for any of the following:
Future research not described in this application (in other words, secondary research)
Submission to a repository, registry, or database managed by the study team, colleagues, or others for research purposes
Sharing with others for their own research
Please consider the broadest possible future plans and whether consent will be obtained now from the subjects for future sharing or research uses (which it may not be possible to describe in detail at this time). Answer YES even if future sharing or uses will use de-identified information or specimens. Answer NO if sharing is unlikely or if the only sharing will be through the NIH Genomic Data Sharing described in question 5.8.
Many federal grants and contracts now require data or specimen sharing as a condition of funding, and many journals require data sharing as a condition of publication. “Sharing” may include (for example): informal arrangements to share banked data/specimens with other investigators; establishing a repository that will formally share with other researchers through written agreements; or sending data/specimens to a third-party repository/archive/entity such as the Social Science Open Access Repository (SSOAR), or the UCLA Ethnomusicology Archive.
 No
 Yes → Answer all of the questions below. (Questions 5.11.a through 5.11.g)
5.11.a. Describe what will be stored for future use, including whether any direct or indirect (e.g., subject codes) identifiers will be stored.
5.11.b. Describe what will be shared with other researchers or with a repository/database/registry, including whether direct identifiers will be shared and (for specimens) what data will be released with the specimens.
5.11.c. Who will oversee and/or manage the sharing?
5.11.d. Describe the possible future uses, including limitations or restrictions (if any) on future uses or users. As stated at the beginning of this question, consider the broadest possible uses.
Examples: data will be used only for cardiovascular research; data will not be used for research on population origins.
5.11.e. Consent. Will consent be obtained now from subjects for the secondary use, banking and/or future sharing?
 No
 Yes → Be sure to include the information about this consent process in the consent form (if there is one) and in the answers to the consent question in Section 8.
5.11.f. Withdrawal. Will subjects be able to withdraw their data/specimens from secondary use, banking or sharing?
 No
 Yes → Describe how, and whether there are any limitations on withdrawal.
Example: data can be withdrawn from the repository but cannot be retrieved after they are released.
5.11.g. Agreements for sharing or release. Confirm by checking the box that the sharing or release will comply with UW (and, if applicable, UW Medicine) policies that require a formal agreement with the recipient for release of data or specimens to individuals or entities other than federal databases.
Data Use Agreements or Gatekeeping forms are used for data; Material Transfer Agreements are used for specimens (or specimens plus data). Do not attach any template agreement forms; the IRB neither reviews nor approves them.
 Confirmed
5.12. Communication with subjects during the study. Describe the types of communication (if any) the research team will have with already-enrolled subjects during the study. Provide a description instead of the actual materials themselves.
Examples: email, texts, phone, or letter reminders about appointments or about returning study materials such as a questionnaire; requests to confirm contact information.
5.13. Future contact with subjects. Is there a plan to retain any contact information for subjects so that they can be contacted in the future?
 No
 Yes → Describe the purpose of the future contact, and whether use of the contact information will be limited to the study team; if not, describe who else could be provided with the contact information. Describe the criteria for approving requests for the information.
Examples: inform subjects about other studies; ask subjects for additional information or medical record access that is not currently part of the study proposed in this application; obtain another sample.
5.14. Alternatives to participation. Are there any alternative procedures or treatments that might be advantageous to the subjects?
If there are no alternative procedures or treatments, select “No”. Examples of advantageous alternatives: earning extra class credit in some time-equivalent way other than research participation; obtaining supportive care or a standard clinical treatment from a health care provider instead of participating in research with an experimental drug.
 No
 Yes → Describe the alternatives.
5.15. Upload to Zipline all data collection forms (if any) that will be directly used by or with the subjects, and any scripts/talking points that will be used to collect the data. Do not include data collection forms that will be used to abstract data from other sources (such as medical or academic records), or video recordings.
Examples: survey, questionnaires, subject logs or diaries, focus group questions.
NOTE: Sometimes the IRB can approve the general content of surveys and other data collection instruments rather than the specific form itself. This prevents the need to submit a modification request for future minor changes that do not add new topics or increase the sensitivity of the questions. To request this general approval, use the text box below to identify the questionnaires/surveys/ etc. for which you are seeking this more general approval. Then briefly describe the scope of the topics that will be covered and the most personal and sensitive questions. The HSD staff person who screens this application will let you know whether this is sufficient or whether you will need to provide more information.
For materials that cannot be uploaded: upload screenshots or written descriptions that are sufficient to enable the IRB to understand the types of data that will be collected and the nature of the experience for the participant. You may also provide URLs (website addresses) or written descriptions below. Examples of materials that usually cannot be uploaded: mobile apps; computer-administered test; licensed and restricted standardized tests.
For data that will be gathered in an evolving way: This refers to data collection/questions that are not pre-determined but rather are shaped during interactions with participants in response to observations and responses made during those interactions. If this applies to the proposed research, provide a description of the process by which the data collection/questions will be established during the interactions with subjects, how the data collection/questions will be documented, the topics likely to be addressed, the most sensitive type of information likely to be gathered, and the limitations (if any) on topics that will be raised or pursued.
Use this text box (if desired) to provide:
Short written descriptions of materials that cannot be uploaded, such as URLs
A description of the process that will be used for data that will be gathered in an evolving way.
The general content of questionnaires, surveys and similar instruments for which general approval is being sought. (See the NOTE bullet point in the instructions above.)
5.16. [DETERMINATION]  SARS-CoV-2 testing. Will the subjects be tested for the SARS-CoV-2 coronavirus?
 No
 Yes → If yes:
Name the testing lab
Confirm that the lab and its use of this test is CLIA certified or certified by the Washington State Department of Health
Describe whether you will return the results to the participants and, if yes, who will do it and how (including any information you would provide to subjects with positive test results).
6. CHILDREN (MINORS) AND PARENTAL PERMISSION
6.1. [DETERMINATION]  Involvement of minors. Does the research include minors (children)?
Minor or child means someone who has not yet attained the legal age for consent for the research procedures, as described in the applicable laws of the jurisdiction in which the research will be conducted. This may or may not be the same as the definition used by funding agencies such as the National Institutes of Health.
In Washington State the generic age of consent is 18, meaning that anyone under the age of 18 is considered a child.
There are some procedures for which the age of consent is much lower in Washington State.
The generic age of consent may be different in other states, and in other countries.
 No → Go to Section 8.
 Yes → Provide the age range of the minor subjects for this study and the legal age for consent in the study population(s). If there is more than one answer, explain.
 Don’t know → This means is it not possible to know the age of the subjects. For example, this may be true for some research involving social media, the Internet, or a dataset that is obtained from another researcher or from a government agency. Go to Section 8.
6.2. Parental permission. Parental permission means actively obtaining the permission of the parents. This is not the same as “passive” or “opt out” permission where it is assumed that parents are allowing their children to participate because they have been provided with information about the research and have not objected or returned a form indicating they don’t want their children to participate.
6.2.a. Will parental permission be obtained for:
 All of the research procedures → Go to question 6.2.b.
 None of the research procedures → Use the table below to provide justification and skip question 6.2.b.
 Some of the research procedures → Use the table below to identify the procedures for which parental permission will not be obtained.
Be sure to consider all research procedures and plans, including screening, future contact, and sharing/banking of data and specimens for future work.
Table footnotes
If the answer is the same for all children groups or all procedures: collapse the answer across the groups and/or procedures.
If identifiable information or biospecimens will be obtained without parent permission, any waiver granted by the IRB does not override parents’ refusal to provide broad consent (for example, through the Northwest Biotrust).
Will parents be informed about the research beforehand even though active permission is not being obtained?
6.2.b. Indicate the plan for obtaining parental permission. One or both boxes must be checked.
 Both parents, unless one parent is deceased, unknown, incompetent, or not reasonably available, or when only one parent has legal responsibility for the care and custody of the child.
 One parent, even if the other parent is alive, known, competent, reasonably available, and shares legal responsibility for the care and custody of the child.
This is all that is required for minimal risk research.
If both are checked explain:
6.3. Children who are wards. Will any of the children be wards of the State or any other agency, institution, or entity?
 No
 Yes → An advocate may need to be appointed for each child who is a ward. The advocate must be in addition to any other individual acting on behalf of the child as guardian or in loco parentis. The same individual can serve as advocate for all children who are wards.
Describe who will be the advocate(s). The description must address the following points:
Background and experience
Willingness to act in the best interests of the child for the duration of the research
Independence of the research, research team, and any guardian organization
6.4. UW Office of the Youth Protection Coordinator. If the project involves interaction (in-person or remotely) with individuals under the age of 18, researchers must comply with UW Administrative Policy Statement 10.13 and the requirements listed at this website. This includes activities that are deemed to be Not Research or Exempt. It does not apply to third-party led research (i.e., research conducted by a non-UW PI). Information and FAQs for researchers are available.
This point is advisory only; there is no need to provide a response.
7. ASSENT OF CHILDREN (MINORS)
Go to Section 8 if your research does not involve children (minors).
When designing assent processes and forms, researchers should first review the GUIDANCE Consent Protected and Vulnerable Populations and TIPSHEET Consent Assent and Legally Authorized Representative.
7.1. Assent of children (minors). Though children do not have the legal capacity to “consent” to participate in research, they should be involved in the process if they are able to “assent” by having a study explained to them and/or by reading a simple form about the study, and then verbally expressing whether they want to participate. They may also provide a written assent if they are older. See GUIDANCE Consent Protected and Vulnerable Populations and WORKSHEET Children for circumstances in which a child’s assent may be unnecessary or inappropriate.
7.1.a. Will assent be obtained for:
 All research procedures and child groups → Go to question 7.2.
 None of the research procedures and child groups → Use the table below to provide justification, then skip to question 7.6.
 Some of your research procedures and child groups → Use the table below to identify the procedures for which assent will not be obtained.
Be sure to consider all research procedures and plans, including screening, future contact, and sharing/banking of data and specimens for future work.
Table footnotes
If the answer is the same for all children groups or all procedures, collapse your answer across the groups and/or procedures.
7.2. Assent process. Describe how assent will be obtained, for each child group. If the research involves children of different ages, answer separately for each group. If the children are non-English speakers, include a description of how their comprehension of the information will be evaluated.
7.3. Dissent or resistance. Describe how a child’s objection or resistance to participation (including non-verbal indications) will be identified during the research, and what the response will be.
7.4. E-consent. Will any electronic processes (email, websites, electronic signatures, etc.) be used to present assent information to subjects/and or to obtain documentation (signatures) of assent? If yes, describe how this will be done.
7.5. Documentation of assent. Which of the following statements describes whether documentation of assent will be obtained?
 None of the research procedures and child groups 	→ Use the table below to provide justification, then go to question 7.5.b.
 All of the research procedures and child groups 		→ Go to question 7.5.a., do not complete the table.
 Some of the research procedures and/or child groups 	→ Complete the table below and then go to question 7.5.a.
Table footnotes
If the answer is the same for all children groups or all procedures, collapse your answer across the groups and/or procedures.
7.5.a. Describe how assent will be documented. If the children are functionally illiterate or are not fluent in English, include a description of the documentation process for them.
7.5.b. Upload all assent materials (talking points, videos, forms, etc.) to Zipline. Assent materials are not required to provide all of the standard elements of adult consent; the information should be appropriate to the age, population, and research procedures. The documents should be in Word, if possible.
7.6. Children who reach the legal age of consent during participation in longitudinal research.
When children are enrolled at a young age and continue for many years, it is best practice to re-obtain assent (or to obtain it for the first time, if it was not obtained at the beginning of their participation).
When children reach the legal age of consent, informed consent must be obtained from the now-adult subject for (1) any ongoing interactions or interventions with the subjects, or (2) the continued analysis of specimens or data for which the subject’s identify is readily identifiable to the researcher, unless the IRB waives this requirement.
7.6.a. Describe the plans (if any) to re-obtain assent from children.
7.6.b. Describe the plans (if any) to obtain consent for children who reach the legal age of consent.
If adult consent will be obtained from them, describe what will happen regarding now-adult subjects who cannot be contacted.
If consent will not be obtained or will not be possible, explain why.
7.7. Other regulatory requirements. (This is for information only; no answer or response is required.) Researchers are responsible for determining whether their research conducted in schools, with student records, or over the Internet comply with permission, consent, and inspection requirements of the following federal regulations:
PPRA – Protection of Pupil Rights Amendment
FERPA – Family Education Rights and Privacy Act
COPPA – Children’s Online Privacy Protection Act
8 CONSENT OF ADULTS
Review the following definitions before answering the questions in this section.
8.1. Groups. Identify the groups to which the answers in this section apply:
 Adult subjects
 Parents who are providing permission for their children to participate in research
→ If you selected PARENTS, the word “consent” below should also be interpreted as applying to parental permission and “subjects” should also be interpreted as applying to the parents.
8.2. The consent process and characteristics. This series of questions is about whether consent will be obtained for all procedures except recruiting and screening, and, if yes, how.
The issue of consent for recruiting and screening activities is address in question 4.7. You do not need to repeat your answer to question 4.7.
8.2.a. Are there any procedures for which consent will not be obtained?
 No
 Yes → Use the table below to identify the procedures for which consent will not be obtained. “All” is an acceptable answer for some studies.
Be sure to consider all research procedures and plans, including future contact, and sharing/banking of data and specimens for future work.
Table footnotes
1. If the answer is the same for all groups, collapse your answer across the groups and/or procedures.
8.2.b. Describe the consent process, if consent will be obtained for any or all procedures, for any or all groups. Address groups and procedures separately if the consent processes are different.
Be sure to include:
The location/setting where consent will be obtained
Who will obtain consent (refer to positions, roles, or titles, not names)
How subjects will be provided sufficient opportunity to discuss the study with the research team and consider participation.
8.2.c. Comprehension. Describe the methods that will be used to ensure or test the subjects’ understanding of the information during the consent process.
8.2.d. Influence. Does the research involve any subject groups that might find it difficult to say “no” to participation because of the setting or their relationship with someone on the study team, even if they aren’t pressured to participate?
Examples: Student participants being recruited into their teacher’s research; patients being recruited into their healthcare provider’s research; study team members who are participants; outpatients recruited from an outpatient surgery waiting room just prior to their surgery.
 No
 Yes → Describe what will be done to reduce any effect of the setting or relationship on the participation decision.
Examples: a study coordinator will obtain consent instead of the subject’s physician; the researcher will not know which subjects agreed to participate; subjects will have two days to decide after hearing about the study.
8.2.e. Information provided is tailored to the needs of the subject population. Describe the basis for concluding that the information that will be provided to subjects (via written or oral methods) is what a reasonable member of the subject population(s) would want to know. If the research consent materials contain a key information section, also describe the basis for concluding that the information present in that section is that which is most likely to assist the selected subject population with making a decision. See GUIDANCE Consent Key Information and EXAMPLE Key Information.
For example: Consultation with publications about research subjects’ preferences, disease-focused nonprofit groups, patient interest groups, or other researchers/study staff with experience with the specific population. It may also involve directly consulting selected members of the study population.
8.2.f. Ongoing process, new information, and reconsent.
For research that involves multiple or continued interaction with subjects over time, describe the opportunities (if any) that will be given to subjects to ask questions or to change their minds about participating.
Throughout the course of the study, subjects may need to be notified about new information. This might take the form of a verbal or written communication or may require subjects to provide reconsent. When a modification is submitted in which subjects need to be informed about new information, describe the method and process the research team will use to provide this information.
See TIPSHEET Consent Reconsent and Ongoing Subject Communication and GUIDANCE Consent Reconsent and Ongoing Subject Communication for details.
8.3. Electronic presentation of consent information. Will any part of the consent-related information be provided electronically for some, or all of the subjects?
This refers to the use of electronic systems and processes instead of (or in addition to) a paper consent form. For example, an emailed consent form, a passive or an interactive website, graphics, audio, video podcasts. See GUIDANCE Consent Electronic Consent and Documentation of Consent for information about electronic consent requirements at UW.
 No → Skip to question 8.4.
 Yes → Answer questions 8.3.a. through 8.3.e.
8.3.a. Describe the electronic consent methodology and the information that will be provided.
All information materials must be made available to the IRB. Website content should be provided as a Word document. It is considered best practice to give subjects information about multi-page/multi-screen information that will help them assess how long it will take them to complete the process. For example, telling them that it will take about 15 minutes, or that it involves reading six screens or pages.
8.3.b. Describe how the information can be navigated (if relevant).
For example, will the subject be able to proceed forward or backward within the system, or to stop and continue at a later time?
8.3.c. In a standard paper-based consent process, the subjects generally have the opportunity to go through the consent form with study staff and/or to ask study staff about any question they may have after reading the consent form. Describe what will be done, if anything, to facilitate the subject’s comprehension and opportunity to ask questions when consent information is presented electronically. Include a description of any provisions to help ensure privacy and confidentiality during this process.
Examples: hyperlinks, help text, telephone calls, text messages or other type of electronic messaging, video conference, live chat with remotely located study team members.
8.3.d. What will happen if there are individuals who wish to participate but who do not have access to the consent methodology being used, or who do not wish to use it? Are there alternative ways in which they can obtain the information, or will there be some assistance available? If this is a clinical trial, these individuals cannot be excluded from the research unless there is a compelling rationale.
For example, consider individuals who lack familiarity with electronic systems, have poor eyesight or impaired motor skills, or who do not have easy email or internet access.
8.3.e. How will the research team ensure continued accessibility of consent materials and information during the study?
8.3.f. How will additional information be provided to subjects during the research, including any significant new findings (such as new risk information). If this is not an issue, explain why.
8.4. Written documentation of consent. Which of the statements below describe whether documentation of consent will be obtained? NOTE: This question does not apply to screening and recruiting procedures which have already been addressed in question 4.7.
Documentation of consent that is obtained electronically is not considered written consent unless it is obtained by a method that allows verification of the individual’s signature. In other words, saying “yes” by email is rarely considered to be written documentation of consent.
8.4.a. Is written documentation being obtained for:
 None of the research procedures 	→ Use the following table to provide justification then go to question 8.5.
 All of the research procedures	→ Do not complete the following table, go to question 8.4.b.
 Some of the research procedures 	→ Use the following table to identify the procedures for which written documentation of consent will not be obtained from adult subjects.
Table footnotes
1. If the answer is the same for all adult groups or all procedures, collapse the answer across the groups and/or procedures.
8.4.b. Electronic consent signature. For studies in which documentation of consent will be obtained, will subjects use an electronic method to provide their consent signature?
See GUIDANCE Consent Documentation of Consent and GUIDANCE Electronic Consent Signatures for information about options (including REDCap e-signature and the DocuSign system) and any associated requirements.
FDA-regulated studies must use a system that complies with the FDA’s “Part 11” requirements about electronic systems and records. Note that the UW-IT supported DocuSign e-signature system does not meet this requirement.
Having subjects check a box at the beginning of an emailed or web-based questionnaire is not considered legally effective documentation of consent.
 No
 Yes → Indicate which methodology will be used
 UW ITHS REDCap (excludes REDCap Mobile application, which is a separate software application for use with a mobile device for consent when internet service is absent or unreliable)
 Other REDCap installation → Please name the institutional version you will be using (e.g., Vanderbilt, Univ. of Cincinnati) in the following field and provide a completed SUPPLEMENT Other REDCap Installation with your submission.
 UW DocuSign
 Other 	→ Please describe in the following field and provide a signed TEMPLATE Other E-signature Attestation Letter with your submission.
8.4.b.1. Is this method legally valid in the jurisdiction where the research will occur?
NOTE: UW ITHS REDCap (excludes REDCap Mobile application) and UW DocuSign have been vetted for compliance with Washington State and federal laws regarding electronic signatures.
 No
 Yes → What is the source of information about legal validity?
8.4.b.2. Will verification of the subject’s identity be obtained if the signature is not personally witnessed by a member of the study team? Note that this is required for FDA-regulated studies.
See the GUIDANCE Consent Documentation of Consent for information and examples
 No → Provide the rationale for why this is not required or necessary to protect subjects or the integrity of the research. Also, what would be the risks to the actual subject if somebody other than the intended signer provides the consent signature?
 Yes → Describe how subject identity will be verified, providing a non-technical description that the reviewer will understand.
8.4.b.3. How will the requirement be met to provide a copy of the consent information (consent form) to individuals who provide an e-signature?
The copy can be paper or electronic and may be provided on an electronic storage device or via email. If the electronic consent information uses hyperlinks or other websites or podcasts to convey information specifically related to the research, the information in these hyperlinks should be included in the copy provided to the subjects and the website must be maintained for the duration of the entire study.
8.4.c. Barriers to written documentation of consent. There are many possible barriers to obtaining written documentation of consent. Consider, for example, individuals who are functionally illiterate; do not read English well; or have sensory or motor impairments that may impede the ability to read and sign a consent form.
8.4.c.1. Describe the plans (if any) for obtaining written documentation of consent from potential subjects who may have difficulty with the standard documentation process (that is, reading and signing a consent form).
Examples of solutions: Translated consent forms; use of the Short Form consent process; reading the form to the person before they sign it; excluding individuals who cannot read and understand the consent form.
8.5. Non-English-speaking or-reading adult subjects. Will the research enroll adult subjects who do not speak English or who lack fluency or literacy in English?
 No
 Yes → Describe the process that will be used to ensure that the oral and written information provided to them during the consent process and throughout the study will be in a language readily understandable to them and (for written materials such as consent forms or questionnaires) at an appropriate reading/comprehension level.
8.5.a. Interpretation. Describe how interpretation will be provided, and when. Also, describe the qualifications of the interpreter(s) - for example, background, experience, language proficiency in English and in the other language, certification, other credentials, familiarity with the research related vocabulary in English and the target language.
8.5.b. Translations. Describe how translations will be obtained for all study materials (not just consent forms). Also, describe the method for ensuring that the translations meet the UW IRB’s requirement that translated documents will be linguistically accurate, at an appropriate reading level for the participant population, and culturally sensitive for the local in which they will be used.
	Check this box to confirm that before using them with subjects, you will upload in Zipline all translated consent materials that will be provided to subjects in written or electronic form (per HSD policy).
If the IRB determines that your study is greater than minimal risk, or otherwise determines it is required, you will need to work with your translator to provide a TEMPLATE Translation Attestation. If the attestation is required, you will be informed by the IRB during the course of the review.
8.6. [DETERMINATION]  Deception. Will information be deliberately withheld, or will false information be provided, to any of the subjects?
NOTE: “Blinding” subjects to their study group/condition/arm is not considered to be deception, but not telling them ahead of time that they will be subjects to an intervention or about the purpose of the procedure(s) is deception.
 No
 Yes → Describe what information and why.
Example: it may be necessary to deceive subjects about the purpose of the study (describe why).
8.6.a. Will subjects be informed beforehand that they will be unaware of or misled regarding the nature or purposes of the research? (Note: this is not necessarily required.)
 No
 Yes
8.6.b. Will subjects be debriefed later? (Note: this is not necessarily required.)
 No → Provide your reasoning for not debriefing subjects.
 Yes → Describe how and when this will occur. Upload any debriefing materials, including talking points or a script, to Zipline.
8.7. [DETERMINATION] Cognitively impaired adults, and other adults unable to consent. Will such individuals be included in the research?
Examples: individuals with Traumatic Brain Injury (TBI) or dementia; individuals who are unconscious, or who are significantly intoxicated.
 No → Go to question 8.8.
 Yes → Answer the following question.
8.7.a. Rationale. Provide the rationale for including this population.
8.7.b. Capacity for consent/decision making capacity. Describe the process that will be used to determine whether a cognitively impaired individual is capable of consent decision making with respect to the research protocol and setting.
8.7.b.1. If there will be repeated interactions with the impaired subjects over a time period when cognitive capacity could increase of diminish, also describe how (if at all) decision-making capacity will be re-assessed and (if appropriate) consent obtained during that time.
8.7.c. Permission (surrogate consent). If the research will include adults who cannot consent for themselves, describe the process for obtaining permission (“surrogate-consent”) from a legally authorized representative (LAR).
For research conducted in Washington State, see GUIDANCE Consent Diminished and Fluctuating Consent Capacity and Use of a Legally Authorized Representative (LAR) to learn which individuals meet the state definition of “legally authorized representative”.
8.7.d. Assent. Describe whether assent will be required of all, some, or none of the subjects. If some, indicate which subjects will be required to assent and which will not (and why not). Describe any process that will be used to obtain and document assent from the subjects.
8.7.e. Dissent or resistance. Describe how a subject’s objection or resistance to participation (including non-verbal) during the research will be identified, and what will occur in response.
8.8. Research use of human fetal tissue obtained from elective abortion. Federal and UW Policy specify some requirements for the consent process. If you are conducing this type of research, check the boxes to confirm these requirements will be followed.
 Informed consent for the donation of fetal tissue for research use will be obtained by someone other than the person who obtained the informed consent for abortion.
 Informed consent for the donation of fetal tissue for research use will be obtained after the informed consent for abortion.
 Participation in the research will not affect the method of abortion.
 No enticements, benefits or financial incentives will be used at any level of the process to incentivize abortion or the donation of human fetal tissue.
 The informed consent form for the donation of fetal tissue for use in research will be signed by both the woman and the person who obtains the informed consent.
8.9. Consent-related materials. Upload to Zipline all consent scripts/talking points, consent forms, debriefing statements, Information Statements, Short Form consent forms, parental permission forms, and any other consent related materials that will be used. Materials that will be used by a specific site should be uploaded to that site’s Local Site Documents page.
Translations must be submitted and approved before they can be used. However, we strongly encourage you to wait to provide them until the IRB has approved the English versions.
Combination forms: it may be appropriate to combine parental permission with consent, if parents are subjects as well as providing permission for the participation of their children. Similarly, a consent form may be appropriately considered an assent form for older children.
For materials that cannot be uploaded: upload screenshots or written descriptions that are sufficient to enable the IRB to understand the types of data that will be collected and the nature of the experience for the participants. URLs (website addresses) may also be provided, or written descriptions of websites. Examples of materials that usually cannot be uploaded: mobile apps; computer-administered text; licensed and restricted standardized tests.
9. PRIVACY AND CONFIDENTIALITY
9.1. [DETERMINATION]  Privacy protections. Describe the steps that will be taken, if any, to address possible privacy concerns of subjects and potential subjects.
Privacy refers to the sense of being in control of access that others have to ourselves. This can be an issue with respect to recruiting, consenting, sensitivity of the data being collected, and the method of data collection.
Examples:
Many subjects will feel a violation of privacy if they receive a letter asking them to participate in a study because they have ____ medical condition, when their name, contact information, and medical condition were drawn from medical records without their consent. Example: the IRB expects that “cold call” recruitment letters will inform the subject about how their information was obtained.
Recruiting subjects immediately prior to a sensitive or invasive procedure (e.g., in an outpatient surgery waiting room) will feel like an invasion of privacy to some individuals.
Asking subjects about sensitive topics (e.g. details about sexual behavior) may feel like an invasion of privacy to some individuals.
9.2. [DETERMINATION]  Identification of individuals in publications and presentations. Will potentially identifiable information about subjects be used in publications and presentations, or is it possible that individual identities could be inferred from what is planned to be published or presented?
 No
 Yes → Will subject consent be obtained for this use?
 Yes
 No → Describe the steps that will be taken to protect subjects (or small groups of subjects) from being identifiable.
9.3. [DETERMINATION]  State mandatory reporting. Each state has reporting laws that require some types of individuals to report some kinds of abuse, and medical conditions that are under public health surveillance. These include:
Child abuse
Abuse, abandonment, neglect, or financial exploitation of a vulnerable adult
Sexual assault
Serious physical assault
Medical conditions subject to mandatory reporting (notification) for public health surveillance
Are you or a member of the research team likely to learn of any of the above events or circumstances while conducting the research AND feel obligated to report it to state authorities?
 No
 Yes → The UW IRB expects subjects to be informed of this possibility in the consent form or during the consent process, unless you provide a rationale for not doing so:
9.4. [DETERMINATION]  Retention of identifiers and data. Check the box below to indicate assurance that any identifiers (or links between identifiers and data/specimens) and data that are part of the research records will not be destroyed until after the end of the applicable records retention requirements (e.g. Washington State; funding agency or sponsor; Food and Drug Administration). If it is important to say something about destruction of identifiers (or links to identifiers) in the consent form, state something like “the link between your identifier and the research data will be destroyed after the records retention period required by state and/or federal law.”
See the “Research Data” sections of the following website for UW Records management for the Washington State research records retention schedules that apply in general to the UW (not involving UW Medicine data): http://f2.washington.edu/fm/recmgt/gs/research?title=R
See the “Research Records and Data” information in Section 8 of this document for the retention schedules for UW Medicine Records: https://www.uwmedicine.org/recordsmanagementuwm-records-retention-schedule.pdf
 Confirm
9.5. [DETERMINATION]  Certificates of Confidentiality. Will a federal Certificate of Confidentiality be obtained for the research data? NOTE: Answer “No” if the study is funded by NIH or the CDC, because most NIH-funded and CDC-funded studies automatically have a Certificate.
 No
 Yes
9.6. [DETERMINATION]  Data and specimen security protections. Identify the data classifications and the security protections that will be provided for all sites where data will be collected, transmitted, or stored, referring to the GUIDANCE Data and Security Protections for the minimum requirements for each data classification level. It is not possible to answer this question without reading this document. Data security protections should not conflict with records retention requirements.
9.6.a. Which level of protections will be applied to the data and specimens? If more than one level will be used, describe which level will apply to which data and which specimens and at which sites.
9.6.b. Use this space to provide additional information, details, or to describe protections that do not fit into one of the levels. If there are any protections within the level listed in 9.6.a which will not be followed, list those here, including identifying the sites where this exception will apply. For example, if you intend to store subject identifiers with study data (not permitted under requirement U9 for Risk Levels 3-5), then indicate this in the box below (e.g., “We will not adhere to requirement U9 for screening data”).
10. RISK / BENEFIT ASSESSMENT
10.1. [DETERMINATION]  Anticipated risks. Describe the reasonably foreseeable risks of harm, discomforts, and hazards to the subjects and others of the research procedures. For each harm, discomfort, or hazard:
Describe the magnitude, probability, duration, and/or reversibility of the harm, discomfort, or hazard, AND
Describe how the risks will be reduced or managed. Do not describe data security protections here, these are already described in question 9.6.
Consider possible physical, psychological, social, legal, and economic harms, including possible negative effects on financial standing, employability, insurability, educational advancement or reputation. For example, a breach of confidentiality might have these effects.
Examples of “others”: embryo, fetus, or nursing child; family members; a specific group.
Ensure applicable risk information from any Investigator Brochures, Drug Package Inserts, and/or Device Manuals is included in your description.
Do not include the risks of non-research procedures that are already being performed.
If the study design specifies that subjects will be assigned to a specific condition or intervention, then the condition or intervention is a research procedure - even if it is a standard of care.
Examples of mitigation strategies: inclusion/exclusion criteria; taking blood samples to monitor something that indicates drug toxicity.
As with all questions on this application, you may refer to uploaded documents.
10.2. [DETERMINATION]  Reproductive risks. Are there any risks of the study procedures to subjects or partner of subjects related to pregnancy, fertility, lactation or effects on a fetus or neonate?
Examples: direct teratogenic effects; possible germline effects; effects on fertility; effects on a woman’s ability to continue a pregnancy; effects on future pregnancies.
 No → Go to question 10.3.
 Yes → Answer the following questions:
10.2.a. Risks. Describe the magnitude, probability, duration and/or reversibility of the risks.
10.2.b. Steps to minimize risk. Describe the specific steps that will be taken to minimize the magnitude, probability or duration of these risks.
Examples: inform the subjects about the risks and how to minimize them; require a pregnancy test before and during the study; require subjects to use contraception; advise subjects about banking of sperm and ova.
If the use of contraception will be required, describe the allowable methods and the time period when contraception must be used.
10.2.c. Pregnancy. Describe what will be done if a subject (or a subject’s partner) becomes pregnant.
For example; will subjects be required to immediately notify study staff, so that the study procedures can be discontinue or modified, or for a discussion of risks, and/or referrals or counseling?
10.3. [DETERMINATION]  MRI risk management. A rare but serious adverse reaction called nephrogenic systemic fibrosis (NSF) has been observed in individuals with kidney disease who received gadolinium-based contrast agents (GBCAs) for the scans. Also, a few healthy individuals have a severe allergic reaction to GBCAs.
10.3.a. Use of gadolinium. Will any of the MRI scans involve the use of a gadolinium-based contrast agent (GBCA)?
 No
 Yes → Which agents will be used? Check all that apply.
10.3.a.1. The FDA has concluded that gadolinium is retained in the body and brain for a significantly longer time than previously recognized, especially for linear GBCAs. The health-related risks of this longer retention are not yet clearly established. However, the UW IRB expects researchers to provide a compelling justification for using a linear GBCA instead of a macrocyclic GBCA, to manage the risks associated with GBCAs.
Describe why it is important to use a GBCA with the MRI scan(s). Describe the dose that will be used and (if it is more than the standard clinical dose recommended by the manufacturer) why it is necessary to use a higher dose. If a linear GBCA will be used, explain why a macrocyclic GBCA cannot be used.
10.3.a.2. Information for subjects. Confirm by checking this box that subjects will be provided with the FDA-approved Patient Medication Guide for the GBCA being used in the research or that the same information will be inserted into the consent form.
 Confirmed
10.3.b. Who will (1) calculate the dose of GBCA; (2) prepare it for injection; (3) insert and remove the IV catheter; (4) administer the GBCA; and (5) monitor for any adverse effects of the GCBA? Also, what are the qualifications and training of these individual(s)?
10.3.c. Describe how the renal function of subjects will be assessed prior to MRI scans and how that information will be used to exclude subjects at risk for NSF.
10.3.d. Describe the protocol for handling a severe allergic reaction to the GBCA or any other medical event/emergency during the MRI scan, including who will be responsible for which actions.
10.4. [DETERMINATION]  Unforeseeable risks. Are there any research procedures that may have risks that are currently unforeseeable?
Example: using a drug that hasn’t been used before in this subject population.
 No
 Yes → Identify the procedures.
10.5. Subjects who will be under regional or general anesthesia. Will any research procedures occur while patients are under general or regional anesthesia, or during the 3 hours preceding general or regional anesthesia (supplied for non-research reasons)?
 No
 Yes → Check all the boxes that apply.
 Administration of any drug for research purposes
 Inserting an intra-venous (central or peripheral) or intra-arterial line for research purposes
 Obtaining samples of blood, urine, bone marrow or cerebrospinal fluid for research purposes
 Obtaining a research sample from tissue or organs that would not otherwise be removed during surgery.
 Administration of a radio-isotope for research purposes**
 Implantation of an experimental device
 Other manipulations or procedures performed solely for research purposes (e.g., experimental liver dialysis, experimental brain stimulation)
If any of the boxes are checked:
Provide the name and institutional affiliation of a physician anesthesiologist who is a member of the research team or who will serve as a safety consultant about the interactions between the research procedures and the general or regional anesthesia of the subject-patients. If the procedures will be performed at a UW Medicine facility or affiliate, the anesthesiologist must be a UW faculty member, and  the Vice Chair of Clinical Research in the UW Department of Anesthesiology and Pain Medicine must be consulted in advance for feasibility, safety and billing.
** If the box about radio-isotopes is checked, the study team is responsible for informing in advance all appropriate clinical personnel (e.g., nurses, technicians, anesthesiologists, surgeons) about the administration and use of the radio-isotope, to ensure that any personal safety issues (e.g., pregnancy) can be appropriately addressed. This is a condition of IRB approval.
10.6. Data and Safety Monitoring. A Data and Safety Monitoring Plan (DSMP) is required for clinical trials (as defined by NIH). If required for this research, or if there is a DSMP for the research regardless of whether it is required, upload the DSMP to Zipline. If it is embedded in another document being uploading (for example, a Study Protocol) use the text box below to name the document that has the DSMP. Alternatively, provide a description of the DSMP in the text box below. For guidance on developing a DSMP, see the ITHS webpage on Data and Safety Monitoring Plans.
10.7. Un-blinding. If this is a double-blinded or single-blinded study in which the participant and/or relevant study team members do not know the group to which the participant is assigned, describe the circumstances under which un-blinding would be necessary, and to whom the un-blinded information would be provided.
10.8. Withdrawal of participants. If applicable, describe the anticipated circumstances under which participants will be withdrawn from the research without their consent. Also, describe any procedures for orderly withdrawal of a participant, regardless of the reason, including whether it will involve partial withdrawal from procedures and any intervention but continued data collection or long-term follow-up.
10.9. [DETERMINATION]  Anticipated direct benefits to participants. If there are any direct research-related benefits that some or all individual participants are likely to experience from taking part in the research, describe them below:
Do not include benefits to society or others, and do not include subject payment (if any). Examples: medical benefits such as laboratory tests (if subjects receive the results); psychological resources made available to participants; training or education that is provided.
10.10. [DETERMINATION]  Return of individual research results.
In this section, provide your plans for the return of individual results. An “individual research result” is any information collected, generated or discovered in the course of a research study that is linked to the identity of a research participant. These may be results from screening procedures, results that are actively sought for purposes of the study, results that are discovered unintentionally, or after analysis of the collected data and/or results has been completed.
See the GUIDANCE Return of Individual Results for information about results that should and should not be returned, validity of results, the Clinical Laboratory Improvement Amendment (CLIA), consent requirements and communicating results.
10.10.a. Is it anticipated that the research will produce any individual research results that are clinically actionable?
“Clinically actionable” means that there are established therapeutic or preventive interventions or other available actions that have the potential to change the clinical course of the disease/condition, or lead to an improved health outcome.
In general, every effort should be made to offer results that are clinically actionable, valid and pose life-threatening or severe health consequences if not treated or addressed quickly. Other clinically actionable results should be offered if this can be accomplished without compromising the research.
 No
 Yes → Answer the following questions (10.10.a.1 through 10.10.a.3.)
10.10.a.1. Describe the clinically actionable results that are anticipated and explain which results, if any, could be urgent (i.e. because they pose life-threatening or severe health consequences if not treated or addressed quickly).
Examples of urgent results include very high calcium levels, highly elevated liver function test results, positive results for reportable STDs.
10.10.a.2. Explain which of these results will be offered to subjects.
10.10.a.3. Explain which results will not be offered to subjects and provide the rationale for not offering these results.
Reasons not to offer the results might include:
There are serious questions regarding validity or reliability
Returning the results has the potential to cause bias
There are insufficient resources to communicate the results effectively and appropriately
Knowledge of the result could cause psychosocial harm to subjects
10.10.b. Is there a plan for offering subjects any results that are not clinically actionable?
Examples: non-actionable genetic results, clinical tests in the normal range, experimental and/or uncertain results.
 No
 Yes → Explain which results will be offered to subjects and provide the rationale for offering these results.
10.10.c. Describe the validity and reliability of any results that will be offered to subjects.
The IRB will consider evidence of validity such as studies demonstrating diagnostic, prognostic, or predictive value, use of confirmatory testing, and quality management systems.
10.10.d. Describe the process for communicating results to subjects and facilitating understanding of the results. In the description, include who will approach the participant with regard to the offer of results, who will communicate the result (if different), the circumstances, timing, and communication methods that will be used.
10.10.e. Describe any plans to share results with family members (e.g. in the event a subject becomes incapacitated or deceased).
10.10.f. Check the box to indicate that any plans for return of individual research results have been described in the consent document. If there are no plans to provide results to participants, this should be stated in the consent form.
See the GUIDANCE Return of Individual Results for information about consent requirements.
 Confirmed
10.11. Commercial products or patents. Is it possible that a commercial product or patent could result from this study?
 No
 Yes → Describe whether subjects might receive any remuneration/compensation and, if yes, how the amount will be determined.
 11. ECONOMIC BURDEN TO PARTICIPANTS
11.1. Financial responsibility for research-related injuries. Answer this question only if the lead researcher is not a UW student, staff member, or faculty member whose primary paid appointment is at the UW.
For each institution involved in conducting the research: Describe who will be financially responsible for research-related injuries experienced by subjects, and any limitations. Describe the process (if any) by which participants may obtain treatment/compensation.
11.2. Costs to subjects. Describe any research-related costs for which subjects and/or their health insurance may be responsible (examples might include: CT scan required for research eligibility screening; co-pays; surgical costs when a subject is randomized to a specific procedure; cost of a device; travel and parking expenses that will not be reimbursed).
 12. RESOURCES
12.1. [DETERMINATION]  Faculty Advisor. (For researchers who are students or residents.) Provide the following information about the faculty advisor.
Advisor’s name
Your relationship with your advisor (for example: graduate advisor; course instructor)
Your plans for communication/consultation with your advisor about progress, problems, and changes.
12.2. UW Principal Investigator Qualifications. Upload a current or recent Curriculum Vitae (CV), Biosketch (as provided to federal funding agencies), or similar document to the Local Site Documents page in Zipline. The purpose of this is to address the PI’s qualifications to conduct the proposed research (education, experience, training, certifications, etc.).
For help with creating a CV, see http://adai.uw.edu/grants/nsf_biosketch_template.pdf and https://medicine.uw.edu/faculty/academic-human-resources/curriculum-vitae-cv
 The CV will be uploaded.
12.3. UW Study team qualifications. Describe the qualifications and/or training for each UW study team member to fulfill their role on the study and perform study procedures. (You may be asked about non-UW study team members during the review; they should not be described here.) You may list these individuals by name, however if you list an individual by name, you will need to modify this application if that individual is replaced. Alternatively, you can describe study roles and the qualifications and training the PI or study leadership will require for any individual who might fill that role. The IRB will use this information to assess whether risks to subjects are minimized because study activities are being conducted by properly qualified and trained individuals.
Describe: The role (or name of person), the study activities they will perform, and the qualifications or training that are relevant to performing those study activities.
12.4. Study team training and communication. Describe how it will be ensured that each study team member is adequately trained and informed about the research procedures and requirements (including any changes) as well as their research-related duties and functions.
 There is no study team
13. OTHER APPROVALS, PERMISSIONS, AND REGULATORY ISSUES
13.1. [DETERMINATION]  Approvals and permissions. Identify any other approvals or permissions that will be obtained. For example: from a school, external site/organization, funding agency, employee union, UW Medicine clinical unit.
Do not attach the approvals and permissions unless requested by the IRB.
13.2. Financial Conflict of Interest. Does any UW member of the team have ownership or other Significant Financial Interest (SFI) with this research as defined by UW policy GIM 10?
 No
 Yes → Has the Office of Research made a determination regarding this SFI as it pertains to the proposed research?
 No → Contact the Office of Research (206.616.0804, research@uw.edu) for guidance on how to obtain the determination.
 Yes → Upload the Conflict Management Plan for every UW team member who has a FCOI with respect to the research, to Zipline. If it is not yet available, use the text box to describe whether the Significant Financial Interest has been disclosed already to the UW Office of Research and include the FIDS Disclosure ID if available.
APPLICATION IRB Protocol
AI Ready and Exploratory Atlas for Diabetes Insights (AI-READI)
University of Washington
Multiple dates via email
The newly-funded multisite NIH study requires IRB approval within 6 months of the funding receipt.
The study will collect a cross-sectional dataset of 4600 people across the US from diverse racial/ethnic groups who are either 1) healthy, or 2) belong in one of the three stages of diabetes severity (pre-diabetes/diet controlled, oral medication and/or non-insulin-injectable medication controlled, or insulin dependent), forming a total of four groups of patients. Clinical data (social determinants of health surveys/PhenX surveys, continuous glucose monitoring data, biomarkers, genetic data, retinal imaging, cognitive testing, etc.) will be collected. The purpose of this project is data generation to allow future creation of artificial intelligence/machine learning (AI/ML) algorithms aimed at defining disease trajectories and underlying genetic links in different racial/ethnic cohorts. A smaller subgroup of participants will be invited to come for a follow-up visit in year 4 of the project (longitudinal arm of the study). Data will be placed in an open-source repository and samples will be sent to the study sample repository and used for future research.
Cross-sectional and longitudinal observational study with data and sample collections.
Check all that apply	Descriptor
Class project or other activity whose purpose is to provide an educational experience for the researcher (for example, to learn about the process or methods of doing research).
Part of an institution, organization, or program’s own internal operational monitoring.
Improve the quality of service provided by a specific institution, organization, or program.
Designed to expand the knowledge based of a scientific discipline or other scholarly field of study, and produce results that:
Are expected to applicable to a larger population beyond the site of data collection or the specific subjects studied, or
Are intended to be used to develop, test, or support theories, principles, and statements of relationships, or to inform policy beyond the study.
Focus directly on the specific individuals about whom the information or biospecimens are collected through oral history, journalism, biography, or historical scholarship activities, to provide an accurate and evidence-based portrayal of the individuals.
A quality improvement or program improvement activity conducted to improve the implementation (delivery or quality) of an accepted practice, or to collect data about the implementation of the practice for clinical, practical, or administrative purposes. This does not include the evaluation of the efficacy of different accepted practices, or a comparison of their efficacy.
Public health surveillance activities conducted, requested, or authorized by a public health authority for the sole purpose of identifying or investigating potential public health signals or timely awareness and priority setting during a situation that threatens public health.
Preliminary, exploratory or research development activities (such as pilot and feasibility studies, or reliability/validation testing of a questionnaire).
Expanded access use of a drug or device not yet approved for this purpose.
Use of a Humanitarian Use Device.
Other. Explain:
The ability to understand and affect the course of complex, multisystem diseases has been limited by a lack of well-designed, high quality, large, and inclusive multimodal datasets. We propose to create such a dataset allowing ML approaches to provide critical insights int the endemic condition, type 2 diabetes mellitus.

Our approach is to collect a cross-sectional dataset of 4,600 people across the US with dual
balancing for self-reported race/ethnicity and four stages of diabetes severity. Building balanced
training datasets is critical for the development of unbiased ML models. Approximately 10% of the participants will be invited to return for follow-up data collection. Thus, rather than targeting the demographic distribution of the US population, we intentionally will recruit
equal numbers of four racial/ethnic groups. The same rationale applies for balancing diabetic severity. AI/ML ready data will include physical measurements, medical history, motor vehicle driving history, health surveys, continuous glucose monitoring, serological testing for endocrine, cardiac, and renal biomarkers, genome-wide polymorphism assessment, visual function testing, retinal imaging, ECG, cognitive testing, 24-hr activity monitoring, and environmental sensing.

The dataset will be specifically designed to permit downstream pseudotime manifold analysis. This approach has been utilized in developmental biology successfully to identify differentiation pathways from complex gene expression datasets of individual cells. An expanded approach may be used to define disease trajectories by collecting complex, multimodal data from participants with differing disease severity. This approach relies on dimensionality reduction, is agnostic to existing classification criteria or biases, and can be used to reconstruct a temporal atlas of pathogenesis and healing.
The study team includes a diverse group of investigators from multiple institutions across the US. Data collection will take place at the University of Washington (UW), the University of Alabama Birmingham (UAB), and the University of California San Diego (UCSD). In the process of preparing the proposal for NIH funding and in planning the study protocol, we have had numerous meetings as well as asynchronous communication to assess the feasibility and safety of the protocol. Investigators at all sites have had prior experience with the proposed study procedures. The ophthalmic imaging devices are non-invasive and do not pose significant risk to the participants. We have teams of study coordinators and research personnel with extensive prior experience conducting research study visits and experience working with participants from an array of diverse backgrounds.
Check all that apply	Type of Research	Supplement Name and Link
Department of Defense
The research involves Department of Defense funding, facilities, data, or personnel.	SUPPLEMENT Department of Defense
Department of Energy
The research involves Department of Energy funding, facilities, data, or personnel.	SUPPLEMENT Department of Energy
Drug, biologic, botanical, supplement
Procedures involve the use of any drug, biologic, botanical or supplement, even if the item is not the focus of the proposed research.	SUPPLEMENT Drugs
Emergency exception to informed consent
Research that requires this special consent waiver for research involving more than minimal risk.	SUPPLEMENT Exception from Informed Consent for Emergency Research (EFIC)
Genomic data sharing
Genomic data are being collected and will be deposited in an external database (such as the NIH dbGaP database) for sharing with other researchers, and the UW is being asked to provide the required certification or to ensure that the consent forms can be certified.	SUPPLEMENT Genomic Data Sharing
Medical device
Procedures involve the use of any medical device, even if the device is not the focus of the proposed research, except when the device is FDA-approved and is being used through a clinical facility in the manner for which it is approved.	SUPPLEMENT Devices
Multi-site or collaborative study
The UW IRB is being asked to review on behalf of one or more non-UW institutions in a multi-site or collaborative study.	SUPPLEMENT Multi-site or Collaborative Research
Non-UW Individual Investigators
The UW IRB is being asked to review on behalf of one or more non-UW individuals who are not affiliated with another organization for the purpose of the research.	SUPPLEMENT Non-UW Individual Investigators
Other REDCap Installation Attestation for Electronic Consent
The research will use a non-UW installation of REDCap for conducting and/or documenting informed consent.	SUPPLEMENT Other REDCap Installation
None of the above.
Adult patients will be recruited into one of four groups: 1) healthy/no diabetes, 2) pre-diabetes/borderline diabetes/diet-controlled diabetes, 3) oral medication and/or non-insulin injectable medication  controlled type 2 diabetes, or 4) insulin dependent type 2 diabetes. We aim to recruit approximately 1000 patients into each of the four groups. Patients will be recruited from University of Washington (UW), University of California at San Diego (UCSD), and University of Alabama at Birmingham (UAB). The study aims to recruit 1,000 subjects from each of the following racial and ethnic groups: white, Asian, Hispanic, and Black. Subjects will be age- and sex-matched within and between groups.

Note; We are submitting an NIH supplement to obtain funding for translation of all study materials to enable recruitment and participation of Spanish-speaking participants who are no fluent in English. Hispanic participants enrolled in the meantime will need to be fluent in English to participate.
Adults (≥ 40 years old)
Patients with and without type 2 diabetes
Able to provide consent
Must be able to read and speak English
Adults older than 85 years of age
Pregnancy
Gestational diabetes
Type 1 diabetes
Check all that apply	Population	Worksheet Name and Link
Fetuses in utero	WORKSHEET Pregnant Women
Neonates of uncertain viability	WORKSHEET Neonates
Non-viable neonates	WORKSHEET Neonates
Pregnant women	WORKSHEET Pregnant Women
Group name/description	Maximum desired number or individuals (or other subject unit) who will complete the research
Provide numbers for the site(s) reviewed by the UW IRB and for the study-wide total number; example: 20/100
No type 2 diabetes	1000
Type 2 diabetes controlled by diet/lifestyle and/or has a diagnosis of prediabetes or borderline diabetes	1000
Type 2 diabetes controlled by oral and/or non-insulin injectable medications	1000
Type 2 diabetes dependent on insulin	1000
The research sites have been chosen based on strong prior experience with conducting longitudinal clinical research studies, experience with ophthalmic study procedures, and available collaborators with commitment to participation. All three sites will recruit people of diverse backgrounds. However, based on local/regional demographics, UW will concentrate in recruiting Asian>Hispanic>Black participants, while the University of California at San Diego site will prioritize Hispanic>Asian>Black participants, and the University of Alabama at Birmingham site will prioritize Black>Hispanic>Asian participants. All sites will recruit white participants. All race and ethnicity information will be self-reported at the time of screening and enrollment.
UCSD – State of California requires inclusion of the Bill of Rights in the ICF.
UCSD has a health Data Oversight Committee (HDOC) that oversees data sharing with other institutions. For sharing research data with another academic (i.e., non-profit/non-commercial) institution, this is typically a straightforward process that is expedited. Approval is contingent upon IRB approval and will be sought immediately after IRB approval is obtained.
We will first screen for eligible participants using medical records data search from each institution including ICD diagnosis codes, race, and age.

Information about the study may be provided to potential subjects as follows:

Mailed recruitment letters
Over the 3+ years of study recruitment, approximately 50,000 invitation letters from each site will be sent by both mail and email to participants identified through the electronic medical record search as meeting study eligibility criteria.

UW and UCSD will provide a list of potential participants (names, addresses, age, sex, race/ethnicity, phone number, email addresses, diabetes status) identified through the electronic medical record search to UAB. UAB will add information for their own eligible patients. Using the combined UW, UCSD, and UAB dataset, UAB will select a stratified sample of potential participants based on the sampling scheme. Sampling is necessary to ensure steady enrollment through the recruitment process and to not overwhelm the study staff by inviting all potential participants at one. Once a sample has been selected, a dataset of patient names, clinical information, and patient contact information will be uploaded to a secure site at UAB for Oregon Health and Science University (OHSU) to download to an OHSU server. Although the data will be housed in the UW REDCap system, OHSU is responsible for uploading the combined UW, UCSD, and UAB dataset into REDCap for the purpose of creating patient-specific REDCap survey links and the distribution of REDCap email invitations. After generating patient-specific REDCap survey links, the complete dataset (now including patient-specific REDCap survey links) will be sent back to UAB to send participant information to Adapt Inc., an SSAE16 SOC SecurityCertified and HIPAA-compliant printing and fulfillment service, to send out the recruitment letters. It is necessary for the letters to contain the patient-specific REDCap survey links. The aforementioned process will occur on a recursive basis with UAB coordinating the selection of participants to achieve balancing of race/ethnicity, sex, and diabetes severity until the recruitment goals have been achieved.



Participants may also be E-mailed the recruitment letter.

Follow-up phone calls will be made by research coordinators to those
Who did not respond to emails and/or mailings
Who initiated contact with the study teams.
AI-READI Website
Participants will also be directed to the AI-READI website (https://aireadi.org, under construction) where they can learn more about the study, including the study FAQ.

If we are having recruitment difficulties, we may rely on the following options:
Records from specific clinics or satellite clinics affiliated with each institution
Referrals from community centers
Referrals from clinicians/clinic staff from each institution
Databases (e.g., Leaf), existing research studies and registries
Self-reports of diabetes status and diagnosis

As social media messages, flyers, and advertisements are not the primary methods of planned recruitment, the information for them have not yet been developed. If we need to rely on these sources, the materials will be submitted to IRB at a later date. However, the general information will be similar to the information in the vitiation letter sent by mail or email, study website (https://aireadi.org) and telephone talking points.
Letters sent out by email or emails
Follow-up phone conversations
Follow-up texts
Websites

If we are having recruitment difficulties, we may rely on any of the following options:
Social media messages
Flyers
Advertisements

These materials, if needed, will be developed at a later date and submitted to the IRB.
As of now, our primary method of recruitment is based on electronic medical records search and contacting them via email, letter, and follow-up phone calls (please refer to 4.1). However, we may need to utilize other methods of recruitment listed in 4.1 if our recruitment rate is too low. In case we start relying on referrals from other clinics, registries, other community centers, and because we are trying to recruit a large number of participants, it is possible that the recruitment happens by word of mouth and we may recruit some participants who may be friends, colleagues, families, and relatives of our study staff. In addition, investigators who are also clinicians may have patients in their clinical practices who are eligible for participation.
There will be a $200 stipend for participation, which will be provided when we receive the final data from participants and return of the study devices. The timing of the payment will typically be at least 2 weeks after the study visit, taking into account 10 days of home-based data collection using the devices provided in the study and additional time for devices to be returned to the study team vial mail. The payment will not be prorated. The amount may be changed in future years depending on the study funding. We will notify IRB in case there is any change in the stipend amount.

Study will cover reasonable costs for parking, public transit, or rideshare fees (e.g., Uber or Lyft)  related to the study visit.

The same type and amount of compensation will be given to patients who are asked to come back in Year 4 of the study.
N/A
N/A
Electronic medical records will be accessed for pre-screening purposes (i.e., to determine age, diabetes status, and other eligibility criteria). Contact information of the participants (phone number, address, email) will be used to contact potential prospective participants. Data gathered during the screening process may be retained as part of the study data to facilitate future screening efforts and/or to contribute to study procedures and reduce duplication efforts.
tudy visit will take approximately 3-4 hours. If requested by subjects, the visit can be performed over a couple of days instead of a single visit.

Sample and Data Collection:
Following consent, we will collect demographic information, basic measurements (weight and height), and vitals. At the earliest possible time, we will collect blood and urine for testing. We will collect approximately 50-60 ml of blood (about 3-4 spoonfuls). In the event that the original blood and/or urine sample quantities were insufficient for the planned testing, we may bring the participant(s) back for an additional blood draw and/or urine sample collection. The list of tests based on blood and urine samples is in the section 5.4.

Participants will not need to be fasting. We will ask the participant about their last per oral food or drink intake at the time of the study visit. A small snack and drink will be available for participants at any time during the study or after the visit.

Questionnaires:
Participants will fill out the following questionnaires:
Center for Epidemiological Studies – Depression (CES-D-10) survey
CES-10 survey
Diabetes self-care activities measurements (SDSCA)
Diet questionnaire
Medication Log
General Health Survey
Problem Areas In Diabetes questionnaire (PAID-5)
Social Determinants of Health questionnaires (PhenX selection of surveys)
Visual Impairment and Access to Eye Care Questionnaires
Alcohol, Smoking, Marijuana and Vaping Use and History Questionnaire (UW and UCSD sites) OR Alcohol, Smoking and Vaping use and History Questionnaire (UAB site only)

To reduce the visit time, questionnaires may be filled out electronically or on paper before or after the in-person visit. Paper-based questionnaires may be brought to the study visit or mailed back with the devices.

Testing:
Memory and cognitive function will be assess with the MoCA cognitive test (Montreal Cognitive Assessment). We will obtain a 12-lead ECG and perform the monofilament test (https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2775618/).

Vision Exam and Retinal Imaging:
Visual function testing will be performed as follows:
Auto-refraction
Best-corrected visual acuity
Best-corrected low luminance visual activity
Contrast sensitivity
Low luminance contrast sensitivity

To perform eye imaging, participants will receive on or two types of eye drops (see Drugs) for pupil dilation. Participants may receive numbing drops (see Drugs) if needed.

Optical coherence tomography (OCT), fundus photography, fluorescence lifetime imaging ophthalmoscopy (FLIO), and optical coherence tomography angiography (OCTA) will be collected using multiple devices. All imaging will be performed using non-invasive imaging methods.

Wearable and at-home devices:
Study participants will have a continuous glucose monitoring device (CGM) applied to their abdomen and will be given a fitness tracking device for activity monitoring. We will use Dexcom CGM devices that collect glucose data over a 10-day period. Participants who are currently wearing a Dexcom G6 will have two options. The first option is to place an additional study Dexcom G6 sensor at least 3 inches away from their current sensor and injection site, wear it for 10 days and send back the transmitter to study staff to download data. The second option does not place an additional study sensor and instead uses the participant’s current Dexcom G6 device to get data through Clarity. The study will invite the participant with a unique code to share their Dexcom Clarity data with the study clinic. At the end of the 10 days, study staff will log into Dexcom Clarity and download the participant’s data. Patients who already wear Dexcom G6 sensor at the time of the visit will be given an option to allow download of the previously collected glucose values for up to 90 days prior to the study visit. After we capture the 10 days of data, up to 100 days total of glucose data will be collected through the application for subjects who agree to the additional data download, and the study clinic will terminate data sharing with the participant. Sharing with the study clinic will not disrupt their current sharing with any other clinic.


Study participants will receive a small sensor (for environmental data monitoring) and place it at home for data collection for 10 days, and will be asked to return it with the CGM device by mail in a box provided by the study team. The participants will be sent home with instruction for the removal of the CGM device and return the devices back to the research site.

At-home Participation:
The participants will wear the CGM and fitness tracker for 10 days. The participants will also place the environmental sensor in their homes for 10 days. At the end of the 10-day period, participants will need to mail all three devices back so we can obtain the data. Subjects will receive a reminder to return the CGM, fitness tracker, and environmental sensor via phone, text, or email, per their stated preference.
Accessing Medical Records:
We will be extracting medical records to obtain data on general health, eye health, medical and surgical history, diabetes, medications, laboratory records, and any previous medical imaging such as eye imaging. We will also collect information about any car accidents they may have experienced and traffic tickets the participant has received in the last 3 years.

Participants will be asked to complete an exit survey.

Future Participation:
Some study participants who were recruited during Year 1 and 2 (we intend to invite 10% of the study population) will be asked to come for additional follow up visit in Year 4. The follow-up visit would be similar to the visit described above, such as requesting wearing of CGM, fitness tracking device, and environmental sensor. No new variables will be added for collection but not all variables will be collected due to budget constraints.
List of variables is uploaded and listed below:

List of variables
Data will be obtained from subjects, devices they will be asked to use, and their specimens, as well as medical records. Driving and accident records will be obtained from the Department of Licensing (no additional approvals are required beyond the IRB approval). Glucose monitoring data for subjects already using specific CGM will be given an option to provide access to glucose data for up to 90 days prior to the study visit, in addition to the study data collection, for a total of up to 100 days, through the CGM vendor’s Clarity application. The additional data collection is optional.
Name, MRN, DOB, age, sex, race/ethnicity, address, phone number, driver’s license number, and email
Age (DOB), sex, race/ethnicity, ZIP code (5 digits), driver’s license number
All data.
Collaborators at all institutions involved in the study will collect, analyze, and interpret data; collect specimens; be involved in preparation of manuscripts, presentations, grants, and other materials stemming from this research.
UAB:
Cynthia Owsley
Nathan E. Miles Endowed Chair of Ophthalmology
University of Alabama, Birmingham
cynthiaowsley@uabmc.edu
205.325.8635

JHU:
T Y. Alvin Liu
Assistant Professor of Ophthalmology
Johns Hopkins University
tliu25@jhmi.edu
410.287.1874

OHSU:
Shannon McWeeney
Professor of Medicine, School of Medicine
Oregon Health Sciences University
mcweeney@ohsu.edu
503.494.8347

CALMI:
Bhavesh Patel
Assistant Research Professor
California Medical Innovations Institute
bpatel@calmi2.org

UCSD:
Linda Zangwill (Contact PI)
Professor of Ophthalmology
UC San Diego
lzangwill@health.ucsd.edu

Sally Baxter (Co-PI)
Assistant Professor of Ophthalmology and Biomedical Informatics
UC San Diego
s1baxter@health.ucsd.edu


Stanford:
Michael Snyder
Stanford W. Ascherman, Professor in Genetics
Stanford University
mpsnyder@stanford.edu
650.723.4668
See 4.6. and 5.4 sections. PHI are required for the study purposes, both for initial screening/recruitment as well for study data collection. We will collect PHI related to participant’s general health, diabetes, cognitive functions, and eye health through prospective data collection as well as review of EHRs as a part of the dataset.
PHI will be collected by directly viewing electronic medical records and direct attainment from the subjects during the study procedures as described.
All PHI accessed or collected following enrollment.
PHI obtained for the prescreening activities
Subject codes, contact information/identifiers, deidentified data, unused specimen
Data will be shared with the investigators in this project who are part of developing software for de-identifying, data uploading and data sharing.

The data includes responses to questionnaires, vitals and demographic information, cognitive testing, motor vehicle collision involvement, results from blood and urine tests, retinal images, read-out from wearables (CGM and fitness tracker), read-out from home environmental sensor, ECG, and cardiology, peripheral neuropathy, eye health, and endocrinology measures, genomic sequencing, raw imaging data such as retina scans. This will also include any previous relevant medical history that was obtained.

Biobank blood specimens will be stored at UAB as a repository. During this project we will create an infrastructure to decide who is allowed to request and receive the specimen for future research.

There will be two databases for data distribution:
The first will be deidentified, which will be publicly available.
The second one has some identifiable or sensitive information: genomic data and race/ethnicity, and will require data user agreement between the researchers to obtain the data and approval process developed by the project committee.
Principal investigators and lead co-investigators at each participating institution. All the research staff under the PI and lead site PIs will be able to see data under the supervision of the PIs and Site PIs..
Manuscript and presentation preparations, grant preparations, future research studies, future AI/ML algorithms, repository/data sharing platform creations, future research using the stored specimen and data.
Language is included in the ICF.
Emails, texts, phone calls, letters will be used for communication about upcoming appointments and reminders to wear/use and return continuous glucose monitoring device, fitness tracker, and environmental sensor. 10% of the study population will be invited to participate in a second research visit.
Examples include: To invite for a follow up study visit, to ask for interests in participating in additional studies related to this study, to obtain additional samples (ex: blood, retinal imaging), to invite to participate in the community advisory board.
Children Group1	Describe the procedures or data/specimen collection (if any) for which there will be NO parental permission2	Reason why parental permission will not be obtained	Will parents be informed about the research?3	Will parents be informed about the research?3
YES	NO
Children Group1	Describe the procedures or data/specimen collection (if any) for which assent will not be obtained	Reason why assent will not be obtained
Children Group1	Describe the procedures or data/specimen collection (if any) for which assent will NOT be documented
When designing consent process and forms, researchers should first review the GUIDANCE Consent and TIPSHEET Consent. The topics of A Foundation for Meaningful Consent and The Key Information Requirement are particularly important for ensuring subject comprehension and voluntary participation in research. Information about parental permission can be found in the GUIDANCE Consent Protected and Vulnerable Populations.
TERM	DEFINITION
CONSENT	is the process of informing potential subjects about the research and asking them whether they want to participate. It does not necessarily include the signing of a consent form.
CONSENT DOCUMENTATION	refers to how a subject’s decision to participate in the research is documented. This is typically obtained by having the subject sign a consent form.
CONSENT FORM	is a document signed by subjects, by which they agree to participate in the research as described in the consent form and in the consent process.
ELEMENTS OF CONSENT	are specific information that is required to be provided to subjects.
CHARACTERISTICS OF CONSENT	are the qualities of the consent process as a whole. These are:
Consent must be legally effective.
The process minimizes the possibility of coercion or undue influence.
Subjects or their representatives must be given sufficient opportunity to discuss and consider participation.
The information provided must:
Begin with presentation of key information (for consent materials over 2,000 words).
Be what a reasonable person would want to have.
Be organized and presented so as to facilitate understanding.
Be provided in sufficient detail.
Not ask or appear to ask subjects to waive their rights.
PARENTAL PERMISSION	is the parent’s active permission for the child to participate in the research. Parental permission is subject to the same requirements as consent, including written documentation of permission and required elements.
SHORT FORM CONSENT	is an alternative way of obtaining written documentation of consent that is most commonly used for the unanticipated enrollment of individuals who are illiterate or whose language is one for which translated consent forms are not available.
WAIVER OF CONSENT	means there is IRB approval for not obtaining consent or for not including some of the elements of consent in the consent process.
NOTE: if you plan to obtain identifiable information or identifiable biospecimens without consent, any waiver granted by the IRB does not override a subject’s refusal to provide broad consent (for example, the Northwest Biotrust).
WAIVER OF DOCUMENTATION OF CONSENT	means that there is IRB approval for not obtaining written documentation of consent.
Group1	Describe the procedures of data/specimen collection (if any) for which there will be NO consent process	Reason why consent will not be obtained	Will subjects be provided with info about the research after they finish? (Check Yes or No)	Will subjects be provided with info about the research after they finish? (Check Yes or No)
YES	NO
In-person consent:
Consent will be obtained in a private room by a study team member. Patients will be given opportunity to ask questions prior to signing the consent. If they require more time, we will provide them with the consent form and schedule additional time to review their questions and concerns prior to signing the consent.

Electronic consent:
Subjects will be contacted by email, letter, or telephone call to ascertain their interest. Participants will be directed to our study interface (REDCap) where they will have the opportunity to read the consent documentation and give e-consent, request a phone call or video conference via a CRC to ask questions, or wait to sign the consent when they come to their in-person visit. No study-specific questionnaires, with the exception of the screening questionnaire, will be available to participants until consent has been given. Once signed, electronic consent will be automatically emailed to the subjects. Subjects will be able to request an additional copy from the study coordinator for the duration of the study.

Our preferred method will be obtaining e-consent so that the participants can fill out surveys prior to the study visit in order to decrease the study visit duration. However, participants who wish to be consented in person will not be excluded and a copy will be mailed to them before their study visit.
Subjects will be asked to confirm their understanding of study activities and requirements to ensure comprehension.
We now included a table at the beginning of the consent form that clarifies which information will be publicly available, and which would require an application and approval process that would contain privacy protection. Since the beginning of the consent form describes the database, the PI believes that this is the best placement for the table and the preceding clarification.
Our first line of recruitment is through medical records database search, so it is unlikely that we will have any previous relationship with participants. However, it is possible that we will have to rely on other forms of recruitment, such as referral based. Then, if the referral to the study comes from a research staff member who a has relationship with the potential participant, the actual recruitment and consenting process will be performed by a research staff member that does NOT have any pre-existing relationship with the participant.
We have years of experience preparing consent forms for study subjects. There will also be a community advisory board that evaluates the consent forms and research protocols and their feedback will be considered in the modifications of the IRB as needed
Study subjects will have the opportunity to ask questions throughout the study and will be reminded that they have the right to change their mind about study participation
Potential participants will receive an email with a unique code to access their e-consent, supported by UW REDCap.
Web access will provide free navigation, forward and backward, within the system, as well as the ability to come back at a later time. Individuals will have the ability to exit out of the consent document and request more information from the study team.
Contact information for the study staff will be given in the recruitment materials, as well as the electronic consent form itself. Potential participants are encouraged to call, email, or set up a Zoom meeting with the study staff. The study website may also have a live chat function.
E-consent is just an additional modality to the standard paper-based consent; subjects can choose which is easier and more accessible to them.
Upon signing, the signed consent form will be automatically emailed to subjects. Subjects will be able to receive an additional copy of their signed e-consent by contacting study coordinator(s) at any time during the study.
Study participants will have access to some of their clinical laboratory data via an online portal. This is not an interventional study, so significant new findings are not expected. Subjects will receive a card with the study visit information (e.g., vital signs, visual acuity), as described.
Adult subject group1	Describe the procedures or data/specimen collection (if any) for which there will be NO documentation of consent	Will they be provided with a written statement describing the research (optional)?
(Check Yes or No)	Will they be provided with a written statement describing the research (optional)?
(Check Yes or No)
YES	NO
For subjects who have difficulty viewing the study consent, a study team member will read the consent form for them. We still expect that a subject would be capable of signing the consent form. If a person is functionally illiterate, a study team member will sign and note the reason for missing subject signature in the presence of an impartial witness.
Based on the feedback from the community advisory group, the study team realized that to truly enroll non-English speaking participants, it would be necessary to translate all materials (not just the consent form), which is currently not possible. Therefore, non-English speaking participants will not be initially enrolled and all patients will need to be proficient enough to understand instructions of the study procedures such as using fitness tracking and CGM.

¼ of our target study population is Hispanic/LatinX. We will first rely on recruiting Hispanic patients who can consent and fully communicate in English. If our recruitment goal is not being met, then the consent form will be translated into Spanish to be able to recruit patients who prefer to be consented in Spanish. We have native Spanish speaking research staff at UCSD who can translate the instructions, but we do not have research staff who speak other languages.
Native Spanish speaking research staff will provide interpretation and instructions as needed.
Spanish language translations will be obtained if we are not meeting our recruitment goals for Hispanic participants. The consent form will be translated into Spanish by UCSD translation services.
All study procedures will be conducted in private rooms.
All study data and identifiable information will be kept on encrypted servers, or in locked cabinets/room with restricted access (at all sites)
Level 3
Confidentiality breach – low risk – all identifiable information will be kept on encrypted computers/servers or in a locked cabinet/room with restricted access at all sites. Only authorized research personnel will have access to identifiable information.

Eye drops – low risk – we will use one or two types of eye drops to dilate subject’s pupils for the eye imaging. Applying eye drops for pupil dilation can cause mild stinging. Following the procedure, subjects will have blurry vision for several hours and increased sensitivity to light. We will inform them not to operate a vehicle until their vision returns to normal, and provide eye shades to protect their eyes from excessive sensitivity to light. Rarely, patients may experience prolonged dilation lasting 1-2 days. There are very rare, but serious adverse events associated with dilation drops, including angle closure attacks, allergic reactions, increased blood pressure, and arrhythmias. In the event that a participant experiences any of these rare events, emergency medical attention would be sought immediately.

Eye imaging – low risk - imaging can occasionally take a long time, or the light source used during imaging may cause slight discomfort or a headache

COVID-19 exposure – due to the current pandemic, patients may be exposed to COVID-19. We have in place an in-person research plan, per HSD guidelines, and will complete the HSD COVID-19 risk assessment tool before and during the study to ensure subjects’ safety.

Emotional stress – low risk – could be caused by any conditions that may be revealed by study participation.

Fatigue – low risk – subjects may be fatigued from the study visit, which may last several hours. We will readily accommodate any subject’s request for a break.

Blood draw – low risk – may cause bruising, fainting, pain, or infection. We will provide a small snack following blood collection to reduce a chance of fainting or weakness. Extremely rarely the puncture site may become infected, which could lead to hospitalization or death.

CGM – low risk – may cause bruising, pain, or infection at the point of insertion; irritation due to prolonged presence of the monitor/adhesive on the skin (e.g., due to excessive sweating or tight clothing), or skin discoloration. If the CGM is uncomfortable, we will let the patients remove it and mail it back before the full 10 days have elapsed.

Fitness tracking – low risk – there is no known harm to wearing a fitness tracking device. The subjects will be educated by staff on sizing the devices to minimize the discomfort caused by ill-fitted wristbands.

Environmental sensor – low risk – there is no known risk to having an environment sensor in the home.
Check all that apply	Brand Name	Generic Name	Chemical Structure	Chemical Structure
Dotarem	Gadoterate meglumine	Macrocylic	Macrocylic
Eovist / Primovist	Gadoxetate disodium	Linear	Linear
Gadavist	Gadobutro	Macrocyclic	Macrocyclic
Magnevist	Gadpentetate dimeglumine	Linear	Linear
MultiHance	Gadobenate dimeglumine	Linear	Linear
Omniscan	Gadodiamide	Linear	Linear
OptiMARK	Gadoversetamide	Linear	Linear
ProHance	Gadoteridol	Macrocyclic	Macrocyclic
Other, provide name:
N/A
N/A
N/A
Subject may receive some direct benefit by being able to access the data generated by the study procedures (see section 10.10.b).
Results obtained during the study visit may be clinically actionable, such as blood pressure or visual acuity/eye exam data and will be shared with the study participants in real time. Each participant will receive a card with these results at the end of the visit. This includes vision or life threatening actionable incidental findings, e.g., retinal detachment, disc edema, or tumors, identified during the visit. In such cases, patients will be referred for treatment.

Due to delay in obtaining the laboratory results and lack of standard interpretation methods for other data (e.g., imaging), it is unlikely that these data would be clinically actionable. All clinical laboratory results will be uploaded on a yearly basis, and participants will be able to access them (but there may be a substantial lag between the study visit and the yearly upload).

CGM data will be shared with the participants, which may include data for up to 100 days of collection.
Subjects will receive results obtained on site during the visit (e.g., blood pressure or vision exam results), CGM data, and laboratory testing results; however, laboratory testing results may not be available within the clinically actionable timeline. The laboratory testing results will be available to subjects through the study database that is updated and uploaded yearly. Each subject will receive an access code. Normal range of values will be available for the laboratory results. All abnormal laboratory results will be marked/flagged, and it will be recommended that subjects use the information and discuss the results with their primary care provider or specialist. Study will not provide interpretation of the data due to the high number of subjects.

The portal will contain a disclaimer, instructing the subjects to review the provided information with their physician(s).
Only data obtained during the study visit exam (e.g., blood pressure, eye exam information), CGM data, and laboratory test results obtained from collected blood and urine samples will be offered to the participants. Since we are recruiting from the EMRs, not every patient is seen by a specialist at the study sites or has access to care from a specialty clinic. Thus, we will return the data that can be interpreted by a primary care doctor. In addition, we are not distributing any test results that do not have standard interpretation methods (e.g., retinal images, fitness tracking).
Subjects will receive results of the tests listed under 10.10.a.2, some of which may be within the normal range. The study team will not curate the results in any way and therefore cannot exclude results that are normal and do not require treatment or intervention. Knowing that their values are normal may be reassuring to the subjects. Furthermore, these normal findings may serve as baseline in case of any future abnormal findings.
Reference ranges for laboratory testing will be provided.
As stated in 10.10.b, participants will be able to access the data via an electronic database and discuss any test results with their clinical care providers. Data access will be provisioned upon completion of all study procedures, and a secure e-mail will be sent by study staff with database access details. Study staff will communicate database access details when data are available, but will not provide any clinical interpretations.
We do not have plans to share results with family members.
There are no plans for the researchers involved in collecting the data set to develop any commercial products or file for patents. However, the data set will be publicly available with a commercial-friendly open-source license and there may be products or patents that are developed from it in the future.
Subjects may incur costs for transport that will not be reimbursed if transportation fees exceed $25.
Examples:
Research Study Coordinator: Obtain consent, administer surveys, blood draw. Will have previous experience coordinating clinical research and be a certified phlebotomist in WA.
Undergraduate Research Assistant: Obtain consent, perform all study procedures. Will have had coursework in research methods, complete an orientation to human subjects protections given by the department, and will receive training from the PI or the graduate student project lead on obtaining consent and debriefing subjects.
Acupuncturist: Perform acupuncture procedures and administer surveys. Must be licensed with WA State DoH and complete training in administering research surveys given by the project director, an experienced survey researcher.
Co-Investigator: Supervise MRI and CT scan procedures and data interpretation, obtain consent. MD, specialty in interventional radiology and body imaging. 5-years clinical research experience.
Principal Investigators/ Co-Principal/ Co-Investigator: Make executive decision pertaining to the scope of the project including data collection, curation, and processing

Program/project manager: Oversee data collection and management of research coordinators/project managers, make project related decisions in collaboration with the investigators, oversee integration and collaboration between all study members and modules

Database manager: Ingest data into the Redcap database and extract for processing

Research/Data scientist/Postdoctoral Fellow: Process data for release in database and development of standards

Clinical research coordinator/study coordinator: Responsible for contacting and consenting the research participants, ushering participants through the research visit, conducting study visits including collection of ophthalmologic images, providing training on study procedures such as continuous glucose monitoring or fitness tracker, and ensuring the return of all study materials (ex. Environmental sensor)

Research assistant: Conduct study visits including collection of ophthalmologic images, provide training on study procedures such as continuous glucose monitoring or fitness tracker, ensure the return of all study materials (ex. Environmental sensor) under the direction of the study coordinator or project manager; responsible for data upload and management from each clinic visit

Software engineer: Responsible for designing, implementing, and testing the different data
curation pipelines

Cloud engineer: Responsible for cloud infrastructure customization, configuration and deployment including integrating, configuring, deploying, and managing centrally provided common cloud services

Phlebotomist: Responsible for blood collection from research participants

Lab Assistant: Responsible for processing and shipping samples

Domain experts: Provide guidance on issues arising from collecting data from each domain of medicine such as endocrinology and cardiology.
The PI assumes the responsibility that each study team member is trained and receives timely information required for performing their research-related duties.
