D4D Property	Description
acquisition_methods	(slot exists but no description)
addressing_gaps	Was there a specific gap that needed to be filled by creation of the dataset?
anomalies	(slot exists but no description)
annotation_analyses	(slot exists but no description)
bytes	Size of the data in bytes.
cleaning_strategies	Was any cleaning of the data done (e.g., removal of instances, processing of missing values)?
collection_mechanisms	What mechanisms or procedures were used to collect the data (e.g., hardware, manual curation, software APIs)? Also covers how these mechanisms were validated.
collection_timeframes	Over what timeframe was the data collected, and does this timeframe match the creation timeframe of the underlying data?
compression	compression format used, if any. e.g., gzip, bzip2, zip
confidential_elements	(slot exists but no description)
conforms_to	(slot exists but no description)
content_warnings	Does the dataset contain any data that might be offensive, insulting, threatening, or otherwise anxiety-provoking if viewed directly?
created_by	(slot exists but no description)
created_on	(slot exists but no description)
creators	Who created the dataset (e.g., which team, research group) and on behalf of which entity (e.g., company, institution, organization)? This may also be considered a team.
data_collectors	Who was involved in the data collection (e.g., students, crowdworkers, contractors), and how they were compensated.
data_protection_impacts	Has an analysis of the potential impact of the dataset and its use on data subjects (e.g., a data protection impact analysis) been conducted? If so, please provide a description of this analysis, including the outcomes, and any supporting documentation.
description	(slot exists but no description)
dialect	(slot exists but no description)
discouraged_uses	Are there tasks for which the dataset should not be used?
distribution_dates	When will the dataset be distributed?
distribution_formats	How will the dataset be distributed (e.g., tarball on a website, API, GitHub)?
doi	digital object identifier
download_url	URL from which the data can be downloaded. This is not the same as the landing page, which is a page that describes the dataset. Rather, this URL points directly to the data itself.
encoding	the character encoding of the data
errata	(slot exists but no description)
ethical_reviews	Were any ethical or compliance review processes conducted (e.g., by an institutional review board)? If so, please provide a description of these review processes, including the frequency of review and documentation of outcomes, as well as a link or other access point to any supporting documentation.
existing_uses	Has the dataset been used for any tasks already?
extension_mechanism	If others want to extend/augment/build on/contribute to the dataset, is there a mechanism for them to do so? If so, please describe how those contributions are validated and communicated.
external_resource	Is the dataset self-contained or does it rely on external resources (e.g., websites, other datasets)? If external, are there guarantees that those resources will remain available and unchanged?
funders	(slot exists but no description)
future_use_impacts	Is there anything about the dataset's composition or collection that might impact future uses or create risks/harm (e.g., unfair treatment, legal or financial risks)? If so, describe these impacts and any mitigation strategies.
hash	hash of the data
human_subject_research	Information about whether the dataset involves human subjects research and what regulatory or ethical review processes were followed.
imputation_protocols	Description of data imputation methodology, including techniques used to handle missing values and rationale for chosen approaches.
informed_consent	Details about informed consent procedures used in human subjects research.
instances	What do the instances that comprise the dataset represent (e.g., documents, photos, people, countries)?
intended_uses	Explicit statement of intended uses for this dataset. Complements FutureUseImpact by focusing on positive, recommended applications rather than risks. Aligns with RO-Crate "Intended Use" field.
ip_restrictions	(slot exists but no description)
is_deidentified	(slot exists but no description)
is_tabular	(slot exists but no description)
issued	(slot exists but no description)
keywords	(slot exists but no description)
known_biases	(slot exists but no description)
known_limitations	(slot exists but no description)
labeling_strategies	Was any labeling of the data done (e.g., part-of-speech tagging)? This class documents the annotation process and quality metrics.
language	language in which the information is expressed
last_updated_on	(slot exists but no description)
license	(slot exists but no description)
license_and_use_terms	Will the dataset be distributed under a copyright or other IP license, and/or under applicable terms of use? Provide a link or copy of relevant licensing terms and any fees.
machine_annotation_analyses	(not found in schema)
maintainers	Who will be supporting/hosting/maintaining the dataset?
md5	md5 hash of the data
media_type	The media type of the data. This should be a MIME type.
missing_data_documentation	Documentation of missing data in the dataset, including patterns, causes, and strategies for handling missing values.
modified_by	(slot exists but no description)
other_tasks	What other tasks could the dataset be used for?
page	(slot exists but no description)
path	(slot exists but no description)
preprocessing_strategies	Was any preprocessing of the data done (e.g., discretization or bucketing, tokenization, SIFT feature extraction)?
prohibited_uses	Explicit statement of prohibited or forbidden uses for this dataset. Stronger than DiscouragedUse - these are uses that are explicitly not permitted by license, ethics, or policy. Aligns with RO-Crate "Prohibited Uses" field.
publisher	(slot exists but no description)
purposes	For what purpose was the dataset created?
raw_data_sources	Description of raw data sources before preprocessing, cleaning, or labeling. Documents where the original data comes from and how it can be accessed.
raw_sources	(slot exists but no description)
regulatory_restrictions	(slot exists but no description)
resources	Sub-resources or component datasets. Used in DatasetCollection to contain Dataset objects, and in Dataset to allow nested resource structures.
retention_limit	(slot exists but no description)
sampling_strategies	Does the dataset contain all possible instances, or is it a sample (not necessarily random) of instances from a larger set? If so, how representative is it?
sensitive_elements	Does the dataset contain data that might be considered sensitive (e.g., race, sexual orientation, religion, biometrics)?
sha256	sha256 hash of the data
status	(slot exists but no description)
subpopulations	Does the dataset identify any subpopulations (e.g., by age, gender)? If so, how are they identified and what are their distributions?
subsets	(slot exists but no description)
tasks	Was there a specific task in mind for the dataset's application?
title	the official title of the element
updates	(slot exists but no description)
use_repository	Is there a repository that links to any or all papers or systems that use the dataset? If so, provide a link or other access point.
version	(slot exists but no description)
version_access	Will older versions of the dataset continue to be supported/hosted/maintained? If so, how? If not, how will obsolescence be communicated to dataset consumers?
vulnerable_populations	Information about protections for vulnerable populations in human subjects research.
was_derived_from	(slot exists but no description)
