Metadata Completeness for Four FAIR Use Cases for all DataCite Repositories
收藏资源简介:
FAIR Principle R1 recommends that "metadata are richly described with a plurality of accurate and relevant attributes", but the specifics of the metadata elements are left up to the community. Several years ago we proposed sets of DataCite metadata attributes that could be used to support several FAIR Use Cases and described how measuring completeness of those attributes could provide quantitative metrics that could be used to measure metadata FAIRness and track improvements through time. These ideas have been described and demonstrated in many contexts (see related identifiers). The current implementation includes four FAIR Use Cases: FAIR Text, FAIR Identifiers, FAIR Connections, and FAIR Contacts that contain 63 DataCite metadata elements. During May 2025 we calculated completeness for these use cases for 3066 DataCite repositories. This dataset is the result of that effort. It includes repository ids, number of records, completeness (% of records) for each use case and the total completeness. This version adds provider and consortium identifiers to the dataset to support repository grouping and completeness totals from May to compare to the totals during January 2026.
FAIR(可发现、可访问、可互操作、可复用)原则R1建议“元数据应通过丰富多样的准确且相关的属性进行详尽描述”,但元数据元素的具体细则由相关社区自主确定。数年前,我们提出了多组DataCite元数据属性集,可用于支撑多项FAIR应用场景,并阐述了如何通过度量这些属性的完备性,得到可用于评估元数据FAIR水平、并随时间追踪改进进程的量化指标。相关理念已在多种场景中得到阐述与验证(详见相关标识符)。当前实现涵盖四项FAIR应用场景:FAIR文本、FAIR标识符、FAIR关联及FAIR联系人,共包含63项DataCite元数据属性。2025年5月,我们针对3066个DataCite知识库,计算了上述应用场景的属性完备性。本数据集即为该项工作的产出成果,其中包含知识库ID、记录总数、各应用场景的完备性(以符合要求的记录占比计)以及总完备性。本版本新增了知识库提供商与联盟标识符,便于对知识库进行分组,并可将2025年5月的完备性统计结果与2026年1月的统计结果进行对比。



