遇见数据集

FAIR Science for Social Machines: Let’s Share Metadata Knowlets in the Internet of FAIR Data and Services

收藏
Mendeley Data2022-11-10 更新2024-06-29 收录
官方服务:

资源简介:

11 figures of this paper. Figure 1 is the hourglass model of the Internet architecture. Figure 2 shows the merge of three hourglasses (data-infrastructure, tools-infrastructure and compute-infrastructure) into the image of a propeller with three blades and the underlying infrastructure. The narrow waist of the hourglass (minimal essential standards and protocols) is comparable to the center of this picture. Figure 3 is the simple Digital Object picture. The smallest conceivable Digital Object is a persistent identifier (PID) (a digital symbol referring to a particular concept). Each digital object that contains “information” should be adorned with metadata asserting things about the nature of that information. Typical intrinsic metadata describe the factual information that is “indisputable” about the digital object itself. Intrinsic metadata containers, expanded metadata containers and the actual containers holding the data elements or the core (in case of for instance a workflow) could also be treated as separate but permanently-linked digital objects, each with their own unique, persistent and resolvable identifier (UPRI) and thus form a stack of related metadata containers that contain (machine readable, FAIR) metadata of different nature, all asserting, however, relevant information about the data container. Figure 4 shows how in the developing Internet of FAIR Data and Services, a linked-data-compliant query in a virtual machine format could automatically find the most relevant databases. Figure 5 shows the semiotic triangle, based on the concept of cancer. Figure 6 shows the single meaningful assertion in machine readable format, which is called a nanopublication. The smallest conceivable assertion has the structure of a subject, a predicate and an object. To form a nanopublication this “triple” needs to be published in machine readable format with full provenance and publication information (also in machine readable format). Figure 7 shows the Knowlet as a collection of cardinal assertions “about” a given subject. The objects effectively form the “conceptual context” of explicitly associated concepts. The predicates can range from very specific and explicit relationship descriptions such as “inhibits” or “is married to” to more generic and less explicit connections, such as “co-occurs in the same sentence as”. Figure 8 shows that the Knowlet is a digital object and needs to be findable, accessible, interoperable and reusable (i.e., FAIR) in its own right. It also may change over time, when more assertions are collected about the core concept. Therefore, each Knowlet in the Internet of FAIR Data and Services (IFDS) needs a unique, persistent and resolvable identifier (UPRI). Figure 9 shows that the Knowlet can be seen as a metadata container for the concept it represents. It can represent many different things from plain concepts like a gene or a person (ORCID record), to a data set, a data base, a work flow or any other thing in the Internet of Things. Figure 10 shows three ways in which Knowlets can be used to connect dispersed digital objects. In Figure 11, A: Concepts, physical objects or things of different semantic types (and thus also intrinsically meaningless unique, persistent and resolvable identifiers (UPRIs)) can cluster based on contextual similarity without ever being explicitly connected (drug might treat disease). B: Nearly identical concepts that are nevertheless in certain circumstances to be seen as distinct, will automatically cluster as one if the resolution of search or matching is lowered, while they will separate out when the resolution is made higher. C: Conceptual and semantic drift occur.

本文包含11幅插图。图1为互联网架构的沙漏模型。图2展示了将三类沙漏——数据基础设施(data-infrastructure)、工具基础设施(tools-infrastructure)与计算基础设施(compute-infrastructure)——整合为一幅三叶片螺旋桨搭配底层基础设施的示意图,沙漏的窄腰(极简核心标准与协议)与该图的中心区域相对应。图3为简易数字对象(Digital Object)示意图。可设想的最小数字对象为持久标识符(persistent identifier, PID),即指向特定概念的数字符号。每一项承载“信息”的数字对象,都应附加元数据,用以阐明该信息的本质属性。典型的固有元数据(intrinsic metadata)用于描述关于该数字对象本身“无可辩驳”的事实性信息。固有元数据容器、扩展元数据容器,以及实际承载数据元素或核心内容(例如工作流场景)的容器,均可视为相互独立但永久关联的数字对象,各自拥有唯一、持久且可解析的标识符(unique, persistent and resolvable identifier, UPRI),由此形成一系列关联的元数据容器栈,其中包含不同类型的(机器可读、可发现、可访问、可互操作、可重用(Findable, Accessible, Interoperable, Reusable, FAIR))元数据,且均用于阐明对应数据容器的相关信息。图4展示了在FAIR数据与服务物联网(Internet of FAIR Data and Services, IFDS)的发展进程中,符合关联数据规范的虚拟机格式查询,如何自动定位到最相关的数据库。图5为基于癌症概念构建的符号三角(semiotic triangle)。图6展示了机器可读格式下的单条有效断言,即纳米出版物(nanopublication)。可设想的最小断言具备主体、谓词与客体的三元结构。要形成纳米出版物,该“三元组”需以机器可读格式发布,并附带完整的溯源信息与出版元数据(同样采用机器可读格式)。图7将Knowlet展示为针对特定主题的核心断言集合。其客体可有效构成显式关联概念的“概念上下文”。谓词的范围可从“抑制”“与……结为配偶”这类具体明确的关系描述,延伸至“与……同句共现”这类泛化且语义模糊的关联。图8表明,Knowlet本身属于数字对象,需具备独立的可发现、可访问、可互操作与可重用(FAIR)特性。随着针对核心概念的断言不断新增,Knowlet也可随时间演进。因此,FAIR数据与服务物联网(Internet of FAIR Data and Services, IFDS)中的每一个Knowlet,都需要一个唯一、持久且可解析的标识符(UPRI)。图9展示了Knowlet可作为其所代表概念的元数据容器,可表征多种不同对象:从基因、个人(Open Researcher and Contributor ID, ORCID)记录这类基础概念,到数据集、数据库、工作流,乃至物联网中的任意其他实体。图10展示了利用Knowlet连接分散式数字对象的三种方式。在图11中:A:不同语义类型的概念、物理实体或事物(因此其本身也附带无固有语义含义的唯一、持久且可解析标识符(UPRI)),可基于上下文相似度进行聚类,而无需显式关联(例如药物可治疗疾病)。B:在特定场景下应视为不同的近乎一致的概念,当搜索或匹配的分辨率降低时会自动聚为一类,而当分辨率提高时则会彼此分离。C:会出现概念与语义漂移现象。

创建时间:
2022-02-10
二维码
社区交流群
二维码
科研交流群
商业服务