structural-isomorphism-benchmark
收藏资源简介:
SIBD(结构同构基准数据集)是一个包含1,214条自然语言描述的数据集,涵盖84种不同的结构类型。每种结构类型在10多个不同的现实领域中以通俗语言(不含领域特定术语)进行描述。该数据集旨在训练和评估能够识别跨领域结构相似性的模型,例如识别恒温器和血糖调节共享相同的反馈循环结构的能力。数据集采用JSON格式,每个条目包含type_id(结构类型标识符)、type_name(人类可读的类型名称)、domain(领域)和description(现象描述)字段。数据集总条目数为1,214,平均每种结构类型约14.5个条目,语言为中文,涵盖物理学、化学、生物学、经济学、法学、教育学、医学、农业、工程学、体育等70多个领域。此外,数据集还提供了一个包含500个现实世界现象的补充知识库,分为自然科学、社会科学与人文科学以及跨学科现象三类。该数据集适用于结构相似性的嵌入模型训练、跨领域类比识别评估、结构同构和知识迁移研究以及跨领域灵感搜索引擎构建。
SIBD (Structural Isomorphism Benchmark Dataset) is a dataset containing 1,214 natural language descriptions, covering 84 distinct structural types. Each structural type is described in plain language (without domain-specific jargon) across more than 10 different real-world domains. The dataset aims to train and evaluate models capable of recognizing cross-domain structural similarity, such as the ability to identify that a thermostat and blood glucose regulation share the same feedback loop structure. The dataset is formatted in JSON, with each entry containing the fields of type_id (structural type identifier), type_name (human-readable type name), domain, and description. The total number of entries in the dataset is 1,214, with an average of approximately 14.5 entries per structural type. The language of the dataset is Chinese, and it covers more than 70 domains including physics, chemistry, biology, economics, law, education, medicine, agriculture, engineering, sports and others. In addition, the dataset provides a supplementary knowledge base consisting of 500 real-world phenomena, which are categorized into three classes: natural sciences, social sciences and humanities, and interdisciplinary phenomena. This dataset is suitable for training embedding models for structural similarity, cross-domain analogy recognition evaluation, research on structural isomorphism and knowledge transfer, as well as the construction of cross-domain inspiration search engines.




