KG-SaF-Data
收藏资源简介:
KG-SaF-Data是由巴里大学团队构建的综合性知识图谱数据集套件,包含10个基于6种不同知识图谱的数据集。这些数据集不仅包含传统的事实三元组,还整合了丰富的模式层知识(如OWL本体),并通过模块化和推理服务确保数据一致性与完备性。数据集规模涵盖数万至数十万条三元组,来源包括DBpedia、YAGO等知名知识图谱及特定领域图谱(如文化遗产、水资源等)。其构建流程通过SPARQL查询、本体合并、模块化等技术实现,并支持PyTorch等机器学习框架的张量表示。该资源旨在解决神经符号推理(NeSy)领域缺乏模式增强型基准数据集的问题,为知识图谱补全、链接预测等任务提供标准化评估平台。
KG-SaF-Data is a comprehensive knowledge graph dataset suite constructed by the team from the University of Bari, which includes 10 datasets based on 6 distinct knowledge graphs. These datasets not only contain traditional factual triples but also integrate rich schema-level knowledge such as OWL ontologies, and ensure data consistency and completeness through modularization and inference services. The scale of these datasets ranges from tens of thousands to hundreds of thousands of triples, with sources covering well-known knowledge graphs like DBpedia and YAGO, as well as domain-specific graphs such as cultural heritage and water resources. Its construction pipeline is implemented via technologies including SPARQL queries, ontology merging and modularization, and supports tensor representation in machine learning frameworks like PyTorch. This resource aims to address the shortage of schema-augmented benchmark datasets in the field of neural-symbolic reasoning (NeSy), providing a standardized evaluation platform for tasks such as knowledge graph completion and link prediction.



