Wikidata5M
收藏arXiv2025-09-30 收录
官方服务:
资源简介:
该数据集名为Wikidata5M,是一个更接近实际规模的数据库,包含了大约500万个实体和822种关系。Wikidata5M设有两种模式:转导模式和归纳模式,它们在训练集和测试集的配置上有所不同。其规模达到了500万个实体和822种关系,所面临的任务是知识图谱的完善。
This dataset, named Wikidata5M, is a database that closely aligns with real-world scale, containing approximately 5 million entities and 822 distinct relations. Wikidata5M provides two experimental settings: transductive and inductive, which differ in the configuration of training and test sets. With its scale of 5 million entities and 822 relations, the core task of this dataset is knowledge graph completion.
搜集汇总
数据集介绍

背景与挑战
背景概述
Wikidata5M是一个百万规模的知识图谱数据集,整合了Wikidata知识图谱和维基百科页面,每个实体都有对应的维基百科描述,支持转导和归纳两种数据分割方式,用于评估未见实体的链接预测。数据集包含知识图谱、语料库和别名三部分,遵循Wikidata标识符系统(实体以'Q'为前缀,关系以'P'为前缀),适用于知识表示和自然语言处理研究。
以上内容由遇见数据集搜集并总结生成



