EMERGE
收藏资源简介:
EMERGE是一个自动构建的基准数据集,用于将文本来源中的新知识与知识图谱(KG)中的变化进行对齐。具体来说,我们将维基数据KG(Vrandeˇci´c和Krötzsch,2014)中的演变更新与反映相同时期新兴知识的维基百科文本段落进行关联。该数据集包括376K个维基百科段落,与从2019年到2025年维基数据的10个不同快照中的总计1.25M个KG编辑相匹配。我们的实验结果突出了基于新兴文本知识更新KG快照的挑战,并将该数据集定位为未来研究的宝贵基准。我们还将公开发布我们的数据集和模型实现。
EMERGE is an automatically constructed benchmark dataset designed to align new knowledge from textual sources with changes in Knowledge Graphs (KGs). Specifically, we associate evolutionary updates from the Wikidata KG (Vrandečić and Krötzsch, 2014) with Wikipedia textual passages that reflect emergent knowledge from the same time period. This dataset comprises 376K Wikipedia passages, matched with a total of 1.25M KG edits across 10 distinct snapshots of Wikidata spanning from 2019 to 2025. Our experimental results highlight the challenges of updating KG snapshots based on emergent textual knowledge, and position this dataset as a valuable benchmark for future research. We will also publicly release our dataset and model implementations.




