Information Updation dataset
收藏资源简介:
Information Updation数据集是由印度古瓦哈蒂理工学院等机构的研究人员创建的,旨在模拟现实世界中更新过时的Wikipedia表格的过程。该数据集包含950个经过人工注释的实例,跨越9个类别和14种语言。它通过从不同时间点提取同一实体的两个版本Wikipedia表格来构建,其中一个版本作为源表格,另一个版本作为参考表格,同时还有一个由人工同步创建的金标准表格。该数据集用于评估信息更新任务,即在源表格中更新行信息,使其与金标准表格中的信息相匹配。
The Information Updation dataset was created by researchers from institutions including the Indian Institute of Technology Guwahati, aiming to simulate the real-world process of updating outdated Wikipedia tables. This dataset contains 950 manually annotated instances, spanning 9 categories and 14 languages. It is constructed by extracting two versions of Wikipedia tables for the same entity from different time points, where one version serves as the source table and the other as the reference table, alongside a gold-standard table manually created through synchronization. This dataset is used to evaluate the information update task, which involves updating the row-level information in the source table to match the content in the gold-standard table.




