KNOWREF
收藏资源简介:
KNOWREF数据集由Mila/麦吉尔大学计算机科学学院和微软研究蒙特利尔共同创建,包含8724个需要大量常识和背景知识来解决的Winograd风格的文本样本。数据集通过从2018年英文维基百科、OpenSubtitles和Reddit评论中筛选和标注得到,旨在解决现有指代消解方法依赖性别和数量线索的问题。KNOWREF数据集的应用领域主要集中在提升模型对文本情境的推理能力,特别是在缺乏明显性别或数量线索的情况下。
The KNOWREF dataset was co-created by the School of Computer Science, Mila / McGill University and Microsoft Research Montreal. It contains 8,724 Winograd-style text samples that require substantial common sense and background knowledge to solve. The dataset is obtained by filtering and annotating from the 2018 English Wikipedia, OpenSubtitles and Reddit comments, aiming to address the problem that existing coreference resolution methods rely heavily on gender and number cues. The main application areas of the KNOWREF dataset focus on improving models' reasoning abilities for textual contexts, especially when there are no obvious gender or number cues.

- 1The Knowref Coreference Corpus: Removing Gender and Number Cues for Difficult Pronominal Anaphora ResolutionMila/麦吉尔大学计算机科学学院 2微软研究蒙特利尔 · 2019年



