LitBank
收藏资源简介:
LitBank数据集由加州大学伯克利分校信息学院创建,包含210,532个标记,来自100部英语小说,旨在解决文学文本中的指代问题。数据集的平均文档长度为2,000字,远超其他基准数据集,包含文学中常见的复杂指代问题。该数据集不仅用于评估指代消解系统的性能,还用于分析长距离文档内指代的特点。此外,数据集的应用领域广泛,包括文学分析、角色研究等,为研究文学文本中的指代现象提供了重要资源。
The LitBank dataset, developed by the School of Information at the University of California, Berkeley, consists of 210,532 annotated tokens derived from 100 English novels. It is constructed to tackle coreference-related challenges in literary texts. Boasting an average document length of 2,000 words—far greater than that of most existing benchmark datasets—the dataset encompasses complex coreference issues prevalent in literary works. Beyond serving as a testbed for evaluating the performance of coreference resolution systems, the dataset enables analysis of the characteristics of long-distance intra-document coreference. Furthermore, the dataset has broad applications across fields such as literary analysis and character research, offering a valuable resource for studies on coreference phenomena in literary texts.




