KoCoNovel
收藏资源简介:
KoCoNovel由首尔国立大学的研究团队创建,旨在为韩国文学文本中的字符共指解析提供丰富的数据支持。该数据集包含了50部小说中的178K个Token,是继NIKL语料库之后的第二大公共共指解析语料库,并且是第一个基于文学作品的共指数据集。KoCoNovel的独特之处在于,其24%的角色提及为单个普通名词,没有修饰语,这一特征深受韩国称谓文化的影响,该文化倾向于使用表示社会关系和亲属关系的术语,而非个人姓名。数据集提供了四种不同版本的数据集,从全知视角和读者视角进行注释,以及将多个实体作为独立或重叠实体处理。KoCoNovel的发布,不仅填补了韩国文学文本共指数据集的空白,也为自然语言处理领域的研究者提供了宝贵的资源。
KoCoNovel was developed by a research team at Seoul National University, aiming to provide robust data support for character coreference resolution in Korean literary texts. This dataset contains 178K Tokens across 50 novels, making it the second-largest public coreference resolution corpus after the NIKL Corpus, and the first coreference dataset based on literary works. A distinctive feature of KoCoNovel is that 24% of character mentions are single common nouns without modifiers, a trait deeply influenced by Korean address and kinship cultural conventions, which prefer terms denoting social relationships and kinship over personal names. The dataset offers four distinct versions, annotated from either the omniscient or reader perspective, with multiple entities treated as either independent or overlapping entities. The release of KoCoNovel not only fills the gap in coreference datasets for Korean literary texts, but also provides a valuable resource for researchers in the field of natural language processing.
数据集概述
数据集名称
KoCoNovel
数据集描述
KoCoNovel 是一个基于50部现代和当代韩国小说的角色共指数据集。该数据集包含了经过语法修正的小说版本,并针对角色共指进行了标注,提供了四种选项类型,以及所有直接引语的说话者标注。
数据来源
数据集的文本来源于Wikisource的公共领域文本。预处理包括纠正文本中的拼写错误和错误的换行,以及调整拼写以符合现代韩语语法。
数据和标注
- 标注类型:
- [Reader/Omniscient]:从全知作者或读者的角度
- [Separate/Overlapped]:多个实体被处理为独立实体(例如,[‘我们’], [‘我’], [‘你’])或重叠实体(例如,[‘我们’, ‘我’], [‘我们’, ‘你’])
引用信息
若使用此数据集,请引用以下工作:
@misc{kim2024koconovel, title={KoCoNovel: Annotated Dataset of Character Coreference in Korean Novels}, author={Kyuhee Kim and Surin Lee and Sangah Lee}, year={2024}, eprint={2404.01140}, archivePrefix={arXiv}, primaryClass={cs.CL} }




