Project Dialogism Novel Corpus (PDNC)
收藏资源简介:
Project Dialogism Novel Corpus (PDNC)是由多伦多大学创建的一个针对英语文学文本引用归属问题的数据集。PDNC包含22部完整小说中的35,978条引用标注,是目前同类数据集中最大的。数据集中的每条引用都标注了说话者、对话对象、引用类型、指称表达和引用文本中的角色提及。PDNC的创建旨在通过提供大规模的标注数据,帮助评估和改进文学文本中的引用归属和共指模型。数据集的应用领域包括文学文本的计算分析,如角色提及追踪、角色风格变化分析等。
Project Dialogism Novel Corpus (PDNC), developed by the University of Toronto, is a dataset focused on the task of quotation attribution in English literary texts. PDNC contains 35,978 annotated quotation instances spanning 22 full-length novels, making it the largest publicly available dataset of its kind. Each quotation in the dataset is annotated with speaker identity, addressee, quotation type, referential expressions, and character mentions within the quoted text. The development of PDNC aims to provide large-scale annotated data to help evaluate and improve models for quotation attribution and coreference resolution in literary texts. Application domains of this dataset include computational analysis of literary texts, such as character mention tracking and analysis of character stylistic variation.




