Grammarly Corpus of Discourse Coherence (GCDC)
收藏资源简介:
Grammarly Corpus of Discourse Coherence (GCDC) 是一个用于评估真实世界文本中话语连贯性的数据集,由伊利诺伊大学厄巴纳-香槟分校创建。该数据集包含来自四个不同领域的4800条文本,包括论坛帖子、电子邮件和产品评论等,每条文本都由专家标注者进行了连贯性评分。GCDC数据集的创建旨在通过大规模评估领先的话语连贯性算法,解决现有研究中缺乏真实世界数据评估的问题。数据集的应用领域广泛,旨在提高自然语言生成系统的质量,并为写作提供反馈,如识别话题间的缺失过渡或突出组织不良的段落。
The Grammarly Corpus of Discourse Coherence (GCDC) is a benchmark dataset designed for evaluating discourse coherence in real-world texts, developed by the University of Illinois Urbana-Champaign. It includes 4,800 texts across four distinct domains, such as forum posts, emails, product reviews, and other similar text genres. Each text has been assigned a coherence score by expert annotators. The primary purpose of constructing the GCDC dataset is to address the shortage of real-world evaluation data in existing research, via large-scale benchmarking of state-of-the-art discourse coherence algorithms. This dataset has wide-ranging applications, aiming to improve the quality of natural language generation systems and provide writing feedback, for example, identifying missing transitions between topics or highlighting poorly organized paragraphs.




