CEDCC
收藏资源简介:
本论文介绍了中国作文语篇连贯性语料库(CEDCC),这是一个用于评估语篇连贯性的多任务数据集。现有研究往往关注语篇连贯性的孤立维度,CEDCC通过整合连贯性评分、主题连续性和语篇关系来填补这一空白。这种方法,连同详细的标注,捕捉了现实世界文本的细微差别,并促进了中文语篇连贯性分析的进展。我们的贡献包括CEDCC的开发、为后续研究建立的基准,以及连贯性对语篇关系识别和自动作文评分影响的展示。
This paper presents the Chinese Essay Discourse Coherence Corpus (CEDCC), a multi-task dataset designed for evaluating discourse coherence. Existing studies typically focus on isolated dimensions of discourse coherence, and CEDCC fills this critical gap by integrating coherence scores, topic continuity, and discourse relations. This approach, paired with fine-grained annotations, captures the subtle nuances of real-world texts and advances research in Chinese discourse coherence analysis. Our contributions encompass the development of the CEDCC corpus, the establishment of standardized benchmarks for subsequent research, and the demonstration of the impact of coherence on discourse relation recognition and automatic essay scoring.
数据集概述
- 数据集名称: CEDCC_corpus
- 数据集目的: 用于评估中文作文的语篇连贯性,包括结构、主题和逻辑分析。
- 数据集特点:
- 整合了语篇连贯性的多个维度,如语篇评分、主题连续性和语篇关系。
- 包含详细的标注,以捕捉真实文本的细微差别。
- 数据集贡献:
- 开发了CEDCC数据集。
- 建立了研究基线。
- 展示了连贯性对语篇关系识别和自动作文评分的影响。
- 相关论文:
- 论文标题: "A Multi-Task Dataset for Assessing Discourse Coherence in Chinese Essays: Structure, Theme, and Logic Analysis"
- 发表于: EMNLP 2023
- 论文链接: EMNLP 2023论文
- 数据集可用性:
- 数据集及相关代码可在GitHub获取。




