EXPECT, EXPECT-denoised
收藏资源简介:
EXPECT数据集是由清华大学等机构创建的一个解释型语法错误纠正数据集,包含大约2万个经过人工标注的样本。数据集来源于W&I+LOCNESS,涵盖了不同英语水平层次的句子。EXPECT通过将含有多个错误的句子拆分为单个错误来简化任务。但由于原始数据集中存在未识别的语法错误,作者创建了EXPECT-denoised数据集以去除这些噪声,保证训练和评估的公正性。该数据集主要用于帮助语言学习者在语法纠正过程中理解纠正的原理。
The EXPECT dataset is an explainable grammatical error correction dataset created by Tsinghua University and other institutions, containing approximately 20,000 manually annotated samples. The dataset is sourced from W&I+LOCNESS, covering sentences of varying English proficiency levels. EXPECT simplifies the task by splitting sentences with multiple errors into single-error instances. However, since the original dataset contained unrecognized grammatical errors, the authors created the EXPECT-denoised dataset to remove such noise and ensure fairness in training and evaluation. This dataset is primarily designed to help language learners understand the principles behind grammatical correction during the correction process.

- 1Corrections Meet Explanations: A Unified Framework for Explainable Grammatical Error Correction清华大学 · 2025年



