C2EC
收藏资源简介:
C2EC数据集是由苏州大学计算机科学与技术学院的研究人员构建的,旨在解决通用汉字错误校正任务。该数据集通过整合CCTC和Lemon两个数据集中的真实世界错误,并经过人工审核,确保数据质量和标注一致性。它包含1995句开发数据和5711句测试数据,大约一半的句子是无错误的。测试集中72.6%的错误是拼写错误,14.0%是多余字符,13.4%是缺失字符。
The C2EC dataset was developed by researchers from the School of Computer Science and Technology, Soochow University, aiming to address the task of general Chinese character error correction. It integrates real-world errors from two existing datasets, CCTC and Lemon, and has undergone manual review to ensure data quality and annotation consistency. The dataset contains 1995 development sentences and 5711 test sentences, with approximately half of the sentences being error-free. Among the errors in the test set, 72.6% are spelling errors, 14.0% are extra character errors, and 13.4% are missing character errors.

- 1A Training-free LLM-based Approach to General Chinese Character Error Correction苏州大学计算机科学与技术学院 · 2025年



