NaSGEC
收藏资源简介:
NaSGEC是一个针对中文母语者文本的多领域语法错误修正数据集,由苏州大学人工智能研究院创建。该数据集包含来自社交媒体、科学写作和考试三个领域的12500个句子,旨在解决中文语法错误修正的跨领域问题。数据集的创建过程包括数据收集、独立标注和专家评审,确保了数据质量。NaSGEC的应用领域广泛,包括写作辅助、论文校对和中文教学,为中文语法错误修正提供了丰富的测试平台。
NaSGEC is a multi-domain grammatical error correction dataset for texts written by native Chinese speakers, developed by the Institute of Artificial Intelligence at Soochow University. This dataset comprises 12,500 sentences sourced from three domains: social media, scientific writing, and examinations, aiming to address cross-domain challenges in Chinese grammatical error correction. The dataset creation workflow includes data collection, independent annotation, and expert review, which guarantees the high quality of the data. NaSGEC covers a wide range of application scenarios, including writing assistance, thesis proofreading, and Chinese language teaching, serving as a rich testbed for Chinese grammatical error correction research.




