FlaCGEC
收藏资源简介:
FlaCGEC是一个由华东师范大学创建的中文语法错误修正数据集,包含10,000个句子,涵盖78个具体的语法点和3种编辑类型。数据集通过从汉语专家定义的语言模式中收集原始语料,通过规则进行句子编辑,并手动精炼生成样本。FlaCGEC旨在为解释和诊断中文语法错误修正方法提供深入的语言拓扑结构,适用于自然语言处理中的写作辅助和搜索引擎等场景,以解决复杂的语法错误问题。
FlaCGEC is a Chinese grammatical error correction dataset developed by East China Normal University. It contains 10,000 sentences, covering 78 specific grammatical points and 3 types of editing operations. The dataset is constructed by collecting raw corpora from language patterns defined by Chinese language experts, conducting sentence edits via rule-based methods, and then manually refining the generated samples. FlaCGEC aims to provide in-depth linguistic topological structures for interpreting and diagnosing Chinese grammatical error correction methods, and is applicable to scenarios such as writing assistance and search engines in the field of natural language processing to address complex grammatical error issues.

- 1FlaCGEC: A Chinese Grammatical Error Correction Dataset with Fine-grained Linguistic Annotation华东师范大学 · 2023年



