GECTurk
收藏资源简介:
GECTurk是首个针对土耳其语的语法错误修正和检测数据集,由科奇大学的计算机工程系创建。该数据集包含超过138,000条高质量平行句,源自专业编辑的文章,并覆盖了20多种专家策划的语法和拼写规则。数据集的创建过程涉及复杂的转换函数,以模拟土耳其语的复杂书写规则。GECTurk不仅用于开发和评估土耳其语的语法错误修正工具,还旨在解决土耳其语在自然语言处理领域的资源稀缺问题。
GECTurk is the first grammatical error correction and detection dataset for Turkish, created by the Department of Computer Engineering at Koç University. This dataset contains over 138,000 high-quality parallel sentence pairs sourced from professionally edited articles, and covers more than 20 grammar and spelling rules curated by domain experts. The construction process of GECTurk involves complex transformation functions to simulate the intricate orthographic rules of the Turkish language. GECTurk is not only utilized for developing and evaluating Turkish grammatical error correction tools, but also aims to address the resource scarcity problem faced by Turkish in the field of natural language processing.

- 1GECTurk: Grammatical Error Correction and Detection Dataset for Turkish计算机工程系,科奇大学,伊斯坦布尔,土耳其 · 2023年



