syauqie/IGED
收藏资源简介:
IGED(Indonesian Grammar Error correction Dataset)是一个大规模印尼语语法错误纠正(GEC)数据集,包含1,345,096个句子对,是公开可用的最大印尼语GEC数据集。每个句子对由一个不合语法的源句子和一个语法正确的目标句子组成。数据集覆盖印尼语的三大错误类别:形态学错误(包括词缀、构词和重叠)、句法错误(包括短语结构、介词和句子完整性)和语义错误(包括措辞、冗言和歧义)。IGED是CASTLE系统的一部分,CASTLE是一个用于低资源语法纠正的上下文感知语义Transformer,带有知识图谱增强,相关研究发表于《Expert Systems With Applications》(2026年)。
IGED is a large-scale Indonesian Grammar Error Correction (GEC) dataset containing 1,345,096 sentence pairs — the largest publicly available GEC dataset for Indonesian. Each pair consists of an ungrammatical source sentence and its grammatically correct target. The dataset covers three major error categories in Indonesian: morphological (affixation, word formation, reduplication), syntactic (phrase structure, preposition, sentence completeness), and semantic errors (diction, pleonasm, ambiguity). IGED was introduced as part of the CASTLE system — a Context-Aware Semantic Transformer with Knowledge Graph Enhancement for low-resource grammar correction, published in Expert Systems With Applications (2026).




