OmniGEC
收藏资源简介:
OmniGEC是一个用于语法错误纠正(GEC)的多语言银标准数据集,覆盖了11种语言:捷克语、英语、爱沙尼亚语、德语、希腊语、冰岛语、意大利语、拉脱维亚语、斯洛文尼亚语、瑞典语和乌克兰语。这些数据集有助于开发多语言GEC解决方案,并有助于弥合将英语GEC解决方案适应多语言GEC的数据差距。数据集的文本来自三个来源:11种目标语言的维基百科编辑、11种目标语言的Reddit子版块和仅乌克兰的UberText 2.0社交媒体语料库。维基百科编辑是从人工纠正中派生出来的,而Reddit和UberText 2.0数据是使用GPT-4o-mini模型自动纠正的。数据集中的校正质量既通过自动方式也通过手动方式进行评估。最后,我们对两个开源大型语言模型——Aya-Expanse (8B)和Gemma-3(12B)——进行了微调,并在多语言OmniGEC语料库上取得了段落级多语言GEC的最新(SOTA)成果。数据集收集和表现最佳的模型可在Hugging Face上获得。
OmniGEC is a multilingual silver-standard dataset for Grammatical Error Correction (GEC) covering 11 languages: Czech, English, Estonian, German, Greek, Icelandic, Italian, Latvian, Slovenian, Swedish, and Ukrainian. It facilitates the development of multilingual GEC solutions and helps bridge the data gap when adapting English-centric GEC systems to cross-lingual scenarios. The text data of the dataset is sourced from three origins: 1) Wikipedia edits of the 11 target languages, 2) Reddit subreddits of the 11 target languages, and 3) the UberText 2.0 social media corpus exclusive to Ukrainian. Wikipedia edits are derived from human-curated corrections, while data from Reddit and UberText 2.0 were automatically corrected using the GPT-4o-mini model. The correction quality of the dataset is evaluated via both automatic and manual methodologies. Finally, we fine-tuned two open-source large language models — Aya-Expanse (8B) and Gemma-3 (12B) — and achieved state-of-the-art (SOTA) results for paragraph-level multilingual GEC on the OmniGEC corpus. The dataset and the best-performing fine-tuned model are publicly available on Hugging Face.




