proxectonos/galician-gec-corpora
收藏资源简介:
Galician GEC Corpora 是一个加利西亚语语法和拼写校正数据集的集合,该仓库将多个句子级校正对来源分组到一个Hugging Face数据集仓库中,每个来源作为一个单独的配置公开。每个实例包含一个错误的或非标准的加利西亚语句子及其校正版本,部分子集还包括错误标签、校正标签、编辑元数据或解释。该数据集旨在用于加利西亚语的语法错误校正、拼写校正和文本校正系统的训练、评估和分析。数据集包含五个配置:che_pairs(专注于加利西亚语附着代词放置的小型目标子集)、cortegal(基于CORTEGAL风格校正示例的标准化版本)、gec_synthetic(包含错误类别、标签和解释的合成语法错误校正数据集)、parlamint_punctuation(源自加利西亚议会文本的标点校正对)和wikipedia_breobot(源自加利西亚维基百科/Breobot编辑的校正对)。数据以JSONL格式分发,总共有约120,817个校正对。
Galician GEC Corpora is a collection of Galician grammatical and orthographic correction datasets. The repository groups several sources of sentence-level correction pairs in a single Hugging Face dataset repository, with each source exposed as a separate configuration. Each instance contains an incorrect or non-standard Galician sentence and its corrected version. Some subsets also include error labels, correction tags, edit metadata, or explanations. The dataset is intended for training, evaluating, and analysing grammatical error correction, orthographic correction, and text correction systems for Galician. The dataset contains five configurations: che_pairs (a small targeted subset focused on Galician clitic placement), cortegal (a normalized version of CORTEGAL-style correction examples), gec_synthetic (a synthetic Galician grammatical error correction dataset with error categories, tags, and explanations), parlamint_punctuation (punctuation-focused correction pairs derived from Galician parliamentary text), and wikipedia_breobot (correction pairs derived from Galician Wikipedia/Breobot edits). The data is distributed in JSONL format, with a total of approximately 120,817 correction pairs.




