Ro551/WikiCorrupted_spanish_to_GEC-GED_min
收藏资源简介:
该数据集是一个用于语法和拼写错误检测的语料库,包含原始句子、错误句子、标记、错误标签、错误类型、错误跨度、注释等特征。错误标签覆盖了多种类型,如正确、语法错误(如单复数、动词形式、冠词使用等)、拼写错误(如缺失、重音错误、拼写错误等)。数据集分为训练集(1178个示例)、验证集(397个示例)和测试集(408个示例),总大小约为4.3MB,适用于自然语言处理任务,如错误纠正和文本分析。
This dataset is a corpus for grammar and spelling error detection, containing features such as original sentences, corrupted sentences, tokens, error tags, error types, error spans, annotations, and more. The error tags cover various types including correct, grammar errors (e.g., singular/plural, verb form, article usage), and spelling errors (e.g., missing, accent errors, mistakes). It is split into training set (1178 examples), validation set (397 examples), and test set (408 examples), with a total size of approximately 4.3MB, suitable for natural language processing tasks like error correction and text analysis.




