遇见数据集

Ro551/WikiCorrupted_spanish_to_GEC-GED_large

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个用于语法错误检测和文本纠错任务的数据集,包含原始句子、带错误的句子、分词结果、错误类型标签(涵盖语法错误如单复数、动词形式、冠词使用等,以及拼写错误如缺少重音、拼写错误等)、错误位置跨度、修正标注、带标签的错误句子、辅助标签列表和空格标记。数据集分为训练集、验证集和测试集,分别包含150,228、1,521和1,514个样本。

This dataset is designed for grammatical error detection and text correction tasks, containing original sentences, corrupted sentences with errors, tokenized words, error type labels (covering grammatical errors such as singular/plural, verb form, article usage, etc., and spelling errors like missing accents, misspellings, etc.), error span positions, correction annotations, tagged corrupted sentences, auxiliary tagged lists, and space markers. The dataset is split into training, validation, and test sets with 150,228, 1,521, and 1,514 examples respectively.

提供机构:
Ro551
二维码
社区交流群
二维码
科研交流群
商业服务