遇见数据集

mashey/dv-synthetic-errors

收藏
Hugging Face2026-05-21 更新2026-05-31 收录
官方服务:

资源简介:

迪维希语文本错误纠正数据集,包含正确句子和合成生成的错误句子。该数据集旨在测试迪维希语错误纠正模型和工具。数据集结构为输入输出对:correct字段为原始正确句子,incorrect字段为带有合成错误的句子。数据集创建基于迪维希语文章收集,通过字符和变音符号替换生成错误,错误率为每单词30%概率。

Dhivehi text error correction dataset containing correct sentences and synthetically generated errors. The dataset aims to test Dhivehi language error correction models and tools. The dataset structure consists of input-output pairs: correct for original correct sentences and incorrect for sentences with synthetic errors. It was created using a collection of Dhivehi articles with error generation via character and diacritic substitutions at a 30% probability per word.

提供机构:
mashey
二维码
社区交流群
二维码
科研交流群
商业服务