遇见数据集

mashey/dv-synthetic-errors-lg

收藏
Hugging Face2026-05-21 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含迪维希语(Dhivehi)的句子对:原始正确的句子及其合成的含错误版本。数据集总计约720万个句子对,分为训练集(80%,约580万对)、验证集(10%,约72万对)和测试集(10%,约72万对)。每个句子对包括以下字段:original(原始正确句子)、error(含合成错误的句子)和error_types(引入的错误类型)。合成错误通过基于规则的转换生成,包括变音符号错误、字符错误、时态标记错误、格标记错误、词序问题、一致性错误和后缀错误。

This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts. The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of a correct Dhivehi sentence and the same sentence with synthetic errors. Errors are synthetically generated using rule-based transformations including diacritic errors, character mistakes, tense marker errors, case marking mistakes, word order issues, agreement errors, and suffix errors.

提供机构:
mashey
二维码
社区交流群
二维码
科研交流群
商业服务