Validation Data for the Seq2Turk Spell Check Model
收藏资源简介:
This is the validation set used to train the seq-to-seq RoBERTa model in our paper. We collected approximately 38 gigabytes of text from various internet sources from Turkish Wikipedia, Turkish OSCAR, and some news sites. From these data, the error-free ones were filtered out, and then data was generated using the functions that produce typos, as mentioned in the paper. In the dataset, the [ORG] token is used as a delimiter. The part written before [ORG] contains the sentence with typos, while the part after [ORG] contains the correctly written version of the sentence.
本数据集为本论文中用于训练序列到序列(seq-to-seq)RoBERTa模型的验证集。 我们从多个互联网数据源中收集了约38吉字节的文本,数据源涵盖土耳其语维基百科、土耳其语OSCAR以及部分新闻站点。我们首先从这批数据中筛选出无错误的文本,随后按照论文所述,通过拼写错误生成函数生成了目标数据集。 本数据集中以[ORG]这个Token作为分隔符。[ORG]之前的内容为带有拼写错误的句子,而[ORG]之后的内容则为该句子的正确拼写版本。




