simonko912/autocorrect-16M
收藏官方服务:
资源简介:
这是一个简单的自动纠错数据集,格式为:<|start|><|user|>错误拼写<|user|><|model|>正确拼写<|model|><|end|>。数据集单词组成包括:手动/基本单词1,386个、维基短语9,975个、NLTK单词88,639个,总计100,000个单词。错误拼写通过最近键匹配、双击、缺失字符、交换两个字符、数字和随机插入等方法生成,每个单词或句子生成164个错误拼写示例。建议使用至少5,000个示例,示例已随机排序。
A simple autocorrect dataset with the format: <|start|><|user|>misspelled<|user|><|model|>correct<|model|><|end|>. The dataset word composition includes: manual/basic words (1,386), wiki phrases (9,975), NLTK words (88,639), totaling 100,000 words. Misspellings are generated using nearest key match, double tap, missing char, swap 2 chars, numbers, and random insert methods, with 164 misspelling examples per word/sentence. It is recommended to use at least 5,000 examples, and the examples are already randomly sorted.
提供机构:
simonko912


