遇见数据集

simonko912/autocorrect-16M

收藏
Hugging Face2026-05-16 更新2026-05-31 收录
官方服务:

资源简介:

这是一个简单的自动纠错数据集,格式为:<|start|><|user|>错误拼写<|user|><|model|>正确拼写<|model|><|end|>。数据集单词组成包括:手动/基本单词1,386个、维基短语9,975个、NLTK单词88,639个,总计100,000个单词。错误拼写通过最近键匹配、双击、缺失字符、交换两个字符、数字和随机插入等方法生成,每个单词或句子生成164个错误拼写示例。建议使用至少5,000个示例,示例已随机排序。

A simple autocorrect dataset with the format: <|start|><|user|>misspelled<|user|><|model|>correct<|model|><|end|>. The dataset word composition includes: manual/basic words (1,386), wiki phrases (9,975), NLTK words (88,639), totaling 100,000 words. Misspellings are generated using nearest key match, double tap, missing char, swap 2 chars, numbers, and random insert methods, with 164 misspelling examples per word/sentence. It is recommended to use at least 5,000 examples, and the examples are already randomly sorted.

提供机构:
simonko912
二维码
社区交流群
二维码
科研交流群
商业服务