遇见数据集

prli/wiki-full_draft_chopped-rare-qwen-alpha1-00_perturb

收藏
Hugging Face2026-04-27 更新2026-05-03 收录
官方服务:

资源简介:

该数据集包含文本数据,每个样本包含原始文本以及通过多种扰动方法(如同义词替换、随机删除、字符大小写变化等)生成的变体,适用于自然语言处理任务中的数据增强或鲁棒性测试。数据集仅包含验证分割,共4000个样本。

This dataset contains textual data. Each sample includes the original text and variants generated via multiple perturbation methods such as synonym replacement, random deletion, character case variation, etc. It is applicable to data augmentation or robustness testing in natural language processing (NLP) tasks. The dataset only has a validation split, with a total of 4000 samples.

提供机构:
prli
二维码
社区交流群
二维码
科研交流群
商业服务