anabury/conv_dataset
收藏官方服务:
资源简介:
升级版Cosmopedia v0.2数据集,包含提示信息(prompt)、文本内容(text)、种子数据(seed_data)等字段。数据集分为训练集,共有约195万个样本。该数据集适用于文本分类任务,使用Apache-2.0许可。具体数据集描述信息在README中未提供。
Upgrade Cosmopedia v0.2 dataset, including fields such as prompt, text, seed_data, etc. The dataset is split into a training set with approximately 1.95 million samples. It is suitable for text classification tasks and is licensed under Apache-2.0. Detailed dataset description information is not provided in the README.
提供机构:
anabury


