官方服务:
资源简介:
trimmed data
应用场景:
创建时间:
2021-04-12
相关数据集
anon8231489123/ShareGPT_Vicuna_unfiltered
该数据集来源于ShareGPT对话,经过筛选后保留了约53k条英文对话。清理过程包括移除非英文对话、过多的Unicode字符、重复字符以及包含特定道德化短语的对话。数据集提供了两个版本:一个移除了包含Im sorry, but的对话,另一个保留了这些对话。数据集已准备好用于训练未经过滤的Vicuna模型。
Hugging Face2023-04-12 更新530
datajuicer/the-pile-pubmed-central-refined-by-data-juicer
--- license: apache-2.0 task_categories: - text-generation language: - en tags: - data-juicer - pretraining size_categories: - 1M<n<10M --- # The Pile -- PubMed Central (refined by Data-Juicer) A re
Hugging Face2023-10-23 更新230
mfeat-factors
**Author**: **Source**: Unknown - Date unknown **Please cite**: Binarized version of the original data set (see version 1). The multi-class target feature is converted to a two-class nominal
OpenML2014-10-04 更新30



