登录后查看消息通知
搜索
常见问题
消息
登录
首页
/
数据集
/
skrishna/toxicity_preprop
skrishna/toxicity_preprop
收藏
Hugging Face
2023-10-05 更新
2024-03-04 收录
内容安全
数据预处理
数据链接:
https://hf-mirror.com/datasets/skrishna/toxicity_preprop
数据链接
链接失效反馈
官方服务:
问题咨询
购买咨询
在线客服
NEW
资源简介:
--- license: mit ---
许可证:MIT许可证(MIT License)
应用场景:
提供机构:
skrishna
原始信息汇总
数据集许可证
许可证类型
: MIT
相关数据集
Zyda
视觉语言模型
数据预处理
Zyda是由Zyphra创建的一个包含1.3万亿tokens的大型语言模型预训练数据集。该数据集整合了多个高质量的开源数据集,通过严格的过滤和去重过程,确保数据质量。Zyda不仅在性能上超越了其他开源数据集,如Dolma和RefinedWeb,还显著提高了Pythia系列模型的表现。数据集的创建过程包括综合多个数据源、应用先进的过滤和去重技术。Zyda的应用领域广泛,主要用于提升大型语言模型的训练
arXiv
2024-06-04 更新
178
0
Clean Cyclistic Data
共享单车
数据预处理
Clean Cyclistic Data (June 2020 through May 2021)
kaggle
2021-09-24 更新
9
0
skrishna/coin_flip_4_transformed
随机序列建模
概率实验模拟
--- dataset_info: features: - name: targets dtype: string - name: targets_vec sequence: int64 - name: inputs dtype: string - name: text dtype: string - name: label dtyp
Hugging Face
2024-03-03 更新
5
0
Data extracted from GitHub repositories (training and test data-sets)
GitHub项目
数据预处理
This dataset contains the SQL tables of the training and test datasets used in our experimentation. These tables contain the preprocessed textual data (in a form of tokens) extracted from each trainin
Mendeley Data
2024-01-31 更新
21
0
skrishna/coin_flip_16_transformed
--- dataset_info: features: - name: targets dtype: string - name: targets_vec sequence: int64 - name: inputs dtype: string - name: text dtype: string - name: label dtyp
Hugging Face
2024-03-20 更新
7
0
© 2023-2026 上海数据发展科技有限责任公司 版权所有
沪ICP备17003045号-15
沪公网安备31010402336585号
热门搜索
社区交流群
科研交流群
商业服务
数据资源
寻源服务
数据采集
标注服务
数据产品
代理销售
数据领域
凭证登记
数据产品
介绍推广