Zyda - 包含1.3万亿Token的开源预训练数据集
收藏资源简介:
Zyda数据集是由Zyphra公司创建的一个大型语言模型预训练数据集。该数据集通过整合多个开源数据集并进行深度处理来构建,包含了1.3万亿Token,其质量接近商业语料。Zyda数据集的创建过程包括了严格的过滤和去重处理,以保持和提高从原始数据集中派生出的质量。实验结果表明,使用Zyda训练的语言模型在多项评估任务上,性能优于其他同类数据集,如Dolma、FineWeb和RefinedWeb。Zyda的发布为开源社区提供了一个高质量的、大规模的预训练语料库,为开源语言模型研究奠定数据基础。
The Zyda dataset is a large language model pre-training corpus developed by Zyphra Corporation. It is built by integrating multiple open-source datasets and undergoing in-depth processing, boasting a scale of 1.3 trillion Tokens and a quality comparable to commercial corpora. The creation process of the Zyda dataset incorporates strict filtering and deduplication operations to preserve and improve the quality of data derived from its original sources. Experimental results demonstrate that language models trained on the Zyda dataset outperform those trained on other similar datasets including Dolma, FineWeb, and RefinedWeb across a wide range of evaluation tasks. The release of the Zyda dataset offers the open-source community a high-quality, large-scale pre-training corpus, laying a solid data foundation for open-source language model research.
数据集概述
基本信息
- 数据集名称: Zyda
- 许可证: Open Data Commons License (ODC-BY)
- 任务类别: 文本生成
- 语言: 英语
- 大小类别: 大于1TB
数据集结构
- 配置名称: default
- 分割:
- 名称: train
- 样本数量: 1,594,197,267
配置详情
- 默认配置:
- 数据文件路径: data///*
- 其他配置:
- zyda_no_starcoder: data/zyda_no_starcoder//
- zyda_arxiv_only: data/zyda_no_starcoder/zyda_arxiv/*
- zyda_c4-en_only: data/zyda_no_starcoder/c4_en/*
- zyda_peS2o_only: data/zyda_no_starcoder/zyda_peS2o/*
- zyda_pile-uncopyrighted_only: data/zyda_no_starcoder/zyda_pile-uncopyrighted/*
- zyda_refinedweb_only: data/zyda_no_starcoder/zyda_refinedweb/*
- zyda_slimpajama_only: data/zyda_no_starcoder/zyda_slimpajama/*
- zyda_starcoder_only: data/zyda_starcoder//
数据集描述
- 数据集来源: 由七个高质量的开源数据集组成
- 数据字段:
text: 训练文本source: 文本来源filtering_features: 用于过滤的预计算特征值(转换为JSON字符串)source_other: 来源数据集的元数据(转换为JSON字符串)
数据收集和处理
- 处理步骤: 过滤和去重
- 过滤方法: 使用手工调整的过滤器,来源包括C4、RedPajama和Gopher等
- 去重方法: 使用minhash近似去重,基于13-gram和Jaccard相似度阈值0.4
数据集组件
| 组件 | 下载大小 (GB) | 文档数量 (百万) | 令牌数量 (十亿) |
|---|---|---|---|
| zyda_refinedweb_only | 1,712.4 | 920.5 | 564.8 |
| zyda_c4-en_only | 366.7 | 254.5 | 117.5 |
| zyda_slimpajama_only | 594.7 | 142.3 | 242.3 |
| zyda_pile-uncopyrighted_only | 189.4 | 64.9 | 82.9 |
| zyda_peS2o_only | 133.7 | 35.7 | 53.4 |
| zyda_arxiv_only | 8.3 | 0.3 | 4.7 |
| zyda_starcoder_only | 299.5 | 176.1 | 231.3 |
| 总计 | 3,304.7 | 1,594.2 | 1,296.7 |
许可证信息
- 许可证: Open Data Commons License (ODC-BY)
- 使用条款: 使用此数据集还需遵守原始数据源的许可证协议和使用条款
引用信息
@misc{tokpanov2024zyda, title={Zyda: A 1.3T Dataset for Open Language Modeling}, author={Yury Tokpanov and Beren Millidge and Paolo Glorioso and Jonathan Pilault and Adam Ibrahim and James Whittington and Quentin Anthony}, year={2024}, eprint={2406.01981}, archivePrefix={arXiv}, primaryClass={cs.CL} }




