dclm-baseline-1.0-parquet
收藏资源简介:
DCLM-baseline 是一个包含4万亿个标记和30亿个文档的预训练数据集,由DCLM团队精心策划,使用英语,并根据CC-by-4.0许可证发布。该数据集源自Common Crawl,经过一系列清洗、过滤和去重步骤处理,特别适用于作为DCLM基准的研究基线。
DCLM-baseline is a pretraining dataset containing 4 trillion tokens and 3 billion documents. Meticulously curated by the DCLM team, it consists of English-language content and is released under the CC BY 4.0 license. Originating from Common Crawl, this dataset has undergone a series of cleaning, filtering, and deduplication steps, and is specifically designed to serve as the research baseline for DCLM.
数据集概述
基本信息
- 名称: DCLM-baseline
- 语言: 英语
- 许可: CC-by-4.0
- 大小: 4T token / 3B document
数据集描述
DCLM-baseline 是一个用于预训练语言模型的大型数据集,包含4万亿个token和30亿个文档,旨在为语言模型基准测试提供强大的性能。
数据集来源
- 团队: DCLM Team
- 来源: Common Crawl
- 论文: DataComp-LM: In search of the next generation of training sets for language models
- 代码: GitHub
使用场景
- 直接使用: 作为DCLM基准测试的研究基线,展示数据筛选在训练高性能语言模型中的重要性。
- 非适用场景: 不适用于训练生产就绪模型或特定领域(如代码和数学)的模型。
数据集创建
- 创建目的: 展示DCLM测试床在开发高质量训练集方面的有效性,作为数据筛选策略的证明。
- 数据处理: 通过一系列清洗、过滤和去重步骤从原始Common Crawl数据(DCLM-Pool)中创建。
偏见、风险和限制
数据集可能包含Common Crawl数据中的偏见,且在代码和数学任务上的表现有限。仅适用于研究目的。
引用
bibtex @misc{li2024datacomplm, title={DataComp-LM: In search of the next generation of training sets for language models}, author={Jeffrey Li and others}, year={2024}, eprint={2406.11794}, archivePrefix={arXiv}, primaryClass={cs.LG} }




