遇见数据集

CofeAI/NanoData

收藏
Hugging Face2024-06-11 更新2024-06-12 收录
官方服务:

资源简介:

--- license: other license_name: other license_link: LICENSE task_categories: - text-generation language: - en size_categories: - 100B<n<1T --- ### Dataset Description To facilitate researchers to use [NanoLM](https://github.com/cofe-ai/nanoLM?tab=readme-ov-file) for comparative analysis across different model designs, we build a curated pre-training dataset from those of existing large-scale models (i.e., Llama, Falcon, GPT-3). It covers diverse domains to improve the generalization capabilities of the resultant models. #### Dataset Creation The data is mainly post-processed and filtered from [RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T) and [RedPajamaV2](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2). We develop a series of cleaning steps to remove redundant formatting, garbled characters, formula errors, duplicated paragraphs, low-quality text, and other unwanted content. After interleaved deduplication on document level of each independent subset, we finally obtain a high-quality dataset. #### Dataset Summary | Dataset | Num Tokens (B) | | -------------- | -------------- | | CommonCrawl | 67.00 | | C4 | 15.00 | | Wikipedia (En) | 5.14 | | Books | 4.48 | | ArXiv | 2.50 | | StackExchange | 2.00 | | Total | 97.12 | We release the data with approximate 100B tokens. Furthermore, we recommend users to add code dataset such as [Starcode](https://huggingface.co/datasets/bigcode/starcoderdata), [The Stack V2](https://huggingface.co/datasets/bigcode/the-stack-v2-dedup) to enrich model's performance on code and reasoning. ### Citation To cite NanoLM, please use: ``` @misc{yao2024nanolm, title={nanoLM: an Affordable LLM Pre-training Benchmark via Accurate Loss Prediction across Scales}, author={Yiqun Yao and Siqi fan and Xiusheng Huang and Xuezhi Fang and Xiang Li and Ziyi Ni and Xin Jiang and Xuying Meng and Peng Han and Shuo Shang and Kang Liu and Aixin Sun and Yequan Wang}, year={2024}, eprint={2304.06875}, archivePrefix={arXiv}, primaryClass={cs.CL} } ``` ### Acknowledgement The data is mainly curated and filtered from [RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T) and [RedPajamaV2](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2). We extend our gratitude to the original authors for their innovative work and for making it available to the community. ### License The code of NanoLM used to process the dataset and loss prediction is licensed under the Apache 2.0 license. For curated data, please refer to the licenses of the original ones. * [Common Crawl Foundation Terms of Use](https://commoncrawl.org/terms-of-use) * [C4 license](https://huggingface.co/datasets/allenai/c4#license) * Books: [the_pile_books3 license](https://huggingface.co/datasets/defunct-datasets/the_pile_books3#licensing-information) and [pg19 license](https://huggingface.co/datasets/deepmind/pg19#licensing-information) * [ArXiv Terms of Use](https://info.arxiv.org/help/api/tou.html) * [Wikipedia License](https://huggingface.co/datasets/legacy-datasets/wikipedia#licensing-information) * [StackExchange license on the Internet Archive](https://archive.org/details/stackexchange)

--- license: 其他 license_name: 其他 license_link: LICENSE task_categories: - 文本生成 language: - 英语 size_categories: - 1000亿 < 令牌数 < 1万亿 --- ### 数据集说明 为助力研究人员使用[NanoLM](https://github.com/cofe-ai/nanoLM?tab=readme-ov-file)开展不同模型设计的对比分析,我们从现有大规模模型(即Llama、Falcon、GPT-3)的预训练数据集中构建了经过精选的预训练数据集。本数据集覆盖多元领域,旨在提升所得模型的泛化能力。 #### 数据集构建流程 本数据集主要源自[RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T)与[RedPajamaV2](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2)数据集,经后处理与筛选得到。我们设计了一系列数据清洗流程,用于移除冗余格式、乱码字符、公式错误、重复段落、低质量文本等无效内容。在对每个独立子集进行文档级交叉去重后,最终得到高质量的预训练数据集。 #### 数据集概览 | 数据集名称 | 令牌数(十亿) | | ---------------- | -------------- | | CommonCrawl | 67.00 | | C4 | 15.00 | | 英文维基百科(Wikipedia (En)) | 5.14 | | 图书语料(Books) | 4.48 | | ArXiv | 2.50 | | StackExchange | 2.00 | | 总计 | 97.12 | 本数据集的发布规模约为1000亿令牌。此外,我们建议用户补充代码类数据集,例如[Starcode](https://huggingface.co/datasets/bigcode/starcoderdata)与[The Stack V2](https://huggingface.co/datasets/bigcode/the-stack-v2-dedup),以提升模型在代码与推理任务上的性能表现。 ### 引用方式 若需引用NanoLM,请使用以下格式: @misc{yao2024nanolm, title={nanoLM: an Affordable LLM Pre-training Benchmark via Accurate Loss Prediction across Scales}, author={Yiqun Yao and Siqi fan and Xiusheng Huang and Xuezhi Fang and Xiang Li and Ziyi Ni and Xin Jiang and Xuying Meng and Peng Han and Shuo Shang and Kang Liu and Aixin Sun and Yequan Wang}, year={2024}, eprint={2304.06875}, archivePrefix={arXiv}, primaryClass={cs.CL} } ### 致谢 本数据集主要源自[RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T)与[RedPajamaV2](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2)数据集的精选与筛选工作。在此,我们向原始数据集的作者致敬,感谢其创新性的研究工作与开源共享精神。 ### 许可证声明 用于数据集处理与损失预测的NanoLM代码采用Apache 2.0许可证进行授权。 对于本精选数据集,请遵循其原始数据集的许可证要求: * [Common Crawl 基金会使用条款](https://commoncrawl.org/terms-of-use) * [C4 许可证](https://huggingface.co/datasets/allenai/c4#license) * 图书语料:遵循[the_pile_books3 许可证](https://huggingface.co/datasets/defunct-datasets/the_pile_books3#licensing-information)与[pg19 许可证](https://huggingface.co/datasets/deepmind/pg19#licensing-information) * [ArXiv 使用条款](https://info.arxiv.org/help/api/tou.html) * [维基百科 许可证](https://huggingface.co/datasets/legacy-datasets/wikipedia#licensing-information) * [互联网档案馆中的StackExchange许可证](https://archive.org/details/stackexchange)

提供机构:
CofeAI
原始信息汇总

数据集描述

数据集创建

本数据集是为了支持研究人员使用NanoLM进行不同模型设计的比较分析而构建的。数据主要从RedPajama和RedPajamaV2中经过一系列清洗步骤处理和过滤得到,包括去除冗余格式、乱码、公式错误、重复段落、低质量文本等。

数据集总结

数据集 令牌数量(B)
CommonCrawl 67.00
C4 15.00
Wikipedia (En) 5.14
Books 4.48
ArXiv 2.50
StackExchange 2.00
总计 97.12

数据集包含约100B令牌。建议用户添加如Starcode和The Stack V2等代码数据集以增强模型在代码和推理方面的性能。

许可证

数据集的原始数据遵循各自原始数据的许可证。具体包括:

搜集汇总
数据集介绍
构建方式
NanoData数据集源自对现有大规模预训练语料库的精炼与整合,其构建基础主要依托于RedPajama与RedPajamaV2两大公开数据集。为提升数据质量,研究团队设计了一套系统化的清洗流程,包括移除冗余格式、乱码字符、公式错误、重复段落及低质量文本等干扰内容。在此基础上,对各独立子集执行文档级别的交错去重操作,最终汇聚成一个涵盖CommonCrawl、C4、Wikipedia、Books、ArXiv及StackExchange六大领域、总规模接近100B token的高质量预训练数据集。这一构建策略既保留了原始数据的多样性,又通过严格过滤确保了语料的纯净度与可用性。
特点
该数据集的核心特点在于其精炼的结构与广泛的领域覆盖。通过从Llama、Falcon、GPT-3等主流模型所依赖的语料中筛选与重组,NanoData在缩减规模的同时,依然保留了跨域知识的丰富性,从而有效提升下游模型的泛化能力。其组成涵盖网页文本、百科全书、学术论文、书籍及问答社区等多种类型,为语言模型的预训练提供了均衡且多元的语义素材。此外,数据集的文档级去重机制显著降低了冗余信息,使得每一token的语义贡献更加高效,尤其适用于对模型设计进行对比分析的研究场景。
使用方法
NanoData专为配合NanoLM框架进行语言模型预训练与对比分析而设计,用户可直接将其作为训练语料加载至NanoLM环境中使用。数据集以标准的文本生成任务格式组织,兼容常见的预训练流程。为增强模型在代码理解与逻辑推理方面的表现,研究团队建议用户额外引入Starcode或The Stack V2等代码数据集进行补充训练。使用时需注意各子集原始数据的许可协议,确保合规引用。数据集的轻量化特性使其在资源受限的场景下仍能支撑有效的模型训练与评估实验。
背景与挑战
背景概述
在大规模语言模型(LLM)研究领域,预训练数据集的构建与质量直接影响模型性能的评估与对比。为应对不同模型架构间缺乏统一基准的困境,CofeAI团队于2024年发布了NanoData数据集,该数据集由Yao等人基于NanoLM项目创建,旨在为跨规模损失预测提供标准化预训练语料。研究团队从RedPajama及RedPajamaV2等公开数据源中筛选并精炼出约1000亿词元的多元领域文本,涵盖CommonCrawl、C4、Wikipedia、Books、ArXiv及StackExchange等子集,从而提升模型的泛化能力。NanoData的推出为轻量级LLM预训练基准设立了新标杆,推动了模型设计比较的可复现性与公平性。
当前挑战
NanoData所面临的核心挑战包括:其一,在领域层面,现有开源数据集常存在冗余格式、乱码、重复段落及低质量文本,需通过精细清洗与去重流程确保语料纯净度,以支撑模型跨领域泛化;其二,构建过程中,从大规模原始数据(如RedPajama的1万亿词元)中高效过滤噪声、平衡各子集规模(如CommonCrawl占67%主导)并保持文档级去重一致性,对计算资源与算法设计提出严苛要求;此外,代码类数据的缺失(如Starcode、The Stack V2)限制了模型在推理任务上的表现,提示未来需补充结构化语料以完善能力覆盖。
常用场景
经典使用场景
NanoData 数据集专为大规模语言模型的预训练而设计,其经典使用场景在于为研究者提供一个经过严格清洗与去重的高质量语料库,用于对比分析不同模型架构(如 Llama、Falcon、GPT-3)的训练效果。该数据集整合了 CommonCrawl、C4、Wikipedia、Books、ArXiv 及 StackExchange 等多源异构数据,覆盖广泛的知识领域,从而有效提升模型的泛化能力。研究者可借助 NanoData 在统一的训练基准上评估模型设计差异,探索预训练数据质量与规模对下游性能的影响。
解决学术问题
NanoData 数据集着力解决了预训练语料质量参差不齐与冗余度高的学术难题。通过系统性的清洗流程——包括去除格式混乱、乱码、公式错误、重复段落及低质文本——并实施文档级交叉去重,显著降低了数据噪声与偏差。这为语言模型预训练提供了可靠的数据基准,使研究者能够更准确地归因模型性能提升的来源,避免因数据污染而误导结论。其意义在于推动了预训练数据工程的可重复性与标准化,为后续高效模型训练奠定了方法论基础。
衍生相关工作
NanoData 数据集衍生了多项经典工作,其中最核心的是其配套的 NanoLM 基准框架,该框架通过跨尺度的准确损失预测实现了低成本的大模型预训练评估。受此启发,后续研究进一步探索了数据筛选策略的优化(如基于质量分数的动态采样)以及多源语料的融合方法。此外,NanoData 的清洗与去重流程已被社区采纳为通用预处理模板,催生了针对特定领域(如科学文献、法律文本)的高质量数据构建工作,推动了预训练数据工程的系统化发展。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务