ZhejiangLab/CPT_Data_Pool
收藏资源简介:
CPT Data Pool是一个用于领域特定大语言模型持续预训练的大规模CPT语料库数据集。它作为F2D-LLM框架(一个端到端的领域特定LLM训练流程)的数据组件。数据集总规模约为2710亿个token,以jsonl文件格式存储,采用统一模式,涵盖数学推理、编程代码、学术知识和通用知识四大类别。所有数据源均标注了质量分数标签(如math-score、code-score、edu-score),并且采样比例经过实验验证以确保有效混合。具体构成包括:数学推理类340亿token(占35%),主要来源于Dolma 1.7-AlgebraicStack、Dolma 1.7-OpenWebMath和MegaMath-Web-Pro;编程代码类1000亿token(占20%),来源于StarCoder;学术知识类770亿token(占15%),来源于Dolma 1.7-RedPajama-arXiv和Pes2o;通用知识类600亿token(占30%),来源于Dolma 1.7-CC News、Dolma 1.7-Wiki、Dolma 1.7-Books、Dolma 1.7-MegaWika、Dolma 1.7-Reddit和Dolma 1.7-Flan。数据集聚合了多个公开可用的数据源,用户在使用或重新分发时必须遵守各原始数据源的许可协议。
CPT Data Pool is a large-scale CPT corpus dataset for continuous pre-training of domain-specific large language models (LLMs). It serves as the data component of the F2D-LLM framework, an end-to-end domain-specific LLM training pipeline. The total scale of the dataset is approximately 271 billion tokens, stored in JSONL file format with a unified schema, covering four major categories: mathematical reasoning, programming code, academic knowledge, and general knowledge. All data sources are annotated with quality score labels (e.g., math-score, code-score, edu-score), and the sampling ratio has been experimentally validated to ensure effective mixture. Its specific composition is as follows: 34 billion tokens (accounting for 35%) of mathematical reasoning data, mainly sourced from Dolma 1.7-AlgebraicStack, Dolma 1.7-OpenWebMath and MegaMath-Web-Pro; 100 billion tokens (accounting for 20%) of programming code data, sourced from StarCoder; 77 billion tokens (accounting for 15%) of academic knowledge data, sourced from Dolma 1.7-RedPajama-arXiv and Pes2o; 60 billion tokens (accounting for 30%) of general knowledge data, sourced from Dolma 1.7-CC News, Dolma 1.7-Wiki, Dolma 1.7-Books, Dolma 1.7-MegaWika, Dolma 1.7-Reddit and Dolma 1.7-Flan. The dataset aggregates multiple publicly available data sources, and users must comply with the license agreements of each original data source when using or redistributing it.




