DeMix_Corpora
收藏资源简介:
DeMix Corpora 是一个全面、高质量、大规模且经过精心混合的资源,可直接用于预训练。该数据集包含15T原始令牌和22T混合令牌,旨在解决现有通用语料库缺乏领域特定强度,而专业语料库又不适用于通用预训练的问题。DeMix Corpora 通过提供经过验证的最佳数据混合比例,平衡了通用语言能力和在复杂任务(如数学推理和代码生成)上的强大性能。数据集结构包括三个部分:用于复现结果的组件模型和参考模型(DeMix_reproduce)、纯数据集(pure_dataset)以及预训练三个阶段的混合样本(mixture_sample)。数据集的创建过程包括从异构开源资源中收集数据,经过全局去重、模糊去重、困惑度过滤、FastText 过滤和中文通用领域过滤等多个步骤,最终按三个阶段进行混合,其中高质量数学和代码数据的比例逐渐增加。
DeMix Corpora is a comprehensive, high-quality, large-scale and meticulously mixed resource directly applicable for pre-training. It contains 15T raw tokens and 22T mixed tokens, aiming to solve the problem that existing general-purpose corpora lack domain-specific depth while specialized corpora are unsuitable for general pre-training. DeMix Corpora balances general language proficiency and robust performance on complex tasks such as mathematical reasoning and code generation by providing validated optimal data mixing ratios. The dataset structure includes three parts: component models and reference models for result reproduction (DeMix_reproduce), the pure dataset (pure_dataset), and mixed samples for the three pre-training stages (mixture_sample). The dataset creation process involves collecting data from heterogeneous open-source resources, followed by multiple steps including global deduplication, fuzzy deduplication, perplexity filtering, FastText filtering and Chinese general-domain filtering, before finally mixing the data across three stages, where the proportion of high-quality mathematical and code data gradually increases.




