LLM_multimodal
收藏资源简介:
LLM_multimodal 是一个为 LLM D6 模型系列设计的官方训练语料库,包含约 3000 亿个高质量 token,旨在平衡语言多样性、数学推理和编程能力。数据集涵盖英语、俄语以及多种编程语言,内容来自通用网络文本、学术/数学论文、代码仓库和对话数据。它分为两个主要子集:预训练(pre-train)子集,用于基础训练,包含大型文本文档,关注高质量、去重后的数据,如网络爬取文本、百科全书、文学作品、源代码以及 LaTeX 格式的方程和 STEM 学术文本;微调(fine-tune)子集,为对话和指令遵循任务精心策划,包含高质量提示-完成对,支持标准 ChatML 提示模板,涵盖物理常识推理、通用推理、多任务准确性和俄语理解等任务。该数据集专为 LLaMA-3 架构(特别是 D6 0.8B 参数定制版本)设计,并使用了一个针对英语、西里尔字母和编程语法优化的定制分词器进行训练。
LLM_multimodal is an official training corpus designed for the LLM D6 model family, containing approximately 300 billion high-quality tokens. It aims to strike a balance across linguistic diversity, mathematical reasoning capabilities, and programming proficiency. The corpus covers English, Russian, and multiple programming languages, with content sourced from general web text, academic and mathematical papers, code repositories, and conversational data. It is divided into two primary subsets: the pre-train subset, used for foundational model training, which comprises large-scale textual documents focusing on high-quality, deduplicated data including web-crawled text, encyclopedias, literary works, source code, LaTeX-formatted equations, and STEM academic texts; the fine-tune subset, meticulously curated for conversational and instruction-following tasks, which contains high-quality prompt-completion pairs, supports the standard ChatML prompt template, and covers tasks such as physical commonsense reasoning, general reasoning, multi-task accuracy, and Russian language comprehension. This corpus is specifically tailored for the LLaMA-3 architecture, particularly the custom 0.8B-parameter variant of the D6 model, and was trained using a custom tokenizer optimized for English, Cyrillic script, and programming syntax.




