遇见数据集

WithinUsAI/Llama_4_Maverick_Distilled_5k

收藏
Hugging Face2026-05-25 更新2026-07-22 收录
官方服务:

资源简介:

Llama_4_Maverick_Distilled – 5,000 Reasoning Traces是一个高质量蒸馏数据集,旨在模拟Llama 4 Maverick类模型的思维和推理轨迹。该数据集包含5,000个无重复的示例,覆盖10个推理领域(每个领域500个示例),包括数学、逻辑推理、编程、科学、阅读理解、常识、规划、定量比较、因果推理和伦理推理。其目的是捕获Llama_4_Maverick_Distilled特有的显式思维链模式:陈述问题、分解为编号步骤、展示中间工作,然后提供简洁的最终答案。数据集文件为JSONL和CSV格式,每个示例包含id、domain、prompt、thinking_trace和response字段。它用于训练学生大型语言模型以复现Llama_4_Maverick_Distilled风格的逐步推理,支持监督蒸馏,推荐在thinking_trace上使用0.7的损失权重,在response上使用0.3的损失权重以鼓励忠实推理。数据集是合成生成的,不包含来自Llama_4_Maverick_Distilled的直接输出,使用时需遵守Llama 4社区许可证。局限性包括推理轨迹是模板化的合成推理、深度限于2-4步以提升训练效率,且不涵盖完整Llama 4 Maverick中的工具使用、长上下文RAG或多模态推理。

Llama_4_Maverick_Distilled – 5,000 Reasoning Traces is a high-quality distilled dataset designed to simulate the thinking and reasoning trajectories of models in the Llama 4 Maverick series. This dataset contains 5,000 unique examples spanning 10 reasoning domains (500 examples per domain), including mathematics, logical reasoning, programming, science, reading comprehension, common sense, planning, quantitative comparison, causal reasoning, and ethical reasoning. Its purpose is to capture the unique explicit chain-of-thought pattern inherent to Llama_4_Maverick_Distilled: stating the problem, breaking it down into numbered steps, displaying intermediate work, and then providing a concise final answer. The dataset is available in JSONL and CSV formats, with each example containing the fields: id, domain, prompt, thinking_trace, and response. It is used to train student large language models to replicate the step-by-step reasoning style of Llama_4_Maverick_Distilled, supporting supervised distillation. A loss weight of 0.7 on thinking_trace and 0.3 on response is recommended to encourage faithful reasoning. The dataset is synthetically generated and does not contain direct outputs from Llama_4_Maverick_Distilled; users must comply with the Llama 4 Community License when using it. Limitations include that the reasoning traces are template-based synthetic reasoning, limited to 2-4 steps in depth to enhance training efficiency, and do not cover tool usage, long-context RAG, or multimodal reasoning present in the full Llama 4 Maverick model.

提供机构:
WithinUsAI
二维码
社区交流群
二维码
科研交流群
商业服务