遇见数据集

WithinUsAI/GLM_5.1_Thinking_Distilled_5k

收藏
Hugging Face2026-05-25 更新2026-07-22 收录
官方服务:

资源简介:

该数据集名为GLM-5.1-Thinking Distilled Dataset,包含5,000个高质量推理轨迹,用于将链式思考能力蒸馏到GLM-5.1-Thinking或兼容语言模型中。每个示例由一个问题、详细的多步推理轨迹(平均长度为305个字符)和一个简洁的最终响应组成。数据集旨在教导模型在回答之前逐步思考,模仿一个强大语言模型的内部推理过程。数据覆盖多个领域,包括算术、文字问题、逻辑、决策分析、数据分析、金融、科学、伦理、谜题、语言和算法等,确保推理的多样性和结构性。数据集以JSONL格式存储,每个示例包含instruction、reasoning_trace和response字段。

This dataset, named GLM-5.1-Thinking Distilled Dataset, contains 5,000 high-quality reasoning trajectories for distilling chain-of-thought capabilities into GLM-5.1-Thinking or compatible large language models. Each example consists of a question, a detailed multi-step reasoning trace with an average length of 305 characters, and a concise final response. This dataset is designed to teach models to think step-by-step before generating answers, mimicking the internal reasoning process of a state-of-the-art large language model. It covers diverse domains including arithmetic, word problems, logic, decision analysis, data analysis, finance, science, ethics, puzzles, linguistics, and algorithms, ensuring the diversity and structural rigor of reasoning. The dataset is stored in JSONL format, with each example containing the fields `instruction`, `reasoning_trace`, and `response`.

提供机构:
WithinUsAI
二维码
社区交流群
二维码
科研交流群
商业服务