SupraLabs/reasoning-summaries-61k
收藏资源简介:
推理摘要数据集是一个用于训练和评估推理摘要模型的数据集,包含约61K个清理样本。其核心目标是将冗长、混乱或过于详细的推理轨迹(例如来自Qwen3.5+、Gemma4、GLM-5.2等模型的推理链或块)转换为简短、用户友好的摘要,以解释模型在做什么,而不暴露完整的原始推理链。数据集不仅关注缩短文本,还包括将非常简短、不完整或粗略的推理轨迹扩展为更完整的摘要(约50词),以保持输出的一致性、可读性和稳定性。这有助于AI产品在保持界面干净和安全的同时,让用户更透明地理解推理模型。数据集使用多个AI模型生成(主要使用Ling 2.6 Flash和GPT-5.4 Mini,约20%由DeepSeek V4 Flash生成),以增加结构、措辞和推理风格的多样性。每个样本围绕推理轨迹及其摘要版本构建,可能包括标题、副标题、类别和简短任务描述等元数据。数据集遵循严格格式:输入为原始推理链/块,输出包括标题、副标题、摘要和当前任务。它覆盖了正常推理案例以及大量边缘案例和粗略案例,如混乱推理轨迹、不完整推理、过于冗长推理、非常简短推理、嘈杂格式以及需要标准化为用户友好摘要的推理。数据集特别适用于模型执行复杂任务(如编程、数学、写作、研究、规划、调试、工具使用或多步问题解决)的应用场景,帮助将混乱、冗长、不完整或不均匀的推理轨迹转换为用户可理解的干净摘要。
The Reasoning Summary dataset is designed for training and evaluating reasoning summarization models, containing approximately 61K cleaned samples. Its primary goal is to convert long, messy, or highly detailed reasoning traces (e.g., reasoning chains or blocks from models like Qwen3.5+, Gemma4, GLM-5.2) into short, user-friendly summaries that explain what the model was doing without exposing the full raw reasoning chain. The dataset not only focuses on shortening text but also includes examples where very short, incomplete, or rough reasoning traces are expanded into a more complete summary of around 50 words, helping outputs stay consistent, readable, and stable. This is useful for AI products that want reasoning models to feel more transparent to users while keeping the interface clean and safe. The dataset was generated using multiple AI models (mainly Ling 2.6 Flash and GPT-5.4 Mini, with around 20% generated by DeepSeek V4 Flash) to add variety in structure, wording, and reasoning styles. Each sample is structured around a reasoning trace and a summarized version of that trace, and may include additional metadata such as a title, subtitle, category, and short task description. The dataset follows a strict format: Input is the raw reasoning chain/block, and output includes title, subtitle, summary, and current task. It covers mostly normal reasoning cases along with a considerable amount of edge cases and rough cases, such as chaotic reasoning traces, incomplete reasoning, overly verbose reasoning, very short reasoning, noisy formatting, and reasoning that needs to be normalized into a stable user-facing summary. The dataset is especially useful for applications where the model performs complex tasks like programming, math, writing, research, planning, debugging, tool use, or multi-step problem solving, helping transform chaotic, verbose, incomplete, or uneven reasoning traces into clean summaries that users can actually understand.




