SupraLabs/chat-titles-12K
收藏资源简介:
Supra Titles 12K 是一个精选的聊天标题生成数据集,源自用于实验性 Supra Title 模型家族的训练流程。该数据集专门设计用于训练、微调和评估模型,这些模型能够从用户对话的第一条消息生成简洁、描述性的标题。与一般的指令遵循数据集不同,Supra Titles 12K 专注于单一任务:将用户消息转换为高质量聊天标题,准确捕捉对话的意图、主题或上下文,同时保持简洁和可读性。数据集构建自过滤的对话数据,原始数据来源于 `sam-mosaic/orca-gpt4-chatml` 数据集。用户消息被提取、清理,并通过多阶段质量流程处理,以提高一致性并减少噪声。生成过程包括数据提取、标题生成(使用 Qwen3.6-35B-A3B 模型)、清理和去重以及质量审查(使用 Qwen3.5-9B-Q8 模型)。每个数据集条目包含两个字段:user(从源数据集中提取的处理后的用户消息)和title(描述用户消息的生成标题)。
Supra Titles 12K is a curated chat title generation dataset derived from the training pipeline used for the experimental Supra Title model family. The dataset is designed specifically for training, fine-tuning, and evaluating models that generate concise, descriptive titles from a users first message in a conversation. Unlike general instruction-following datasets, Supra Titles 12K focuses on a single task: transforming a user message into a high-quality chat title that accurately captures the intent, topic, or context of the conversation while remaining concise and readable. The dataset was constructed from filtered conversation data originating from the `sam-mosaic/orca-gpt4-chatml` dataset. User messages were extracted, cleaned, and processed through a multi-stage quality pipeline designed to improve consistency and reduce noise. The generation process included data extraction, title generation (using Qwen3.6-35B-A3B), cleaning and deduplication, and quality review (using Qwen3.5-9B-Q8). Each dataset entry contains two fields: user (the processed user message extracted from the source dataset) and title (a generated title describing the users message).




