SupraLabs/chat-titles-filtered-115K
收藏资源简介:
Supra Titles 115K 是一个精选的聊天标题生成数据集,源自用于实验性 Supra Title 模型家族的训练流程。该数据集专门设计用于训练、微调和评估模型,这些模型能够从对话中的用户首条消息生成简洁、描述性的标题。与通用的指令遵循数据集不同,Supra Titles 115K 专注于单一任务:将用户消息转换为高质量的聊天标题,准确捕捉对话的意图、主题或上下文,同时保持简洁和可读性。数据集包含约 115,000 个过滤样本,提供了大量高质量的用户-标题对,并强调一致性、可读性和相关性。每个数据集条目包含两个字段:user(从源数据集中提取的已处理用户消息)和title(描述用户消息的生成标题)。数据集创建过程包括数据提取、使用 Qwen3.6-35B-A3B 生成标题、清理和去重以及使用 Qwen3.5-9B-Q8 进行质量审查。
Supra Titles 115K is a curated chat title generation dataset derived from the training pipeline of the experimental Supra Title model family. This dataset is specifically designed for training, fine-tuning, and evaluating models that can generate concise, descriptive titles from the initial user messages in conversations. Unlike general instruction-following datasets, Supra Titles 115K focuses on a single task: converting user messages into high-quality chat titles that accurately capture the intent, topic, or context of the conversation while remaining concise and readable. The dataset contains approximately 115,000 filtered samples, providing a large volume of high-quality user-title pairs, with an emphasis on consistency, readability, and relevance. Each dataset entry includes two fields: `user` (the processed user message extracted from the source dataset) and `title` (the generated title that describes the user message). The dataset creation process encompasses data extraction, title generation using Qwen3.6-35B-A3B, cleaning and deduplication, as well as quality review conducted with Qwen3.5-9B-Q8.




