SupraLabs/chat-titles-unfiltered-150K
收藏资源简介:
Supra Title 150K是一个大规模聊天标题生成数据集,源自Supra Title模型家族的训练流程。该数据集专门设计用于训练、微调和评估模型,以根据用户对话中的第一条消息生成简洁、描述性的标题。与一般的指令跟随数据集不同,Supra Title 150K专注于单一任务:将用户消息转换为高质量聊天标题,准确捕捉对话的意图、主题或上下文,同时保持简洁和可读性。数据集包含约150,000个样本,提供比小规模标题生成数据集更广泛的覆盖范围,并保留了真实的用户提示和对话场景。每个数据条目包含两个字段:user(从源数据集中提取并处理的用户消息)和title(描述用户消息的生成标题)。数据集创建过程包括数据提取、标题生成、清理和去重、质量审查等步骤,确保数据质量。
Supra Title 150K is a large-scale chat title generation dataset derived from the training pipeline used for the experimental Supra Title model family. The dataset is designed specifically for training, fine-tuning, and evaluating models that generate concise, descriptive titles from a users first message in a conversation. Unlike general instruction-following datasets, Supra Title 150K focuses on a single task: transforming a user message into a high-quality chat title that accurately captures the intent, topic, or context of the conversation while remaining concise and readable. Containing approximately 150,000 samples, the dataset provides significantly broader coverage than smaller title-generation datasets while preserving realistic user prompts and conversational scenarios. Each dataset entry contains two fields: user (the processed user message extracted from the source dataset) and title (a generated title describing the users message). The dataset creation process includes data extraction, title generation, cleaning and deduplication, and quality review steps to ensure data quality.




