遇见数据集

Kukedlc/slopus-v3-train-lt2k

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

slopus-v3-train-lt2k是一个包含17121个预格式化示例的数据集,专为推理和思维链任务设计。数据使用Qwen3-4B-Thinking-2507模板进行格式化,并经过过滤(确保token数小于2048)和去重处理。数据来源于五个公开子集:roman(7515例,来自Roman1111111/claude-opus-4.6-10000x的reasoning字段)、jirafa(6644例,来自angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k的think inline字段)、nohurry(2233例,来自nohurry/Opus-4.6-Reasoning-3000x-filtered)、teich(639例,来自TeichAI/Claude-Opus-4.6-Reasoning-887x的thinking字段)和gpt55(90例,来自armand0e/gpt-5.5-chat)。每行数据包含text(完整渲染的聊天模板,包括系统、用户和助手消息,使用<think>...</think>格式)、n_tokens(token计数)、source(来源数据集)和category(源元数据中的类别)。该数据集适用于监督微调(SFT)和推理任务,可通过HuggingFace的datasets库直接加载使用。

slopus-v3-train-lt2k is a dataset containing 17,121 pre-formatted examples, designed for reasoning and chain-of-thought tasks. The data is formatted using the Qwen3-4B-Thinking-2507 template, filtered to have fewer than 2048 tokens, and deduplicated. It is sourced from five public subsets: roman (7,515 examples, from Roman1111111/claude-opus-4.6-10000x, reasoning field), jirafa (6,644 examples, from angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k, think inline field), nohurry (2,233 examples, from nohurry/Opus-4.6-Reasoning-3000x-filtered), teich (639 examples, from TeichAI/Claude-Opus-4.6-Reasoning-887x, thinking field), and gpt55 (90 examples, from armand0e/gpt-5.5-chat). Each row includes text (the fully rendered chat template with system, user, and assistant messages using <think>...</think> format), n_tokens (token count), source (origin dataset), and category (category from source metadata). The dataset is suitable for supervised fine-tuning (SFT) and reasoning tasks, and can be loaded directly via the HuggingFace datasets library.

提供机构:
Kukedlc
二维码
社区交流群
二维码
科研交流群
商业服务