遇见数据集

Qwen-3-1.7B-Reasoning-x500

收藏
魔搭社区2026-07-15 更新2026-07-15 收录
官方服务:

资源简介:

Credits to [LH-Tech AI](https://huggingface.co/LH-Tech-AI) for providing this dataset! # Qwen-3-1.7B-Reasoning-x500 ## Overview This is a high-quality synthetic dataset consisting of 500 diverse samples generated by **Qwen 3 1.7B**. The goal of this dataset is to provide clean, direct, and logical reasoning traces for distilling larger model capabilities into Small Language Models (SLMs) like my Apex models or those of CompactAI. ## Dataset Structure The data is provided in the **Alpaca format**: - `input`: The user prompt - `reasoning`: The reasoning output of the model - `output`: The response generated by Qwen 3 1.7B ## Generation method Generated using Qwen 3 1.7B via Ollama.<br> Prompts were designed to cover 25 distinct domains including logic, coding, science, daily life and much more. The focus was on eliminating conversational filler and maximizing information density and for teaching smaller models reasoning and better english sentence structure. ## Why this dataset? - **Anti-Refusal Focus:** Prompts were answered directly without "preachy" AI disclaimers, so you can train your model with the real output of the model. - **Diverse Domains:** Covers Coding, Science, Philosophy, Ethics, and Logic. - **Optimized for SLMs:** Perfect for finetuning models like Gemma 4-E2B, Llama 3.x, or Phi-4. ## How to use You can load this dataset using the Hugging Face `datasets` library: ```python from datasets import load_dataset dataset = load_dataset("LH-Tech-AI/Qwen-3-1.7B-with-Reasoning-x500") ```

提供机构:
maas
创建时间:
2026-04-28
二维码
社区交流群
二维码
科研交流群
商业服务