遇见数据集

UltraData-SFT-2605

收藏
魔搭社区2026-07-15 更新2026-06-07 收录
官方服务:

资源简介:

# UltraData-SFT-2605 <div align="center"> <img src="assets/logo.png" width="400"/> </div> <p align="center"> <a href="https://huggingface.co/collections/openbmb/ultradata">📦 UltraData Collection</a> | <a href="https://ultradata.openbmb.cn/">🌐 UltraData</a> | <a href="https://huggingface.co/collections/openbmb/minicpm5">🤗 MiniCPM5 Series</a> </p> <p align="center"> English | <a href="https://huggingface.co/datasets/openbmb/UltraData-SFT-2605/blob/main/README_ZH.md">中文</a> </p> ## 📚 Introduction ***UltraData-SFT-2605*** is the full set of core-domain SFT data used in the **post-training of [MiniCPM5-1B-SFT](https://huggingface.co/openbmb/MiniCPM5-1B-SFT)** within the **[MiniCPM5-1B](https://huggingface.co/collections/openbmb/minicpm5)** series, and a key representative of **L3 refined data** in the [UltraData](https://ultradata.openbmb.cn/) [L0-L4 tiered data management framework](https://arxiv.org/pdf/2602.09003). It covers math, code, knowledge, instruction following, and other core domains, containing **over 15 million Deep Thinking and Non-thinking training samples**. Every sample passes through a **High-Quality SFT Data Management Pipeline**—spanning query construction and filtering, answer quality control, training-based validation, and benchmark decontamination—to ensure that data entering final training is clean and genuinely effective. In every domain and at every difficulty level, **UltraData-SFT-2605** constructs both **Deep Thinking** and **Non-thinking** data: - **Non-thinking data** targets the model's ability to respond directly in scenarios where users need fast, immediate answers. - **Deep Thinking data** targets reasoning, planning, and verification capabilities required for complex tasks. This dual coverage ensures the model receives appropriate training signals across diverse usage scenarios—from quick, conversational responses to multi-step reasoning chains. ## 📢 What's New - **[2026.05.28]** The [***UltraData-SFT-2605***](https://huggingface.co/datasets/openbmb/UltraData-SFT-2605) dataset is released! The full set of core-domain SFT data used in the **post-training of [MiniCPM5-1B-SFT](https://huggingface.co/openbmb/MiniCPM5-1B-SFT)** within the **[MiniCPM5-1B](https://huggingface.co/collections/openbmb/minicpm5)** series, and a key representative of **L3 refined data** in the [UltraData](https://ultradata.openbmb.cn/) [L0-L4 tiered data management framework](https://arxiv.org/pdf/2602.09003). It covers math, code, knowledge, instruction following, and other core domains, containing **over 15 million Deep Thinking and Non-thinking training samples**. 🚀🚀🚀 - **[2026.05.25]** ***[MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B) is released!***, the first model in the MiniCPM5 series. It is a dense 1B Transformer built for on-device, local deployment, and resource-constrained scenarios, reaching 1B-class open-source SOTA. UltraData-SFT-2605 serves as the core SFT dataset for MiniCPM5-1B. - **[2026.02.08]** The [***UltraData***](https://ultradata.openbmb.cn/) platform is now live, introducing the [L0-L4 tiered data management framework](https://arxiv.org/pdf/2602.09003). 🔍🔍🔍 ## 🏗️ High-Quality SFT Data Management Pipeline <div align="center"> <img src="assets/ultradata-sft-2605-pipeline.png" width="600"/> </div> UltraData-SFT-2605 follows a six-step high-quality SFT data management pipeline: ### 1. Open-source Data Validation & Query Filtering For data sourced from the open-source community, we perform **query-level filtering** focusing on: - Whether the question has genuine training value - Whether the intent is clear - Whether capability coverage is sufficient - Whether the difficulty is reasonable - Whether the question pushes the model to learn real, useful skills rather than repeating low-value templates ### 2. Internal Query Construction For internally constructed data, we design query sources and construction methods around different capability types: - **Knowledge data** is constructed based on exam syllabi and assessment points - **Instruction-following data** is constructed from atomic instructions - **Self Evolution & Augmentation**: Self-evolving question evolution and augmentation ### 3. High-quality Pre-training Format L3 Data Filtering Post-training also incorporates high-quality pre-training-format data, such as **L3 textbook or wiki style content**, to strengthen the model's knowledge organization, expression, and generalization. For this category, we further filter by structural integrity, information density, and learnability—ensuring suitability for post-training rather than naively converting pre-training text into Q&A format. ### 4. Answer Quality Filtering We focus on: - Whether the answer is **correct** - Whether the expression is **clear** - Whether the format meets requirements For Deep Thinking data, we additionally verify that the reasoning process aids the model in learning **problem decomposition and intermediate verification**, rather than piling up lengthy, vacuous "thinking text". ### 5. Single-data Validation All data undergoes single-data validation. We use a **70% candidate data + 30% instruction-following data** mix for rapid SFT validation training, defaulting to 3 epochs with a training budget capped at **20B tokens**. In this step, we mainly focus on: - Validate the actual capability gain of each data category - Search for the optimal epoch count by combining evaluation results across checkpoints - Determine the data's role in the final training mix ### 6. Benchmark Decontamination All data undergoes decontamination testing against existing benchmarks, minimizing the risk of training-evaluation overlap. This ensures the model's capability gains come from **real data quality improvements**, not from memorizing test items. ## 🎯 Capability Coverage UltraData-SFT-2605 covers seven core capability domains. Most domains provide both Deep Thinking and Non-thinking variants, while multilingual domains are released as Non-thinking only. | Domain (config) | Description | |:---|:---| | **Math** | Mathematical reasoning, problem solving, formula derivation | | **Code** | Code generation, debugging, algorithmic problem solving | | **Knowledge** | Factual knowledge, conceptual understanding, exam-oriented Q&A | | **Chinese-general** | General-purpose Chinese conversational and reasoning data | | **IF** | Instruction following — multi-constraint instructions, format compliance | | **Multi-lang-Math** | Multilingual mathematical reasoning data | | **Multi-lang-Knowledge** | Multilingual knowledge / world-fact Q&A | ## 📊 Dataset Statistics After the full data management pipeline, **UltraData-SFT-2605 contains 15M+ samples in total**. The breakdown by domain and thinking mode: | Domain (config) | Deep Thinking (`think`) | Non-thinking (`no_think`) | Total | |:---|---:|---:|---:| | **Math** | 2,499,830 | 2,999,644 | 5,499,474 | | **Code** | 2,788,465 | 3,000,000 | 5,788,465 | | **Knowledge** | 499,667 | 800,000 | 1,299,667 | | **Chinese-general** | 499,954 | 500,000 | 999,954 | | **IF** | 199,883 | 199,991 | 399,874 | | **Multi-lang-Math** | — | 549,230 | 549,230 | | **Multi-lang-Knowledge** | — | 499,514 | 499,514 | | **Total** | **6,487,799** | **8,548,379** | **15,036,178** | Raw sample counts (before quality filtering) are slightly higher; the table above shows the **final post-data management counts** released here. ## 🚀 Quick Start You can load the dataset directly from Hugging Face: Each config corresponds to a capability domain. Within each config, `think` and `no_think` are two splits (when both are available). ```python from datasets import load_dataset # Math: Deep Thinking split ds = load_dataset("openbmb/UltraData-SFT-2605", "Math", split="think") # Math: Non-thinking split ds = load_dataset("openbmb/UltraData-SFT-2605", "Math", split="no_think") # Code / Knowledge / IF / Chinese-general — same usage ds = load_dataset("openbmb/UltraData-SFT-2605", "Code", split="think") # Multi-lang-Knowledge / Multi-lang-Math: only no_think is available ds = load_dataset("openbmb/UltraData-SFT-2605", "Multi-lang-Math", split="no_think") ``` Available configs: `Chinese-general`, `IF`, `Knowledge`, `Code`, `Math`, `Multi-lang-Knowledge`, `Multi-lang-Math`. ## 💡 Use Cases UltraData-SFT-2605 is not just a "large-scale" SFT dataset—it is a **high-quality post-training resource** that has gone through filtering, decontamination, and training-based validation. It is suitable for: - **Training small-parameter models**: A proven SFT recipe for compact models, validated on MiniCPM5-1B. - **Domain fine-tuning**: Selective use of math, code, knowledge, or instruction-following slices for targeted capability enhancement. - **Mix-ratio research**: Studying how Deep Thinking vs. Non-thinking data ratios affect model behavior, latency, and downstream task performance. - **Benchmarking post-training methodology**: A reference dataset for comparing post-training approaches under controlled conditions. ## 📖 Citation If you find **UltraData-SFT-2605** useful in your research, please consider citing: ```bibtex @misc{ultradata-sft-2605, title={UltraData-SFT-2605}, author={OpenBMB}, year={2026}, url={https://huggingface.co/datasets/openbmb/UltraData-SFT-2605}, publisher={Hugging Face} } ``` ## 📜 License This project is released under the [Apache 2.0](./LICENSE) license. ***UltraData-SFT-2605*** incorporates queries from multiple source datasets; in addition to this repository's license, users must also review and comply with the **license terms of each upstream dataset**. **No unauthorized unchanged redistribution:** Without prior written permission from the original authors (or this organization), any institution, organization, or third-party platform is strictly prohibited from directly reposting, mirroring, re-hosting, or commercially repackaging and republishing any artifacts of this project in any form.

提供机构:
maas
创建时间:
2026-05-29
二维码
社区交流群
二维码
科研交流群
商业服务