遇见数据集

OpenDataArena-scored-data

收藏
魔搭社区2026-04-28 更新2026-07-19 收录
官方服务:

资源简介:

# OpenDataArena-scored-data This repository contains a special **scored and enhanced** collection of over **47** original datasets for Supervised Fine-Tuning (SFT). These datasets were processed using the [OpenDataArena-Tool](https://github.com/OpenDataArena/OpenDataArena-Tool), a comprehensive suite of automated evaluation methods for assessing instruction-following datasets. The scored-dataset collection from the [OpenDataArena (ODA)]((https://arena.opendatalab.org.cn/)) reflects a large-scale, cross-domain effort to quantify dataset value in the era of large language models. ODA offers an open, transparent scoring pipeline that spans multiple domains—such as **instruction-following**, **coding**, **mathematics**, **science** and **reasoning**—and evaluates hundreds of datasets across 20+ scoring dimensions. What makes this collection special is that every subset: * is treated under a unified training-and-evaluation framework, * covers diverse domains and tasks (from code generation to math problems to general language instruction), * has been scored in a way that enables direct comparison of dataset “value” (i.e., how effective the dataset is at improving downstream model performance). > 🚀 **Project Mission**: All data originates from the **OpenDataArena** project, which is dedicated to transforming dataset value assessment from "guesswork" into a "science". This scored version provides rich, multi-dimensional scores for each instruction (question) and instruction-response pair (QA pair). This enables researchers to perform highly granular data analysis and selection, transforming the original datasets from simple text collections into structured, queryable resources for model training. ## 🚀 Potential Use Cases These multi-dimensional scores enable a powerful range of data processing strategies: * 🎯 **High-Quality Data Filtering**: Easily create a "gold standard" SFT dataset by filtering for samples with high `Deita_Quality` and high `Correctness`. * 📈 **Curriculum Learning**: Design a training curriculum where the model first learns from samples with low `IFD` (Instruction Following Difficulty) and gradually progresses to more complex samples with high `IFD`. * 🧐 **Error Analysis**: Gain deep insights into model failure modes and weaknesses by analyzing samples with high `Fail_Rate` or low `Relevance`. * 🧩 **Complexity Stratification**: Isolate questions with high `Deita_Complexity` or `Difficulty` to specifically test or enhance a model's complex reasoning abilities. * ⚖️ **Data Mixture Optimization**: When mixing multiple data sources, use `Reward_Model` scores or `Deita_Quality` as weighting factors to build a custom, high-performance training mix. ## 📊 Dataset Composition (Continuously Expanding...) | Subset | Count | |----------------------------------|-----------| | AM-Thinking-v1-Distilled-code | 324k | | AM-Thinking-v1-Distilled-math | 558k | | Code-Feedback | 66.4k | | DeepMath-309K | 309k | | EpiCoder-func-380k | 380k | | KodCode-V1 | 484k | | Magicoder-OSS-Instruct-75K-shuffle| 75.2k | | Magpie-Reasoning-V1-150K-code | 66.2k | | Magpie-Reasoning-V2-250K-CoT-QwQ-math | 235k | | Maths-College | 970k | | OpenCoder-LLM__opc-sft-stage2 | 436k | | OpenHermes-shuffle | 615k | | OpenR1-Math | 93.7k | | OpenThoughts | 114k | | OpenThoughts3-1_2M-all | 1.2M | | OpenThoughts3-1_2M-code | 250k | | OpenThoughts3-1_2M-math | 850k | | Python-Code-23k-ShareGPT | 22.6k | | QwQ-LongCoT | 133k | | QwQ-LongCoT-130K-math | 90.1k | | R1-Distill-SFT-math | 172k | | RabotniKuma_Fast-Math-R1-SFT_train| 7.9k | | Self-oss-Instruct | 50.7k | | allenai__tulu-3-sft-personas-code | 35k | | alpaca_gpt4 | 52k | | code290k_sharegpt | 289k | | code_instructions_122k | 122k | | code_instructions_18k | 18.6k | | frontedcookbook | 45.4k | | gretel-text-to-python-fintech-en-v1 | 22.5k | | herculesv1 | 463k | | hkust-nlp__dart-math-hard | 585k | | ise-uiuc__Magicoder-Evol-Instruct-110K | 111k | | m-a-p__CodeFeedback-Filtered-Instruction | 157k| | magpiepro_10k_gptmini | 10k | | mathfusionqa_all | 59.9k | | mathgpt4o200k | 200k | | mathplus | 894k | | numinamath-cot | 859k | | numinamath1_5 | 896k | | openmathinstruct-2 | 1M | | orca-agentinstruct-shuffle | 994k | | orca-math-word-problems-200k | 200k | | qihoo360_Light-R1-SFTData_train | 79.4k | | theblackcat102__evol-codealpaca-v1 | 111k | | tulu-3-sft-mixture | 939k | | webinstructcft | 50k | --- ## 💾 Data Format & Structure The dataset is provided in **JSON Lines (JSONL)** format. Each line is a JSON object with the following structure: ```json { "instruction": "The original instruction given to the model.", "output": "The model's response.", "Q_scores": { "Deita_Complexity": 3.5, "Thinking_Prob": null, "Difficulty": "medium", "Clarity": 0.9, ... }, "QA_scores": { "Deita_Quality": 5.0, "IFD": 2.8, "Reward_Model": 0.75, "Fail_Rate": 0.1, "Relevance": 1.0, ... "A_Length": 128 } } ```` * `instruction`: `string` - The original instruction. * `output`: `string` - The model's response. * `Q_scores`: `dict` - A dictionary containing all scores that **evaluate the instruction (Q) only**. * `QA_scores`: `dict` - A dictionary containing all scores that **evaluate the instruction-response pair (QA)**. **Note on `null` values**: Some scores may be `null`, indicating that the metric was not applicable to the given data sample (e.g., `Thinking_Prob` is specialized for math problems) or that the specific scorer was not run on that sample. ## 📊 Scoring Dimensions All scores are organized into two nested dictionaries. For a detailed explanation of each metric, please refer to the tables below. ### `Q_scores` (Instruction-Only) These metrics assess the quality and complexity of the `instruction` itself. | Metric | Type | Description | | :--- | :--- | :--- | | `Deita_Complexity` | Model-based | Estimates instruction complexity on a scale of 1-6. Higher scores indicate greater cognitive demand. | | `Thinking_Prob` | Model-based | (For math problems) Quantifies the necessity of deep reasoning. Higher values suggest harder problems. | | `Difficulty` | LLM-as-Judge | A specialized score for the difficulty of code or math questions. | | `Clarity` | LLM-as-Judge | Evaluates the instruction's clarity. | | `Coherence` | LLM-as-Judge | Evaluates the instruction's logical consistency. | | `Completeness` | LLM-as-Judge | Evaluates if the instruction is self-contained with enough information. | | `Complexity` | LLM-as-Judge | Evaluates the instruction's inherent complexity. | | `Correctness` | LLM-as-Judge | Evaluates the instruction's factual accuracy. | | `Meaningfulness` | LLM-as-Judge | Evaluates the instruction's practical value or significance. | ### `QA_scores` (Instruction-Answer Pair) These metrics assess the quality of the `output` in the context of the `instruction`. | Metric | Type | Description | | :--- | :--- | :--- | | `Deita_Quality` | Model-based | Estimates the overall quality of the instruction-response pair on a scale of 1-6. Higher scores indicate a better response. | | `IFD` | Model-based | (Instruction Following Difficulty) Quantifies how difficult it was for the model to follow the specific instruction. Higher scores indicate greater difficulty. | | `Reward_Model` | Model-based | A scalar reward score from the Skywork-Reward-Model. Higher scores indicate more preferable and aligned responses. | | `Fail_Rate` | Model-based | (For problems like math) Estimates the probability that a model will fail to produce a correct solution. | | `Relevance` | LLM-as-Judge | Evaluates whether the answer remains focused on the question. | | `Clarity` | LLM-as-Judge | Evaluates the response's clarity and ease of understanding. | | `Coherence` | LLM-as-Judge | Evaluates the response's logical consistency with the question. | | `Completeness` | LLM-as-Judge | Evaluates if the response completely addresses the query. | | `Complexity` | LLM-as-Judge | Evaluates the depth of reasoning reflected in the response. | | `Correctness` | LLM-as-Judge | Evaluates the response's factual accuracy. | | `Meaningfulness` | LLM-as-Judge | Evaluates if the response is insightful or valuable. | | `A_Length` | Heuristic | The number of tokens in the `output`, calculated using the `o200k_base` encoder. A simple measure of response length. | ## 💻 How to Use You can easily load any of the scored datasets (as a subset) using the 🤗 `datasets` library and filter it based on the scores. ```python from datasets import load_dataset # 1. Load a specific subset from the Hugging Face Hub # Replace "<subset_name>" with the name of the dataset you want, e.g., "alpaca" dataset_name = "<subset_name>" dataset = load_dataset("OpenDataArena/OpenDataArena-scored-data", name=dataset_name)['train'] print(f"Total samples in '{dataset_name}': {len(dataset)}") # 2. Example: How to filter using scores # Let's filter for a "high-quality and high-difficulty" dataset # - Response quality > 5.0 # - Instruction complexity > 4.0 # - Response must be correct high_quality_hard_data = dataset.filter( lambda x: x['QA_scores']['Deita_Quality'] is not None and \ x['QA_scores']['Correctness'] is not None and \ x['Q_scores']['Deita_Complexity'] is not None and \ x['QA_scores']['Deita_Quality'] > 5.0 and \ x['QA_scores']['Correctness'] > 0.9 and \ x['Q_scores']['Deita_Complexity'] > 4.0 ) print(f"Found {len(high_quality_hard_data)} high-quality & hard samples.") # 3. Access the first filtered sample if len(high_quality_hard_data) > 0: sample = high_quality_hard_data[0] print("\n--- Example Sample ---") print(f"Instruction: {sample['instruction']}") print(f"Output: {sample['output']}") print(f"QA Quality Score: {sample['QA_scores']['Deita_Quality']}") print(f"Q Complexity Score: {sample['Q_scores']['Deita_Complexity']}") ``` # 🌐 About OpenDataArena [OpenDataArena](https://arena.opendatalab.org.cn/) is an open research platform dedicated to **discovering, evaluating, and advancing high-quality datasets for AI post-training**. It provides a transparent, data-centric ecosystem to support reproducible dataset evaluation and sharing. **Key Features:** * 🏆 **Dataset Leaderboard** — helps researchers identify **the most valuable and high-quality datasets across different domains**. * 📊 **Detailed Evaluation Scores** — provides **comprehensive metrics** to assess data quality, complexity, difficulty etc. * 🧰 **Data Processing Toolkit** — [OpenDataArena-Tool](https://github.com/OpenDataArena/OpenDataArena-Tool) offers an open-source pipeline for dataset curation and scoring. If you find our work helpful, please consider **⭐ starring and subscribing** to support our research. ## 📚 Citation Information If you use this scored dataset collection in your work or research, please cite both the **original datasets** (that you use) and the **OpenDataArena-Tool**. **Citing this Dataset Collection** ```bibtex @dataset{opendataarena_scored_data_2025, author = {OpenDataArena}, title = {OpenDataArena-scored-data}, year = {2025}, url = {[https://huggingface.co/datasets/OpenDataArena/OpenDataArena-scored-data](https://huggingface.co/datasets/OpenDataArena/OpenDataArena-scored-data)} } ``` **Citing this Report** ```bibtex @article{cai2025opendataarena, title={OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value}, author={Cai, Mengzhang and Gao, Xin and Li, Yu and Lin, Honglin and Liu, Zheng and Pan, Zhuoshi and Pei, Qizhi and Shang, Xiaoran and Sun, Mengyuan and Tang, Zinan and others}, journal={arXiv preprint arXiv:2512.14051}, year={2025} } ``` **Citing the Scoring Tool (OpenDataArena-Tool)** ```bibtex @misc{opendataarena_tool_2025, author = {OpenDataArena}, title = {{OpenDataArena-Tool}}, year = {2025}, url = {[https://github.com/OpenDataArena/OpenDataArena-Tool](https://github.com/OpenDataArena/OpenDataArena-Tool)}, note = {GitHub repository}, howpublished = {\url{[https://github.com/OpenDataArena/OpenDataArena-Tool](https://github.com/OpenDataArena/OpenDataArena-Tool)}}, } ```

提供机构:
maas
创建时间:
2025-11-13
二维码
社区交流群
二维码
科研交流群
商业服务