遇见数据集

sutra-10M

收藏
魔搭社区2026-04-28 更新2026-07-19 收录
官方服务:

资源简介:

# Sutra 10M Pretraining Dataset A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 7,252 educational entries totaling approximately 10 million tokens. ## Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: - **Clear pedagogical structure**: Content follows proven educational patterns - **Cross-domain connections**: Concepts are linked across disciplines - **Varied complexity levels**: From foundational (level 4) to advanced (level 10) - **Quality-controlled generation**: All entries meet minimum quality thresholds ## Dataset Statistics | Metric | Value | |--------|-------| | Total Entries | 7,252 | | Total Tokens | 10,007,381 | | Avg Tokens/Entry | 1,379 | | Quality Score Range | 60.0 - 70.0 | | Average Quality Score | 65.6 | ### Domain Distribution | Domain | Count | Percentage | |--------|-------|------------| | Science | 1,728 | 23.8% | | Mathematics | 1,585 | 21.9% | | Programming & Systems | 1,322 | 18.2% | | Logic & Reasoning | 666 | 9.2% | | Engineering | 633 | 8.7% | | Applied Fields | 424 | 5.8% | | Language & Communication | 365 | 5.0% | | Social Sciences | 276 | 3.8% | | Humanities | 253 | 3.5% | ### Content Type Distribution | Content Type | Count | Percentage | |--------------|-------|------------| | Reasoning Demonstration | 2,034 | 28.0% | | Meta-Learning | 1,580 | 21.8% | | Concept Introduction | 1,535 | 21.2% | | Synthesis | 1,320 | 18.2% | | Cross-Domain Bridge | 783 | 10.8% | ### Complexity Distribution | Level | Count | Percentage | |-------|-------|------------| | Level 4-5 | 503 | 6.9% | | Level 6 | 1,197 | 16.5% | | Level 7 | 2,060 | 28.4% | | Level 8 | 2,135 | 29.4% | | Level 9-10 | 1,357 | 18.7% | ## Data Fields Each entry contains the following fields: | Field | Description | |-------|-------------| | `id` | Unique identifier for the entry | | `concept_name` | The concept being taught | | `domain` | Primary knowledge domain | | `content_type` | Type of pedagogical content | | `text` | The main educational content | | `quality_score` | Quality assessment score (0-100) | | `information_density` | Measure of information per token | | `complexity_level` | Difficulty level (1-10) | | `prerequisites` | Required prior knowledge | | `builds_to` | Advanced concepts this enables | | `cross_domain_connections` | Links to other domains | | `token_count` | Number of tokens in the entry | | `quality_assessment` | Detailed quality breakdown | ## Generation Details - **Generator**: Sutra Framework - **Quality Threshold**: Minimum score of 60.0 - **Generation Date**: June 2025 ## Intended Use This dataset is designed for: - **LLM Pretraining**: High-quality educational content for foundational model training - **Domain-specific fine-tuning**: Subset by domain for specialized training - **Educational AI research**: Studying pedagogical content generation ## Quality Assurance All entries underwent multi-dimensional quality assessment including: - Clarity and readability - Information density - Pedagogical structure - Reasoning completeness - Practical utility - Connection richness ## Related Datasets - [sutra-100M](https://huggingface.co/datasets/codelion/sutra-100M): 100M token pretraining dataset - [sutra-30k-seeds](https://huggingface.co/datasets/codelion/sutra-30k-seeds): Instruction prompts for post-training - [sutra-magpie-sft](https://huggingface.co/datasets/codelion/sutra-magpie-sft): SFT dataset generated from seed prompts ## Citation ```bibtex @article{sharma2026sutra, title={Scaling Pedagogical Pretraining: From Optimal Mixing to 10 Billion Tokens}, author={Sharma, Asankhaya}, year={2026}, url={https://huggingface.co/blog/codelion/scaling-pedagogical-pretraining-10-billion-tokens} } ``` ## License Apache 2.0

提供机构:
maas
创建时间:
2026-03-04
二维码
社区交流群
二维码
科研交流群
商业服务