遇见数据集

sutra-1B

收藏
魔搭社区2026-04-28 更新2026-07-19 收录
官方服务:

资源简介:

# Sutra 1B Pretraining Dataset A high-quality pedagogical dataset designed for LLM pretraining, containing 948,709 educational entries totaling over 1 billion tokens. ## Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: - **Clear pedagogical structure**: Content follows proven educational patterns - **Cross-domain connections**: Concepts are linked across disciplines - **Varied complexity levels**: From foundational (level 1) to advanced (level 10) - **Quality-controlled generation**: All entries meet minimum quality thresholds - **Diverse content types**: 33 different pedagogical formats ## Dataset Statistics | Metric | Value | |--------|-------| | Total Entries | 948,709 | | Total Tokens | 1,016,114,311 | | Avg Tokens/Entry | 1,071 | | Quality Score Range | 60.0 - 85.0 | | File Size | 6.4 GB | ### Domain Distribution | Domain | Entries | Tokens | Percentage | |--------|---------|--------|------------| | Applied Fields | 355,195 | 228.7M | 22.5% | | Programming & Systems | 136,777 | 175.2M | 17.2% | | Science | 103,385 | 137.1M | 13.5% | | Mathematics | 97,009 | 134.4M | 13.2% | | Logic & Reasoning | 115,567 | 118.2M | 11.6% | | Engineering | 85,794 | 85.1M | 8.4% | | Humanities | 27,331 | 76.4M | 7.5% | | Social Sciences | 16,849 | 40.5M | 4.0% | | Language & Communication | 10,802 | 20.6M | 2.0% | ### Content Type Distribution (Top 15) | Content Type | Count | Percentage | |--------------|-------|------------| | Concept Introduction | 336,974 | 35.5% | | Reasoning Demonstration | 135,529 | 14.3% | | Code Implementation | 118,534 | 12.5% | | Technical Documentation | 59,801 | 6.3% | | Tutorial | 42,857 | 4.5% | | Cross-Domain Bridge | 30,179 | 3.2% | | Worked Examples | 28,124 | 3.0% | | QA Pairs | 23,812 | 2.5% | | Common Misconceptions | 22,013 | 2.3% | | Meta-Learning | 16,573 | 1.7% | | Synthesis | 16,034 | 1.7% | | Prerequisite Scaffolding | 14,186 | 1.5% | | Code Explanation | 13,108 | 1.4% | | Diagnostic Assessment | 12,058 | 1.3% | | Code Debugging | 9,620 | 1.0% | ## Data Sources This dataset combines multiple high-quality sources: | Source | Description | |--------|-------------| | Sutra Pedagogical Generation | Synthetically generated educational content using the Sutra framework | | FineWeb Filtered | Quality-filtered content from HuggingFace FineWeb | | DCLM Rephrased | Pedagogically rephrased content from DCLM baseline | | Nemotron-CC v2.1 | High-quality curated content from NVIDIA Nemotron Common Crawl | ## Data Fields Each entry contains the following fields: | Field | Description | |-------|-------------| | `id` | Unique identifier for the entry | | `concept_name` | The concept being taught | | `domain` | Primary knowledge domain | | `content_type` | Type of pedagogical content | | `text` | The main educational content | | `quality_score` | Quality assessment score (0-100) | | `information_density` | Measure of information per token | | `complexity_level` | Difficulty level (1-10) | | `token_count` | Number of tokens in the entry | | `prerequisites` | Required prior knowledge (optional) | | `builds_to` | Advanced concepts this enables (optional) | | `cross_domain_connections` | Links to other domains (optional) | ## Data Cleaning The dataset underwent comprehensive cleaning: - **Duplicate Removal**: 43,971 exact duplicates and 564 near-duplicates removed - **Quality Filtering**: 109 short/low-quality entries removed - **Text Normalization**: Whitespace normalized, encoding issues fixed - **Field Validation**: All entries validated for required fields ## Generation Details - **Framework**: Sutra Pedagogical Dataset Generator - **Quality Threshold**: Minimum score of 60.0 - **Generation Date**: January 2026 - **Cleaning Pipeline**: Multi-stage deduplication with hash-based and prefix-based detection ## Intended Use This dataset is designed for: - **LLM Pretraining**: High-quality educational content for foundational model training - **Domain-specific fine-tuning**: Subset by domain for specialized training - **Educational AI research**: Studying pedagogical content generation - **Curriculum learning**: Progressive complexity for staged training ## Quality Assurance All entries underwent multi-dimensional quality assessment including: - Clarity and readability - Information density - Pedagogical structure - Reasoning completeness - Practical utility - Cross-domain connection richness ## Related Datasets - [sutra-100M](https://huggingface.co/datasets/codelion/sutra-100M): 100M token pretraining dataset - [sutra-10M](https://huggingface.co/datasets/codelion/sutra-10M): 10M token pretraining dataset - [sutra-30k-seeds](https://huggingface.co/datasets/codelion/sutra-30k-seeds): Instruction prompts for post-training - [sutra-magpie-sft](https://huggingface.co/datasets/codelion/sutra-magpie-sft): SFT dataset generated from seed prompts ## Citation ```bibtex @article{sharma2026sutra, title={Scaling Pedagogical Pretraining: From Optimal Mixing to 10 Billion Tokens}, author={Sharma, Asankhaya}, year={2026}, url={https://huggingface.co/blog/codelion/scaling-pedagogical-pretraining-10-billion-tokens} } ``` ## License Apache 2.0

提供机构:
maas
创建时间:
2026-03-04
二维码
社区交流群
二维码
科研交流群
商业服务