ray0rf1re/FineWeb-Nano
收藏资源简介:
--- license: odc-by language: - en size_categories: - 10B<n<100B --- # FineWeb-Nano ## Dataset Description **FineWeb-Nano** is a highly curated, premium subset extracted from [`nampdn-ai/mini-fineweb`](https://huggingface.co/datasets/nampdn-ai/mini-fineweb). ### How "The Best" Was Determined This dataset was created programmatically by streaming the original dataset and sorting chunks based on a rigorous quality scoring algorithm. The heuristic heavily favors: 1. High `language_score` (if provided by the upstream extraction). 2. Optimal document length (penalizing abnormally short snippets and excessively long, unformatted dumps). 3. Strong structural coherence suitable for pre-training Small Language Models (SLMs) and LLMs. **FineWeb-Nano** represents the elite tier of data. It takes the **BEST 29.8 Gigabytes** directly from the top percentiles of the `Fineweb-Tiny` subset. This is the ultimate, hyper-distilled dataset for rapid model experimentation. ### License This dataset is released under the **Open Data Commons Attribution License (ODC-By) v1.0** to perfectly match the upstream Fineweb mini licensing constraints.
--- 许可证:odc-by 语言: - 英语 规模类别: - 10B < 数据量 < 100B --- # FineWeb-Nano ## 数据集描述 **FineWeb-Nano** 是从 [`nampdn-ai/mini-fineweb`](https://huggingface.co/datasets/nampdn-ai/mini-fineweb) 中提取的经过精心筛选的优质子集。 ### 「最优」数据集的判定依据 本数据集通过流式处理原始数据集,并基于严格的质量评分算法对数据块进行排序生成。该启发式筛选规则优先考量以下几点: 1. 较高的`language_score`(若上游提取流程已提供该指标)。 2. 最优文档长度(对过短的片段与过长的无格式转储均予以惩罚)。 3. 适配小语言模型(Small Language Models, SLMs)与大语言模型(Large Language Models, LLMs)预训练的强结构连贯性。 **FineWeb-Nano** 属于顶级精英层级的数据子集,它从`Fineweb-Tiny`子集的百分位顶端直接选取,最终得到**29.8吉字节**的最优数据。本数据集是面向快速模型实验的极致精炼版预训练数据集。 ### 许可协议 本数据集采用**开放数据共同体署名许可协议(Open Data Commons Attribution License, ODC-By)v1.0**发布,与上游FineWeb mini数据集的许可约束保持一致。



