ecreeth/english-1b
收藏资源简介:
这是一个高质量的10亿词英语语料库(实际总词数为0.09B),专为训练小型语言模型设计。它结合了教育内容、叙事散文、百科全书事实和对话式网络文本,数据来源包括FineWeb-Edu(高质量教育网络内容)、OpenWebText(GPT-2训练数据的复现)、WikiText-103(经过验证的英语维基百科文章)、TinyStories(针对小型模型的叙事语法)、Project Gutenberg(1919年之前的文学书籍)以及SQuAD v2 + Alpaca(问答和指令遵循模式)。数据集经过先进的清理流程处理,包括Gopher启发式(单词和句子长度检查以确保自然散文)、压缩比率(zlib)过滤(去除重复垃圾和随机噪声)、自然语言密度(基于停用词的散文验证)、精确去重(基于哈希去除跨来源的重复文章)以及样板移除(剥离Gutenberg免责声明、页码和目录伪影)。格式为Parquet(分块),主要语言为英语,许可证为开放(基于来源混合)。
This dataset is a high-quality 1-billion-word English corpus designed for training small language models. It combines educational content, narrative prose, encyclopedic facts, and conversational web text. Total words: 0.09B, format: Parquet (chunked), primary language: English, license: Open (Mixed based on sources). Data sources include FineWeb-Edu, OpenWebText, WikiText-103, TinyStories, Project Gutenberg, and SQuAD v2 + Alpaca. The dataset was processed through a cleaning pipeline with Gopher heuristics, compression ratio (zlib) filtering, natural language density checks, exact deduplication, and boilerplate removal.



