Tralalabs/CC-Clean-2026-04
收藏资源简介:
CC-Clean-2026-04是一个经过清洗的Common Crawl数据集子集,包含从Common Crawl的CC-MAIN-2026-04快照(2026年1月)中随机抽取的10,211个高质量英文网页文档。这些文档经过了多阶段的过滤流程,包括语言过滤(仅英文,最小置信度0.70)、NSFW内容过滤(包括显式词汇模式、域名黑名单和特定TLDs)以及质量过滤(基于长度、字符比例、行长度等启发式规则)。数据集以Parquet格式存储,适用于小型语言模型的预训练、自定义分词器训练以及语言建模研究。数据集还提供了详细的统计信息、过滤流程、使用意图、限制和伦理考虑等内容。
CC-Clean-2026-04 is a cleaned subset of the Common Crawl snapshot CC-MAIN-2026-04 (January 2026), containing 10,211 high-quality English web documents randomly sampled from 10 WET segments. The documents have undergone a multi-stage filtering pipeline including language filtering (English only, minimum confidence 0.70), NSFW content filtering (explicit word patterns, domain blocklists, and specific TLDs), and quality filtering (based on length, character ratios, line length heuristics, etc.). The dataset is stored in Parquet format and is suitable for pretraining small language models, training custom tokenizers, and research on language modeling. The README provides detailed statistics, filtering pipeline, intended use cases, limitations, and ethical considerations.




