open-index/open-markdown-v2
收藏资源简介:
Open Markdown是一个大规模网络文本数据集,基于Common Crawl构建。该数据集通过一个处理流程,从原始HTML中提取主要内容,转换为干净的Markdown格式,并将结果打包为Parquet文件,同时保留WARC元数据以支持可追溯性。当前版本包含爬虫CC-MAIN-2026-25,约2,039,988,496个文档,分布在100,000个分片中;处理了约329.3 TB的原始HTML,生成约8.9 TB的干净Markdown,数据量减少了97.3%。数据集采用Open Data Commons Attribution License (ODC-By) v1.0许可,旨在降低训练和检索高质量网络数据的门槛,适用于文本生成和特征提取等多语言任务。
Open Markdown is a large-scale web text dataset built from Common Crawl. Every page goes through a pipeline that extracts the main content from raw HTML, converts it to clean Markdown, and packages the result into Parquet files with WARC metadata for traceability. The dataset currently includes crawl CC-MAIN-2026-25 with ~2,039,988,496 documents across 100,000 shards. It processed ~329.3 TB of raw HTML into ~8.9 TB of clean Markdown, a 97.3% reduction. Released under the Open Data Commons Attribution License (ODC-By) v1.0, it aims to lower the barrier to training and retrieval of high-quality web data and is suitable for multilingual tasks such as text generation and feature extraction.




