open-index/open-markdown
收藏资源简介:
Open Markdown是一个大规模的网络文本数据集,源自Common Crawl的非营利性网络爬取项目。该数据集通过将原始HTML内容转换为干净的Markdown格式,并打包成Parquet文件,同时保留了WARC元数据以便追溯。数据集目前包含CC-MAIN-2026-12爬取的数据,共773,597,197个文档,分布在44,746个分片中。处理过程中,100.3 TB的原始HTML被压缩为6.5 TB的干净Markdown,减少了93.5%。数据集采用Open Data Commons Attribution License (ODC-By) v1.0许可,与Common Crawl相同。
Open Markdown is a large-scale web text dataset built from Common Crawl, a non-profit that crawls the web and freely provides its archives and datasets to the public. The dataset processes raw HTML into clean Markdown, packaging the result into Parquet files with useful WARC metadata for traceability. It currently includes crawl CC-MAIN-2026-12 with 773,597,197 documents across 44,746 shards, processing 100.3 TB of raw HTML into 6.5 TB of clean Markdown — a 93.5% reduction. The dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0, the same license used by Common Crawl.




