Literature-zh
收藏资源简介:
这个数据集是由公共爬虫收集的中文书籍、论文、法律文件和专利组成的复合数据集。经过数据清洗,移除了含有非拉丁、非中文字符比例超过2%的文本,以及含有大量特殊字符的文本,同时将繁体中文转换为了简体中文。数据集中的文档还经过语言质量评估,移除了质量较低的文档。数据集包含47,087,808个样本,磁盘大小为214G的parquet文件。
This composite dataset consists of Chinese books, academic papers, legal documents and patents collected through public crawlers. Subsequent data cleaning steps included removing texts with over 2% proportion of non-Latin and non-Chinese characters, texts containing excessive special characters, as well as converting all Traditional Chinese content to Simplified Chinese. Additionally, all documents in the dataset underwent language quality evaluation, and low-quality documents were removed. The dataset contains 47,087,808 samples and is stored as Parquet files with a total disk size of 214 GB.
Literature-zh 数据集概述
数据集简介
- 类型:中文书籍、论文、法律文件和专利的复合数据集
- 数据来源:Common Crawl
- 许可证:Apache-2.0
数据处理流程
数据清洗
- 移除包含超过2%非拉丁、非中文字符的文本
- 移除包含大量特殊字符的文本
- 将繁体中文转换为简体中文
模型过滤
- 使用Qwen2.5-32B-Instruct模型生成语言质量标注(1-5分)
- 标注样本量:中文398K,英文250K
- 基于标注训练XLM-RoBERT-large分类器(回归任务)
- 移除分类器评分为1或2的文档
数据集统计
- 样本数量:47,087,808
- 磁盘大小:214GB(parquet格式)




