IndustryCorpus_mathematics
收藏资源简介:
本数据集是一个高质量的多行业分类中英文预训练数据集,通过22个行业数据处理操作符从超过100TB的开放源数据集中筛选出3.4TB的高质量数据,包括1TB的中文数据和2.4TB的英文数据。数据集涵盖18个行业类别,并进行了详细的标注和过滤处理,如传统中文转换、电子邮件和IP地址移除、链接移除、Unicode修复等。此外,数据集还进行了模型训练验证,显示了显著的性能提升。
This dataset is a high-quality multi-industry classified Chinese-English pre-training dataset. 3.4 TB of high-quality data was screened out from over 100 TB of open-source datasets using 22 industry-specific data processing operators, including 1 TB of Chinese data and 2.4 TB of English data. The dataset covers 18 industry categories, and has undergone detailed annotation and filtering processes such as simplified-to-traditional Chinese conversion, removal of email addresses and IP addresses, removal of hyperlinks, Unicode repair, and others. In addition, model training validation was performed on the dataset, which showed significant performance improvements.
数据集概述
数据集基本信息
- 许可证:Apache 2.0
- 语言:中文、英文
- 数据量:1TB 中文数据,2.4TB 英文数据
- 任务类别:文本生成
数据集构建
- 原始数据来源:WuDaoCorpora, BAAI-CCI, redpajama, SkyPile-150B
- 原始数据量:超过 100TB
- 处理后数据量:3.4TB
- 行业分类:18个行业类别
- 数据处理操作:22个行业数据处理操作符
数据处理方法
- 基于规则的过滤:繁体中文转换、电子邮件移除、IP地址移除、链接移除、Unicode修复等
- 基于模型的过滤:行业分类语言模型,准确率80%
- 数据去重:MinHash文档级去重
数据标注
- 中文数据标注:12种标签,包括字母数字比、平均行长度、语言置信度分数、最大行长度、困惑度等
数据集验证
- 模型训练:继续预训练、SFT、DPO训练
- 性能提升:客观性能提升20%,主观胜率82%
行业分类数据量
| 行业类别 | 数据量 (GB) | 行业类别 | 数据量 (GB) |
|---|---|---|---|
| 编程 | 4.1 | 政治 | 326.4 |
| 法律 | 274.6 | 数学 | 5.9 |
| 教育 | 458.1 | 体育 | 442 |
| 金融 | 197.8 | 文学 | 179.3 |
| 计算机科学 | 46.9 | 新闻 | 564.1 |
| 技术 | 333.6 | 影视 | 162.1 |
| 旅游 | 82.5 | 医学 | 189.4 |
| 农业 | 41.6 | 汽车 | 40.8 |
| 情感 | 31.7 | 人工智能 | 5.6 |
| 总计 (GB) | 3386.5 |




