IndustryCorpus_technology
收藏资源简介:
该数据集是为了解决行业模型训练中数据量不足、质量低和缺乏领域专业知识的问题而构建的。通过应用22个行业数据处理操作符,从超过100TB的开放源数据集中筛选出3.4TB的高质量多行业分类的中英文预训练数据集,包括1TB的中文数据和2.4TB的英文数据。数据集进行了12种类型的标签标注,并经过了行业分类语言模型的过滤和文档级别的去重处理。数据集涵盖了18个行业类别,并针对每个行业类别提供了数据大小。为了验证数据集的性能,还进行了持续预训练、SFT和DPO训练,结果显示性能有显著提升。
This dataset is developed to address three core challenges in industry-specific model training: insufficient data volume, low data quality, and the absence of domain-specific professional knowledge. Leveraging 22 industry-oriented data processing operators, we screened a 3.4 TB high-quality multi-industry categorized Chinese-English pre-training dataset from over 100 TB of open-source datasets, which comprises 1 TB of Chinese data and 2.4 TB of English data. The dataset has been annotated with 12 types of labels, filtered by industry-classification language models, and processed with document-level deduplication. It covers 18 industry categories, with detailed data size specifications provided for each category. To validate the performance of this dataset, we conducted continuous pre-training, supervised fine-tuning (SFT), and direct preference optimization (DPO) training, and the results demonstrated significant performance improvements.
数据集概述
数据集描述
- 语言: 中文和英文
- 数据量: 1TB中文数据,2.4TB英文数据
- 任务类别: 文本生成
- 行业分类: 18个行业类别,包括医疗、教育、文学、金融、旅行、法律、体育、汽车、新闻等
数据处理
- 数据来源: 从超过100TB的开放源数据集中筛选,包括WuDaoCorpora, BAAI-CCI, redpajama, SkyPile-150B
- 数据处理操作: 应用22个行业数据处理操作符进行清洗和过滤
- 规则过滤: 传统中文转换、电子邮件移除、IP地址移除、链接移除、Unicode修复等
- 模型过滤: 使用行业分类语言模型,准确率80%
- 数据去重: 使用MinHash文档级去重
数据标注
- 中文数据标签: 包括字母数字比、平均行长度、语言置信度分数、最大行长度、困惑度、毒性字符比等12种标签
数据集验证
- 模型训练: 进行了持续预训练、SFT和DPO训练
- 性能提升: 客观性能提升20%,主观胜率82%
行业分类数据量
| 行业类别 | 数据量 (GB) | 行业类别 | 数据量 (GB) |
|---|---|---|---|
| 编程 | 4.1 | 政治 | 326.4 |
| 法律 | 274.6 | 数学 | 5.9 |
| 教育 | 458.1 | 体育 | 442 |
| 金融 | 197.8 | 文学 | 179.3 |
| 计算机科学 | 46.9 | 新闻 | 564.1 |
| 技术 | 333.6 | 电影与电视 | 162.1 |
| 旅行 | 82.5 | 医学 | 189.4 |
| 农业 | 41.6 | 汽车 | 40.8 |
| 情感 | 31.7 | 人工智能 | 5.6 |
| 总计 (GB) | 3386.5 |




