IndustryCorpus_automobile
收藏资源简介:
该数据集是为了解决行业模型训练数据集存在的问题而构建的,包括数据量不足、质量低和缺乏领域专业性。通过应用22个行业数据处理操作符,从超过100TB的开放源数据集中筛选出3.4TB的高质量多行业分类中英文预训练数据集。筛选后的数据包括1TB的中文数据和2.4TB的英文数据,并对中文数据进行了12种类型的标签标注。数据集涵盖18个行业类别,包括医疗、教育、文学、金融等,并进行了基于规则和模型的过滤以及文档级别的去重。数据集被分割成18个行业的子数据集,当前描述的是汽车行业的子数据集。
This dataset was developed to address the core issues plaguing industry-specific model training datasets, namely insufficient data volume, subpar data quality, and lack of domain-specific expertise. By utilizing 22 industry-specific data processing operators, we filtered a 3.4 TB high-quality multilingual (Chinese and English) pre-trained dataset with cross-industry classification from an open-source dataset exceeding 100 TB in total size. The filtered dataset comprises 1 TB of Chinese data and 2.4 TB of English data, with the Chinese portion annotated using 12 distinct label categories. Covering 18 industry categories including healthcare, education, literature, finance and more, the dataset has undergone rule-based and model-based filtering as well as document-level deduplication. The full dataset is split into 18 industry-specific sub-datasets, and the sub-dataset described herein corresponds to the automotive industry.
数据集概述
数据集描述
- 语言: 中文和英文
- 数据大小: 1TB中文数据,2.4TB英文数据
- 任务类别: 文本生成
- 行业分类: 18个行业类别,包括医疗、教育、文学、金融、旅游、法律、体育、汽车、新闻等
数据处理
- 数据来源: 从超过100TB的开放源数据集中筛选,包括WuDaoCorpora, BAAI-CCI, redpajama, SkyPile-150B
- 数据处理操作: 应用22个行业数据处理操作符进行清洗和过滤
- 规则过滤: 传统中文转换、电子邮件移除、IP地址移除、链接移除、Unicode修复等
- 模型过滤: 使用行业分类语言模型,准确率80%
- 数据去重: 使用MinHash文档级去重
数据标注
- 中文数据标签: 包括字母数字比、平均行长度、语言置信度分数、最大行长度、困惑度、毒性字符比等12种标签
数据集性能验证
- 模型训练: 进行了持续预训练、SFT和DPO训练
- 性能提升: 目标性能提升20%,主观胜率82%
行业分类数据大小
| 行业类别 | 数据大小 (GB) | 行业类别 | 数据大小 (GB) |
|---|---|---|---|
| 编程 | 4.1 | 政治 | 326.4 |
| 法律 | 274.6 | 数学 | 5.9 |
| 教育 | 458.1 | 体育 | 442 |
| 金融 | 197.8 | 文学 | 179.3 |
| 计算机科学 | 46.9 | 新闻 | 564.1 |
| 技术 | 333.6 | 影视 | 162.1 |
| 旅游 | 82.5 | 医学 | 189.4 |
| 农业 | 41.6 | 汽车 | 40.8 |
| 情感 | 31.7 | 人工智能 | 5.6 |
| 总计 (GB) | 3386.5 |
数据集分割
- 分割方式: 将大数据集分割成18个行业的子数据集,当前为汽车行业子数据集




