IndustryCorpus_medicine
收藏资源简介:
该数据集是为了解决行业模型训练中数据量不足、质量低和缺乏领域专业知识的问题而构建的。通过应用22个行业数据处理操作符,从超过100TB的开放源数据集中筛选出3.4TB的高质量多行业分类中英文预训练数据集。数据集包括1TB的中文数据和2.4TB的英文数据,并对中文数据进行了12种类型的标签标注。此外,数据集涵盖了18个行业类别,并通过模型和规则基础的过滤方法进行了处理。数据集的大小和行业分类数据大小也在描述中详细列出。
This dataset was constructed to address the challenges of insufficient data volume, low data quality, and lack of domain-specific expertise during industry model training. By applying 22 industry-specific data processing operators, we filtered out a 3.4 TB high-quality multi-industry categorized Chinese-English pre-training dataset from an open-source dataset exceeding 100 TB in size. The dataset comprises 1 TB of Chinese data and 2.4 TB of English data, with the Chinese portion annotated with 12 types of labels. Additionally, the dataset covers 18 industry categories and was processed using model-based and rule-based filtering methodologies. The sizes of the full dataset and the data volumes of each industry category subset are also explicitly detailed in this description.
数据集概述
基本信息
- 许可证: Apache-2.0
- 语言: 中文, 英文
- 数据量: 1TB 中文, 2.4TB 英文
- 任务类别: 文本生成
数据来源与处理
- 原始数据量: 超过 100TB 的开源数据集,包括 WuDaoCorpora, BAAI-CCI, redpajama, SkyPile-150B
- 处理后数据量: 3.4TB 高质量多行业分类中英文预训练数据集
- 数据处理操作: 22 个行业数据处理操作符,用于清洗和过滤数据
- 数据去重: MinHash 文档级去重
- 模型分类: 行业分类语言模型,准确率 80%
数据标签
- 中文数据标签: 包括字母数字比、平均行长度、语言置信度分数、最大行长度、困惑度、毒性字符比等 12 种标签
行业分类数据量
- 行业类别: 18 个类别,包括医疗、教育、文学、金融、旅游、法律、体育、汽车、新闻等
- 具体数据量:
- 编程: 4.1 GB
- 法律: 274.6 GB
- 教育: 458.1 GB
- 金融: 197.8 GB
- 计算机科学: 46.9 GB
- 技术: 333.6 GB
- 旅游: 82.5 GB
- 农业: 41.6 GB
- 情感: 31.7 GB
- 人工智能: 5.6 GB
- 政治: 326.4 GB
- 数学: 5.9 GB
- 体育: 442 GB
- 文学: 179.3 GB
- 新闻: 564.1 GB
- 电影与电视: 162.1 GB
- 医学: 189.4 GB
- 汽车: 40.8 GB
- 总计: 3386.5 GB
数据集应用
- 模型训练: 进行了持续预训练、SFT 和 DPO 训练,验证了数据集的性能,客观性能提升 20%,主观胜率 82%
- 数据集分割: 将大型数据集分割成 18 个行业的子数据集,当前为医疗行业子数据集




