MiChao-HuaFen 1.0
收藏资源简介:
MiChao-HuaFen 1.0是由上海米读科技有限公司和上海人工智能实验室共同创建的预训练语料库数据集,专注于新闻和政府领域。该数据集包含超过7000万条数据,源自2022年的公开互联网数据,经过多轮清洗和处理确保高质量和可靠来源。创建过程中采用了关键词过滤、图像提取、基于规则的过滤和格式转换等方法。该数据集主要用于支持中文垂直领域大型模型的预训练,推动深度学习研究和应用在相关领域的发展,特别适用于AI研究者、学者以及新闻机构和政府部门。
MiChao-HuaFen 1.0 is a pre-trained corpus dataset co-created by Shanghai Midu Technology Co., Ltd. and Shanghai AI Laboratory, focusing on the news and government domains. This dataset contains over 70 million entries, sourced from public internet data in 2022, and has undergone multiple rounds of cleaning and processing to ensure high quality and reliable provenance. Methods including keyword filtering, image extraction, rule-based filtering and format conversion were adopted during its creation. This dataset is primarily used to support the pre-training of large-scale Chinese vertical-domain models, promoting the development of deep learning research and applications in relevant fields, and is particularly suitable for AI researchers, scholars, news agencies and government departments.




