遇见数据集

2020年7月至2023年10月人民日报官网新闻数据集

收藏
官方服务:

资源简介:

数据内容:数据集包含人民日报2020年7月至2023年10月之间的文本数据,总计文本79719件。用以支持新闻知识图谱的构建。不涉及个人隐私、社会机构及公共利益等敏感数据。 采集方案:通过收集公开数据、自动清洗提取的方式进行数据采集。 时间及地点:采集于2022年-2023年,媒体融合生产技术与系统国家实验室智能化视频生产系统研究部,杭州。

Data Content: This dataset contains textual data from People's Daily spanning from July 2020 to October 2023, totaling 79,719 text entries. It is intended to support the construction of news knowledge graphs, and no sensitive data involving personal privacy, social institutions or public interests is included herein. Collection Approach: Data was collected by gathering public data and performing automated cleaning and extraction. Time and Location: The data was collected between 2022 and 2023 at the Research Department of Intelligent Video Production System, National Laboratory of Media Convergence Production Technology and Systems, Hangzhou.

搜集汇总
数据集介绍
2020年7月至2023年10月人民日报官网新闻数据集 数据集图片
背景与挑战
背景概述
该数据集收录了2020年7月至2023年10月期间人民日报官网的新闻文本,共计79719件,旨在支持新闻知识图谱的构建。数据通过公开采集和自动清洗方式获取,由新华智云科技有限公司在2022年至2023年于杭州整理完成。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务