Chinese News Framing dataset
收藏资源简介:
中文新闻报道框架数据集(Chinese News Framing dataset)是由谢菲尔德大学计算机科学学院创建的,该数据集是首个专注于中文新闻框架检测的自动检测数据集。它包含了从13个不同国家的网站收集的约30万篇中文新闻文章,经过精心挑选和标注,用于分析不同类型中文媒体话语中的框架模式。数据集涵盖了从2020年至2024年底发布的新闻,包括全球讨论的各种事件,如COVID-19疫苗、巴以冲突、俄乌战争和美国大选等。数据集通过BERTopic进行主题标注,然后根据主题与新闻框架类别的相似度进行分层抽样和标注,以用于新闻框架检测的多标签多类别分类任务。
The Chinese News Framing dataset was created by the School of Computer Science at the University of Sheffield, and it is the first automated dataset dedicated to Chinese news framing detection. It comprises approximately 300,000 Chinese news articles collected from websites across 13 distinct countries, which have been carefully curated and annotated to analyze framing patterns in the discourse of various Chinese media outlets. The dataset covers news published from 2020 to the end of 2024, encompassing a wide range of globally discussed events including COVID-19 vaccines, the Israel-Hamas conflict, the Russia-Ukraine war, and the U.S. presidential election. For the annotation process, BERTopic is first used to perform topic labeling, followed by stratified sampling and annotation based on the similarity between topics and news framing categories, to support multi-label and multi-class classification tasks for news framing detection.




