Quote Following
收藏资源简介:
This data serves as the companion dataset for the forthcoming paper "Exposing the Obscured Influence of State-Controlled Media: A Causal Estimation of Influence Between Media Outlets Via Quotation Propagation." This partially-sanitized dataset is a compromise between preserving valuable intellectual property and allowing replication. The original dataset has 618,328 quotes from 123,396 articles published between May 2018 and October 2019. It includes articles from 454 outlets labeled with one of 167 topics and 418 sentiments. Each quote has additional information about the date of the article, the country of media origin, any geographic references in the quotation, the quote speaker and the outlet where the quote and article appeared. Each quote concerns geopolitical news. The quotes are drawn from articles published by the most prominent outlets in 24 European countries. The published dataset contains the necessary data to replicate the analysis. We present the data after having conducted quote matching on the original, ~600k quote dataset. Each row represents an instance of "quote following," where one outlet, the "Following Outlet," used a quote in an article after the "Source Outlet" used the quote in an article. The source article using the quote was published on "Leading Article Date" and the following outlet article using the quote was published on "Leading Article Date". The article was hand-labeled with a given "Topic" and "Quote Sentiment". "N Sources Using Quote" gives the number of outlets that published an article using the quote, while "N Followings" gives the number of times the following outlet used the quote. Finally, "Leading English Quote" gives the quote used by the source outlet, while "Following English Quote" gives the quote used by the following outlet. In the case where multiple variations of a quote were used by either the source or following outlet, the quote given is the first version of the quote they used. Topics and sentiments have been sanitized to preserve intellectual property. Topics are labeled as either "Nuclear Cooperation" or "All Topics". This allows for replication on either all topics or nuclear cooperation, as done in the companion paper. Similarly, all sentiments except those taking a stance either for or against Russia or the United States are labeled "All Sentiments". English quotes for Russian-language media have been translated using Google translate.
本数据集为即将发表的论文《揭露官方媒体的隐蔽影响力:基于引语传播(Quotation Propagation)的媒体间影响力因果估计》(Exposing the Obscured Influence of State-Controlled Media: A Causal Estimation of Influence Between Media Outlets Via Quotation Propagation)的配套数据集(companion dataset)。本部分脱敏数据集(partially-sanitized dataset)在保留有价值的知识产权与支持研究复现之间取得了平衡。 原始数据集包含2018年5月至2019年10月期间发表的123396篇文章中的618328条引语(quote)。数据集涵盖来自454家媒体机构(media outlets)的文章,这些机构被标注为167个主题(Topic)与418种情感倾向(sentiment)。每条引语附带文章发布日期、媒体来源国、引语中的地理参考信息、引语发言者以及该引语与文章所属媒体机构的相关信息。所有引语均与地缘政治新闻相关,数据取自24个欧洲国家头部媒体机构发表的文章。 公开可用的数据集包含复现分析所需的全部数据。我们对原始约60万条引语的数据集完成了引语匹配(quote matching)后发布了本数据集。数据集中每一行代表一次引语追随(quote following)实例:即某一追随媒体机构(following outlet)在其文章中使用了某条引语,而该引语此前已被来源媒体机构(source outlet)在其文章中使用过。使用该引语的来源文章发布于"Leading Article Date",追随媒体机构使用该引语的文章则发布于"Leading Article Date"。文章被人工标注了指定的主题(Topic)与引语情感倾向(Quote Sentiment)。使用该引语的来源媒体数量(N Sources Using Quote)表示发布过包含该引语的文章的媒体机构数量,而追随次数(N Followings)表示追随媒体机构使用该引语的总次数。最后,来源引语原文(Leading English Quote)为来源媒体机构使用的引语文本,追随引语原文(Following English Quote)为追随媒体机构使用的引语文本。若来源或追随媒体机构使用了该引语的多个变体,则此处给出的是其首次使用的版本。 为保护知识产权,主题与情感倾向已完成脱敏处理。主题仅标注为核合作(Nuclear Cooperation)或全主题(All Topics),这使得研究人员可以如配套论文中所示,选择针对全主题或核合作主题开展复现研究。类似地,除明确支持或反对俄罗斯(Russia)与美国(United States)的情感倾向外,其余所有情感倾向均被标注为全情感倾向(All Sentiments)。 俄语媒体的英文引语均通过谷歌翻译(Google Translate)完成翻译。



