Data for manuscript "The Prevalence of Terms Denoting Far-right and Far-left Political Extremism in U.S. and U.K. News Media"
收藏资源简介:
This data set belongs to an academic manuscript examining longitudinally (2000-2019) the prevalence of terms denoting far-right and far-left political extremism in a large corpus of more than 32 million written news and opinion articles from 54 news media outlets popular in the United States and the United Kingdom. The textual content of news and opinion articles from the 54 outlets listed in the main manuscript is available in the outlet's online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We used derived word frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions. Targeted textual content was located in HTML raw data using outlet specific xpath expressions. Tokens were lowercased prior to estimating frequency counts. To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1 million words of article content from an outlet. This threshold was chosen to maximize inclusion in our analysis of outlets with sparse amounts of articles text per year. Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time. The list of compressed files in this data set is listed next: -analysisScripts.rar contains the analysis scripts used in the main manuscript -articlesContainingTargetWords.rar contains counts of target words in outlets articles as well as total counts of words in articles Usage Notes In a small percentage of articles, outlet specific XPath expressions failed to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles. Most of the incorrect frequency counts were minor deviations from the actual counts such as for instance counting the word "Facebook" in an article footnote encouraging article readers to follow the journalist’s Facebook profile and that the XPath expression mistakenly included as the content of the article main text. Some additional outlet-specific inaccuracies that we could identify occurred in "The Hill" and "Newsmax" news outlets where XPath expressions had some shortfalls at precisely capturing articles’ content. For "The Hill", in years 2007-2009, XPath expressions failed to capture the complete text of the article in about 40% of the articles. This does not necessarily result in incorrect frequency counts for that outlet but in a sample of articles’ words that is about 40% smaller than the total population of articles words for those three years. In the case of "NewsMax", the issue was that for some articles, XPath expressions captured the entire text of the article twice. Notice that this does not result in incorrect frequency counts. If a word appears x times in an article with a total of y words, the same frequency count will still be derived when our scripts count the word 2x times in the version of the article with a total of 2y words. To conclude, in a data analysis of 32 million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 in the main manuscript for illustration of the accuracy of the frequency counts).
本数据集隶属于一篇聚焦2000-2019年纵向研究的学术论文,该研究针对美英两国54家主流新闻媒体的超3200万篇书面新闻与评论文章组成的大型语料库,考察了指代极右翼与极左翼政治极端主义的术语的使用频次。 上述54家媒体的新闻与评论文章文本内容,可通过各媒体官方域名,以及谷歌缓存(Google cache,https://webcache.googleusercontent.com)、互联网档案馆时光机(The Internet Wayback Machine,https://archive.org/web/web.php)、Common Crawl(https://commoncrawl.org)等公共缓存库获取。本研究使用源自上述渠道的衍生词频统计数据。 本分析纳入的文本范围限定为文章标题与正文主体,不包含图注等其他文章元素。目标文本通过适配各媒体的XPath表达式(XPath)从HTML原始数据中提取。在估算词频前,所有词元(Token)均被转换为小写形式。 为避免单年度文本产出较少的媒体扭曲整体词频统计结果,我们仅纳入当年度该媒体至少产出100万词文章的年份数据。设置该阈值的目的是尽可能将每年文章文本量较少的媒体纳入分析。 某媒体在某年度的目标词年度使用频次,通过该年度所有文章中目标词的总出现次数除以该年度所有文章的总词数计算得出。该频次估算方法可消除不同时期文章总产出量差异带来的影响。 本数据集包含的压缩文件清单如下: - analysisScripts.rar:包含主论文中使用的分析脚本 - articlesContainingTargetWords.rar:包含各媒体文章中的目标词统计量及文章总词数统计量 使用须知 在少数文章中,由于各媒体在线平台的HTML元素与CSS样式组合存在异质性,适配媒体的XPath表达式未能正确捕获文章内容。因此,小部分文章的总词数与目标词词数统计存在误差。在对部分文章与媒体进行随机抽样手动估算后,自动生成的词频统计与手动统计的重合度超过90%。 多数错误的词频统计仅与实际值存在小幅偏差,例如XPath表达式误将文章脚注中鼓励读者关注记者"Facebook"账号的"Facebook"一词纳入正文统计。此外,本研究还发现部分媒体存在特定统计误差:以《国会山报》(The Hill)与Newsmax为例,XPath表达式在捕获文章内容时存在缺陷。 针对《国会山报》,2007至2009年间,约40%的文章未能被XPath表达式完整捕获文本。这并不必然导致该媒体的词频统计错误,仅会使得该三年间的抽样文章词数较实际总词数减少约40%。而就NewsMax而言,部分文章的文本被XPath表达式重复捕获了一次。需注意的是,这并不会导致词频统计错误:若某文章中目标词出现x次,总词数为y,那么当脚本将该文章统计为2x次目标词、总词数为2y时,最终计算出的词频仍保持不变。 综上,在涵盖3200万篇文章的数据分析中,我们无法对每一篇文章的词频统计正确性进行人工核查,且由于在线平台中存在少量难以检测的边界情况(如HTML标记语法错误),实现100%的文章内容捕获准确率并不现实。但总体而言,我们确信本研究的词频统计指标能够反映印刷新闻媒体中的术语使用频次,相关准确性说明详见主论文中的图1。



