A meta analysis of Wikipedia's coronavirus sources during the COVID-19 pandemic
收藏资源简介:
At the height of the coronavirus pandemic, on the last day of March 2020, Wikipedia in all languages broke a record for most traffic in a single day. Since the breakout of the Covid-19 pandemic at the start of January, tens if not hundreds of millions of people have come to Wikipedia to read - and in some cases also contribute - knowledge, information and data about the virus to an ever-growing pool of articles. Our study focuses on the scientific backbone behind the content people across the world read: which sources informed Wikipedia’s coronavirus content, and how was the scientific research on this field represented on Wikipedia. Using citation as readout we try to map how COVID-19 related research was used in Wikipedia and analyse what happened to it before and during the pandemic. Understanding how scientific and medical information was integrated into Wikipedia, and what were the different sources that informed the Covid-19 content, is key to understanding the digital knowledge echosphere during the pandemic. To delimitate the corpus of Wikipedia articles containing Digital Object Identifier (DOI), we applied two different strategies. First we scraped every Wikipedia pages form the COVID-19 Wikipedia project (about 3000 pages) and we filtered them to keep only page containing DOI citations. For our second strategy, we made a search with EuroPMC on Covid-19, SARS-CoV2, SARS-nCoV19 (30’000 sci papers, reviews and preprints) and a selection on scientific papers form 2019 onwards that we compared to the Wikipedia extracted citations from the english Wikipedia dump of <strong>May 2020</strong> (2’000’000 DOIs). This search led to 231 Wikipedia articles containing at least one citation of the EuroPMC search or part of the wikipedia COVID-19 project pages containing DOIs. Next, from our 231 Wikipedia articles corpus we extracted DOIs, PMIDs, ISBNs, websites and URLs using a set of regular expressions. Subsequently, we computed several statistics for each wikipedia article and we retrive Atmetics, CrossRef and EuroPMC infromations for each DOI. Finally, our method allowed to produce tables of citations annotated and extracted infromations in each wikipadia articles such as books, websites, newspapers. Files used as input and extracted information on Wikipedia's COVID-19 sources are presented in this archive. See the WikiCitationHistoRy Github repository for the R codes, and other bash/python scripts utilities related to this project.
在新冠疫情(coronavirus pandemic)最严峻的阶段,2020年3月的最后一天,全球多语言维基百科(Wikipedia)创下了单日访问量的新纪录。自2020年初新冠疫情(Covid-19 pandemic)暴发以来,即便未达数亿,也有数千万用户登录维基百科,阅读乃至参与贡献与该病毒相关的知识、信息与数据,共同充实着日益增长的百科词条库。 本研究聚焦于全球用户所阅读的维基百科内容背后的科学支撑体系:究竟是哪些数据源为维基百科的新冠相关内容提供了信息支撑,而该领域的科学研究又是如何在维基百科中得以呈现的。我们以引用关系作为观测指标,试图梳理新冠病毒相关研究在维基百科中的使用情况,并分析疫情暴发前后这些研究的变迁路径。理解科学与医学信息如何融入维基百科,以及为新冠相关词条提供信息的各类数据源,是解析疫情期间数字知识回音圈(digital knowledge echosphere)的核心关键。 为了划定包含数字对象标识符(Digital Object Identifier, DOI)的维基百科词条语料库,我们采用了两种不同的策略。第一种策略是抓取新冠疫情维基百科专项项目下的所有页面(共计约3000页),并通过筛选仅保留包含DOI引用的词条。第二种策略则是在欧洲PubMed Central(EuroPMC)中针对"Covid-19""SARS-CoV2""SARS-nCoV19"开展检索,共获取30000篇科技论文、综述与预印本;同时筛选2019年以来发表的科技论文,并将其与2020年5月的英文维基百科快照(共计2000000个DOI)中提取的引用内容进行比对。通过该检索流程,我们最终得到231个维基百科词条:这些词条要么至少包含一条EuroPMC检索结果的引用,要么属于前述新冠维基百科专项项目中包含DOI的页面。 随后,我们从这231个词条组成的语料库中,通过正则表达式集合提取了DOI、PubMed ID(PMID)、国际标准书号(ISBN)、网站与统一资源定位符(URL)。紧接着,我们为每个维基百科词条计算了多项统计指标,并为每个DOI检索了替代计量指标(Altmetrics,原文拼写为Atmetics,疑为笔误)、交叉引用数据库(CrossRef)与欧洲PubMed Central(EuroPMC)的相关信息。最终,我们的方法得以生成各维基百科词条中经标注的引用与提取信息表格,涵盖书籍、网站、报纸等各类来源。 本项目所使用的输入文件与维基百科新冠相关数据源的提取信息已归档于此。相关R代码、以及本项目涉及的Bash/Python脚本工具,可参见WikiCitationHistoRy GitHub仓库。



