A corpus designed to study preprints produced during the Covid-19 crisis and to make comparative studies with the pre-pandemic period
收藏资源简介:
This dataset has been created to allow comparative studies of abstracts associated with preprints issued in response to the COVID-19 pandemic (from 01/01/2020 to 12/04/2020) relative to abstracts produced in 2019, the closest pre-pandemic period. The dataset has 2 files: - a txt file with the queries we ran in Dimensions and Lens to create the whole corpus and retrieve metadata - a csv file with the metadata for all preprints in the corpus and the positive, negative and hedge words we extracted with CorTexT Manager tool.
本数据集旨在支撑针对两类摘要开展对比研究:一类为2020年1月1日至2020年4月12日期间为应对COVID-19疫情发布的预印本(preprint)相关摘要,另一类为2019年(距离疫情最近的前疫情时期)产出的同类摘要。 本数据集包含2个文件: - 一个TXT文件,记录了我们在Dimensions与Lens平台中执行的检索策略,用于构建完整语料库并获取元数据; - 一个CSV文件,包含语料库中所有预印本的元数据,以及我们借助CorTexT Manager工具提取的积极词汇、消极词汇与模糊限制语(hedge words)。




