Dataset with Curated Coronavirus-related R&D Outputs (1970 - March 2020)
收藏资源简介:
<strong>The zip file includes metadata about R&D outputs related to all </strong><strong>Coronaviruses</strong>. The two main data sources used for this work include: <strong>Microsoft Academic Graph (MSA) API</strong>: Subset of 10.000+ coronavirus R&D outputs, with data from 1970, including patents and scientific publications. <strong>The Global Research Identifier Database (GRID)</strong><strong>:</strong> Used to enrich the organization data extracted from MSA. After cleaning and enriching the data we extracted a total of <strong>1.100+ organization and</strong> <strong>26.700+ researchers</strong> spread across <strong>90+ countries</strong> and <strong>700+ cities</strong>. Inside the zip file you will find: documents.csv: Full list of documents topics.csv: list of topics connected to documents terms.csv: list of terms connected to documents people.csv: list of people (authors) connected to documents orgsplus.csv: list of organisations (incl their locations) connected to documents More information is available at https://app.gitbook.com/@dataverz/s/coronavirus-r-and-d/
本压缩包包含与所有冠状病毒(Coronaviruses)相关的研发产出元数据。本研究采用的两大核心数据源如下:其一为微软学术图谱(Microsoft Academic Graph, MSA)API:涵盖1970年以来的10000余篇冠状病毒研发产出数据,包含专利与学术出版物;其二为全球研究标识符数据库(Global Research Identifier Database, GRID):用于丰富从MSA中提取的机构数据。经数据清洗与富集处理后,本次研究共提取得到分布于90余个国家、700余座城市的1100余家机构与26700余名研究人员。压缩包内包含以下文件:documents.csv:文档完整列表;topics.csv:与文档关联的主题列表;terms.csv:与文档关联的术语列表;people.csv:与文档或作者关联的人员列表;orgsplus.csv:与文档关联的机构(含其所在地)列表。更多详细信息可访问:https://app.gitbook.com/@dataverz/s/coronavirus-r-and-d/



