遇见数据集

CoronaCentral

收藏
Zenodo2023-09-13 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This describes the output file for the CoronaCentral data. The scripts used to create it are hosted in the corona-ml Github repo. The sources for the documents before processing for CoronaCentral are PubMed and CORD-19. The file is a gzipped JSON document containing one record per document. Each document has at least one of: a PubMed ID, a CORD-19 ID (cord_uid), a DOI or a URL. The fields that documents should have are: pubmed_id: PubMed identifier (optional) pmcid: PubMed Central identifier (optional) doi: Digital object identifier (optional) cord_uid: CORD-19 identifier (optional) url: URL journal: Journal/preprint server publish_year: Year of publication (optional) publish_month: Month of publication (optional) publish_day: Day of publication (optional) title: Title of article abstract: Abstract of article (optional) is_preprint: Whether the article is a preprint categories: Predicted categories for article entities: Extracted entities (e.g. drugs) with identifiers and locations within text Please report issues to the corona-ml Github issues page.

本说明用于介绍CoronaCentral数据集(CoronaCentral)的输出文件。用于生成该文件的脚本托管于corona-ml的GitHub代码仓库中。该数据集的原始文档数据源为PubMed数据库(PubMed)与CORD-19数据集(CORD-19)。该文件为经gzip压缩的JSON格式文档,每份文档对应一条记录。每份文档至少包含以下标识之一:PubMed编号、CORD-19编号(cord_uid)、数字对象标识符(Digital Object Identifier,DOI)或统一资源定位符(Uniform Resource Locator,URL)。 文档应包含的字段如下: - pubmed_id: PubMed标识符(可选) - pmcid: PubMed Central标识符(可选) - doi: 数字对象标识符(Digital Object Identifier,可选) - cord_uid: CORD-19标识符(可选) - url: 统一资源定位符(Uniform Resource Locator,可选) - journal: 期刊/预印本服务器 - publish_year: 发表年份(可选) - publish_month: 发表月份(可选) - publish_day: 发表日期(可选) - title: 文章标题 - abstract: 文章摘要(可选) - is_preprint: 标识文章是否为预印本 - categories: 文章预测分类标签 - entities: 提取得到的实体(例如药物),包含其标识符与在文本中的位置信息 若发现任何问题,请前往corona-ml的GitHub问题反馈页面提交。

提供机构:
Zenodo
创建时间:
2021-03-14
二维码
社区交流群
二维码
科研交流群
商业服务