遇见数据集

Supporting data for "CoVEffect: Interactive System for Mining the Effects of SARS-CoV-2 Mutations and Variants Based on Deep Learning"

收藏
Zenodo2023-04-20 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

This repository contains the datasets created and extracted for the paper: Giuseppe Serna García, Ruba Al Khalaf, Francesco Invernici, Stefano Ceri, and Anna Bernasconi. 2022.<br> "<strong>CoVEffect</strong>: Interactive System for Mining the <strong>Effects of SARS-CoV-2 Mutations and Variants</strong> Based on Deep Learning". (Available online at http://gmql.eu/coveffect) --------------------------------------------------------------------------------<br> LIST OF FILES WITH DESCRIPTION:<br> -------------------------------------------------------------------------------- AdditionalFile1-effects-taxonomy:<br> Descriptions of legal values for the 'Effect' field, based on a categorized taxonomy. AdditionalFile2-levels-taxonomy:<br> Descriptions of legal values for the 'Level' field. AdditionalFile3-training_dataset_target:<br> List of target tuples (manually annotated) of 221 abstracts considered for training the model. For each abstract, target tuples follow the schema ID, DOI, title, entity, effect, level, type (mutation or variant), tuples_count (&gt;1 when an effect/level is shared by multiple entities, #abstracts containing the same effect described in the tuple). AdditionalFile4-validation_dataset_target:<br> List of target tuples (manually annotated) of 50 abstracts considered for validating the prepared prediction model.<br> For each abstract, target tuples follow the schema defined for AdditionalFile3. AdditionalFile5-validation_dataset_highlighted:<br> Textual abstracts of the 50 manuscripts considered for validation; the text used to support the manual target annotations has been highlighted in yellow. AdditionalFile6-validation_dataset_prediction:<br> List of predicted annotations of 50 abstracts considered for validating the prepared prediction model. The file is split in 4 TSV, respectively for entity (a), effect (b), level (c), and whole tuple predictions (d). AdditionalFile7-keywords_query_list:<br> Keyword-based search run on the CORD-19 dataset to extract a relevant subset of abstracts regarding the scope of interest of CoVEffect. The Boolean logic used to combine keywords is explained in the section 'Annotations of the biology-related CORD-19 cluster'. AdditionalFile8-CORD-19_batch_dataset_metadata:<br> Metadata of the 7,230 papers extracted by the keyword-based query in AdditionalFile7.<br> These abstracts have been annotated by the prediction framework. AdditionalFile9-CORD-19_batch_dataset_prediction:<br> List of predicted annotations of 7,230 abstracts extracted from the biology-related cluster of CORD-19. AdditionalFile10-test_dataset_target:<br> List of target tuples (manually annotated) of 100 abstracts randomly selected from the 7,230 extracted as in AdditionalFile8.<br> For each abstract, target tuples follow the schema defined for AdditionalFile3. AdditionalFile11-test_dataset_prediction:<br> List of predicted annotations of 100 abstracts considered for testing the prediction model on a subset of the CORD-19 biology-related cluster. As AdditionalFile6, it is split in 4 TSV, respectively for entity (a), effect (b), level (c), and whole tuple predictions (d).

本仓库包含为以下论文创建并提取的配套数据集:作者为Giuseppe Serna García、Ruba Al Khalaf、Francesco Invernici、Stefano Ceri与Anna Bernasconi,发表于2022年。 **CoVEffect**:基于深度学习的严重急性呼吸综合征冠状病毒2(SARS-CoV-2)突变与变异株影响挖掘交互式系统(可在线访问:http://gmql.eu/coveffect) -------------------------------------------------------------------------------- 文件列表及说明: -------------------------------------------------------------------------------- AdditionalFile1-effects-taxonomy: “效应(Effect)”字段合法取值的说明,基于分类学体系构建。 AdditionalFile2-levels-taxonomy: “层级(Level)”字段合法取值的说明。 AdditionalFile3-training_dataset_target: 用于模型训练的221篇经人工标注的摘要目标元组列表。每篇摘要对应的目标元组遵循以下结构:ID、数字对象标识符(DOI, Digital Object Identifier)、标题、实体、效应、层级、类型(突变或变异株)、元组计数(当多个实体共享同一效应/层级时该值大于1,以及包含该元组所描述效应的摘要总数)。 AdditionalFile4-validation_dataset_target: 用于验证所构建预测模型的50篇经人工标注的摘要目标元组列表。每篇摘要对应的目标元组遵循AdditionalFile3中定义的结构。 AdditionalFile5-validation_dataset_highlighted: 用于验证的50篇论文的文本摘要;用于支撑人工目标标注的文本已以黄色高亮标记。 AdditionalFile6-validation_dataset_prediction: 用于验证所构建预测模型的50篇摘要的预测标注列表。该文件分为4个制表符分隔值(TSV, Tab-Separated Values)文件,分别对应实体(a)、效应(b)、层级(c)与完整元组预测(d)。 AdditionalFile7-keywords_query_list: 基于关键词在CORD-19数据集上运行的检索式,用于提取与CoVEffect研究范围相关的摘要子集。组合关键词的布尔逻辑已在“生物学相关CORD-19集群的标注”章节中说明。 AdditionalFile8-CORD-19_batch_dataset_metadata: 通过AdditionalFile7中的关键词检索式提取的7230篇论文的元数据。这些摘要已通过预测框架完成标注。 AdditionalFile9-CORD-19_batch_dataset_prediction: 从CORD-19生物学相关集群中提取的7230篇摘要的预测标注列表。 AdditionalFile10-test_dataset_target: 从AdditionalFile8提取的7230篇摘要中随机选取的100篇经人工标注的摘要目标元组列表。每篇摘要对应的目标元组遵循AdditionalFile3中定义的结构。 AdditionalFile11-test_dataset_prediction: 用于在CORD-19生物学相关集群子集上测试预测模型的100篇摘要的预测标注列表。与AdditionalFile6一致,该文件分为4个制表符分隔值(TSV)文件,分别对应实体(a)、效应(b)、层级(c)与完整元组预测(d)。

提供机构:
Zenodo
创建时间:
2023-04-11
二维码
社区交流群
二维码
科研交流群
商业服务