An annotated dataset for gene-melanoma relation extraction from scientific literature
收藏资源简介:
Melanoma is the least common but the deadliest of skin cancers. This cancer begins when the genes of a cell suffer damage or fail, and identifying the genes involved in melanoma is crucial for understanding the melanoma tumorigenesis. To date, machine learning for gene-melanoma relation extraction from text has been limited by the lack of annotated resources. To overcome this problem, we have exploited the information of the Melanoma Gene Database (a manually curated database of human melanoma related genes) to build an annotated dataset of binary relations between genes and melanoma entities mentioned in PubMed abstracts. The exploitability of the dataset was tested with both traditional machine learning, and neural network-based models. These models are then used to automatically extract gene-melanoma relations from the biomedical literature. Researchers can use the annotated dataset to develop and compare their own models. Moreover, the relations extracted from the literature can be integrated with existing structured knowledge to facilitate researchers in their data search.
黑色素瘤(Melanoma)是皮肤癌中最为少见但致死性最高的癌种。该癌症起源于细胞基因受损或功能失常,明确与黑色素瘤相关的基因对于解析黑色素瘤的肿瘤发生机制至关重要。迄今为止,基于文本的基因-黑色素瘤关系提取相关机器学习研究,受限于标注资源的匮乏。为解决这一问题,我们借助黑色素瘤基因数据库(Melanoma Gene Database,该数据库为经人工整理标注的人类黑色素瘤相关基因数据库)的相关信息,构建了一套标注数据集,该数据集涵盖PubMed摘要中提及的基因与黑色素瘤实体间的二元关系。我们分别采用传统机器学习模型与基于神经网络的模型,对该数据集的适配性能进行了验证。借助这些训练完成的模型,可从生物医学文献中自动提取基因-黑色素瘤关系。研究人员可使用该标注数据集开发并对比自有模型。此外,从文献中提取的关系可与现有结构化知识进行整合,以助力研究人员开展数据检索工作。



