遇见数据集

An annotated dataset for gene-melanoma relation extraction from scientific literature

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

Melanoma is the least common but the deadliest of skin cancers. This cancer begins when the genes of a cell suffer damage or fail, and identifying the genes involved in melanoma is crucial for understanding the melanoma tumorigenesis. To date, machine learning for gene-melanoma relation extraction from text has been limited by the lack of annotated resources. To overcome this problem, we have exploited the information of the Melanoma Gene Database (a manually curated database of human melanoma related genes) to build an annotated dataset of binary relations between genes and melanoma entities mentioned in PubMed abstracts. The exploitability of the dataset was tested with both traditional machine learning, and neural network-based models. These models are then used to automatically extract gene-melanoma relations from the biomedical literature. Researchers can use the annotated dataset to develop and compare their own models. Moreover, the relations extracted from the literature can be integrated with existing structured knowledge to facilitate researchers in their data search.

黑色素瘤(Melanoma)是皮肤癌中发病率最低却致死性最强的癌种。该癌症起源于细胞基因受损或功能失常,明确与黑色素瘤相关的基因对于解析黑色素瘤的肿瘤发生机制至关重要。迄今为止,从文本中提取基因-黑色素瘤关联关系的机器学习研究,一直受限于标注资源的匮乏。为解决这一问题,我们依托黑色素瘤基因数据库(Melanoma Gene Database,一款人工整理的人类黑色素瘤相关基因数据库),从PubMed摘要中提取提及的基因与黑色素瘤实体,构建了二者间的二元关联关系标注数据集。本数据集的可用性通过传统机器学习模型与基于神经网络的模型分别进行了测试验证。上述模型可用于从生物医学文献中自动提取基因-黑色素瘤关联关系。研究人员可借助该标注数据集开发并对比自研模型。此外,从文献中提取的关联关系可与现有结构化知识进行整合,以助力研究人员开展数据检索工作。

创建时间:
2022-08-30
二维码
社区交流群
二维码
科研交流群
商业服务