Ontology based text mining of gene-phenotype associations: application to candidate gene prediction
收藏资源简介:
Gene-phenotype associations play an important role in understanding<br> the disease mechanisms which is a requirement for treatment<br> development. A portion of gene-phenotype associations are observed<br> mainly experimentally and made publicly available through several<br> standard resources such as MGI. However, there is still a vast<br> amount of gene--phenotype associations buried in the biomedical<br> literature. Given the large amount of literature data, we need<br> automated text mining tools to alleviate the burden in manual<br> curation of gene-phenotype associations and to develop<br> comprehensive resources. We developed an ontology based<br> approach in combination with statistical methods to text mine<br> gene-phenotype associations from literature. Our method achieved<br> AUC values of 0.90 and 0.75 in recovering known gene-phenotype<br> associations from HPO and MGI respectively. We posit that candidate<br> genes and their relevant diseases should be expressed with similar<br> phenotypes in publications. Thus, we demonstrate the utility of our<br> approach by predicting disease candidate genes based on the semantic<br> similarities of phenotypes associated with genes and diseases. We evaluated our disease candidate prediction model on<br> the gene-disease associations from MGI. Our model achieved AUC<br> values of 0.90 and 0.87 on OMIM (human) and MGI (mouse) datasets of<br> gene-disease associations respectively. Our manual analysis on the<br> text mined data revealed that, our method can accurately extract<br> gene-phenotype associations which are not currently covered by the<br> existing public gene-phenotype resources. Overall, results indicate<br> that our method can precisely extract known as well as new<br> gene-phenotype associations from literature. This released dataset at Zenodo covers our gene-phenotype extracts from the literature. All the methods used to extract the data are available at https://github.com/bio-ontology-research-group/genepheno.
基因-表型关联(gene-phenotype associations)在解析疾病机制中发挥着关键作用,而阐明疾病机制是开发治疗手段的必要前提。目前已有部分基因-表型关联通过实验手段得以验证,并通过MGI等多项标准数据库公开发布。然而,仍有大量基因-表型关联潜藏在生物医学文献之中。鉴于现有文献数据体量庞大,亟需借助自动化文本挖掘工具,以减轻基因-表型关联人工编目的工作负担,并助力构建全面的相关资源库。 本研究开发了一种结合统计方法的基于本体的分析方法,用于从文献中开展基因-表型关联的文本挖掘。该方法在从人类表型本体(Human Phenotype Ontology,HPO)和MGI数据集召回已知基因-表型关联的任务中,分别取得了0.90和0.75的AUC(受试者工作特征曲线下面积)值。我们提出假设:在已发表文献中,候选基因与其关联疾病所对应的表型特征应当具有相似性。据此,我们基于基因与疾病所关联表型的语义相似度实现疾病候选基因预测,以此验证本方法的实用性。 我们基于MGI的基因-疾病关联数据集,对疾病候选基因预测模型进行了评估。该模型在基因-疾病关联的在线人类孟德尔遗传数据库(Online Mendelian Inheritance in Man,OMIM,人类)数据集与MGI(小鼠)数据集上,分别取得了0.90和0.87的AUC值。我们对文本挖掘得到的数据开展人工分析后发现,本方法能够精准提取当前已有公共基因-表型资源未收录的基因-表型关联。 综上,实验结果表明,本方法能够从文献中精准提取已知及全新的基因-表型关联。本研究已将从文献中提取的基因-表型关联数据集上传至Zenodo平台并公开。本研究用于提取数据的全部方法已开源至https://github.com/bio-ontology-research-group/genepheno。



