Data for: Discovering Gene-Disease Associations with Biomedical Word Embeddings
收藏资源简介:
This is the dataset supporting the publication Discovering Gene-Disease Associations with Biomedical Word Embeddings. Finding the right target for a disease is critical in the drug development process. This paper presents a machine learning approach for predicting gene-disease associations that (i) employs biomedical word embeddings as features for a classifier trained on Open Targets Platform (OTP) data that (ii) generalises beyond a specific disease or gene class. We train, evaluate and compare different word embedding models and classifiers for the task at hand. In addition, we validate the approach by training on a past OTP release and show that it can assist in identifying probable positive associations among current low evidence associations, confirmed by a recent OTP release. Furthermore, we train word embedding models on different time slices of biomedical articles from ScienceDirect and demonstrate that the trained classifier predicts associations that have not explicitly been mentioned in the training corpus, 5 years into the future. Please send a message to Elsevier describing briefly your request on how you would like to use the assets with a short justification. Elsevier will connect directly with you for the elaboration of a personalized license. The contact information can be found in the license information.
本数据集支撑了题为《基于生物医学词嵌入的基因-疾病关联挖掘》的学术论文发表。在药物研发流程中,为疾病找到合适的靶点至关重要。本研究提出一种用于预测基因-疾病关联的机器学习方法:其一,将生物医学词嵌入作为模型特征,基于开放靶点平台(Open Targets Platform, OTP)的数据训练分类器;其二,该方法可泛化至特定疾病或基因类别之外的场景。针对当前研究任务,我们训练、评估并对比了多款不同的词嵌入模型与分类器。此外,我们利用过往版本的OTP数据训练该方法并开展验证,结果表明,其可助力识别当前证据级别较低的关联中潜在的阳性关联,且该结果已被最新版本的OTP数据所证实。进一步地,我们基于ScienceDirect数据库中不同时间切片的生物医学文献训练词嵌入模型,并证明所训练的分类器能够预测训练语料中未被明确提及、且滞后5年的关联关系。请向爱思唯尔(Elsevier)发送邮件,简要说明您使用该数据集的用途及相关理由。爱思唯尔将直接与您沟通,共同制定个性化授权协议。联系方式可在授权信息中获取。




