Exploiting Pretrained Biochemical Language Models for Targeted Drug Design
收藏资源简介:
This repository contains materials for the paper,<em> Exploiting Pretrained Biochemical Language Models for Targeted Drug Design, </em>which<em> </em>has been accepted for publication in <em>Bioinformatics</em> Published by Oxford University Press. <em>data.zip</em> contains vocabulary files for the pretrained models, additional information regarding proteins (PFAM family, protein similarity) and interactions filtered from BindingDB which are further split into train, validation and test sets and used to train target specific molecule generation models. <em>models.zip </em>includes files for the models trained in this study. <em>predictions.zip </em>comprises the compounds generated with the targeted models and the result of their evaluation with respect to benchmarking metrics. <em>docking.zip </em>contains <em>targets/ </em>including PDB files of the test proteins selected for docking evaluation, <em>ligands/ </em>including SDF files for molecules generated with the targeted models and two decoding strategies (i.e. beam search and sampling) and <em>complex/ </em>including docking outputs.



