Annotation of Allosteric Compounds to Enhance Bioactivity Modeling for Class A GPCRs
收藏资源简介:
Proteins often have both orthosteric and allosteric binding sites. Endogenous ligands, such as hormones and neurotransmitters, bind to the orthosteric site, while synthetic ligands may bind to orthosteric or allosteric sites, which has become a focal point in drug discovery. Usually, such allosteric modulators bind to a protein noncompetitively with its endogenous ligand or substrate. The growing interest in allosteric modulators has resulted in a substantial increase of these entities and their features such as binding data in chemical libraries and databases. Although this data surge fuels research focused on allosteric modulators, binding data is unfortunately not always clearly indicated as being allosteric or orthosteric. Therefore, allosteric binding data is difficult to retrieve from databases that contain a mixture of allosteric and orthosteric compounds. This decreases model performance when statistical methods, such as machine learning models, are applied. In previous work we generated an allosteric data subset of ChEMBL release 14. In the current study an improved text mining approach is used to retrieve the allosteric and orthosteric binding types from the literature in ChEMBL release 22. Moreover, convolutional deep neural networks were constructed to predict the binding types of compounds for class A G protein-coupled receptors (GPCRs). Temporal split validation showed the model predictiveness with Matthews correlation coefficient (MCC) = 0.54, sensitivity allosteric = 0.54, and sensitivity orthosteric = 0.94. Finally, this study shows that the inclusion of accurate binding types increases binding predictions by including them as descriptor (MCC = 0.27 improved to MCC = 0.34; validated for class A GPCRs, trained on all GPCRs). Although the focus of this study is mainly on class A GPCRs, binding types for all protein classes in ChEMBL were obtained and explored. The data set is included as a supplement to this study, allowing the reader to select the compounds and binding types of interest.
蛋白质通常同时具备正构结合位点(orthosteric binding site)与变构结合位点(allosteric binding site)。内源性配体(如激素、神经递质)可结合于正构位点,而合成配体则可结合于正构或变构位点,这一特性已成为药物研发的核心焦点。通常而言,此类变构调节剂会以非竞争性方式与蛋白质的内源性配体或底物相结合。 学界对变构调节剂的研究兴趣与日俱增,使得化学库与数据库中这类分子及其结合数据等相关特征的数量大幅攀升。尽管数据量的激增有力推动了变构调节剂相关研究,但遗憾的是,多数结合数据并未明确标注其属于变构还是正构结合类型。因此,从同时混有变构与正构化合物的数据库中检索变构结合数据极具难度,这会在应用机器学习模型等统计方法时显著降低模型的预测性能。 在既往研究中,我们曾生成ChEMBL 14版的变构数据子集。本研究采用改进后的文本挖掘方法,从ChEMBL 22版的文献数据中检索化合物的变构与正构结合类型。此外,我们构建了卷积深度神经网络,用于预测A类G蛋白偶联受体(G protein-coupled receptors, GPCRs)的化合物结合类型。时序拆分验证结果显示,该模型的预测性能为:马修斯相关系数(Matthews correlation coefficient, MCC)=0.54,变构灵敏度为0.54,正构灵敏度为0.94。 最后,本研究证实,将准确的结合类型作为描述符纳入模型,可有效提升结合预测性能(MCC从0.27提升至0.34;该验证针对A类GPCRs开展,训练集涵盖所有GPCRs)。尽管本研究的核心聚焦于A类GPCRs,但我们仍获取并探索了ChEMBL数据库中所有蛋白质类别的结合类型。本数据集作为本文的补充材料提供,便于读者按需选取感兴趣的化合物与结合类型。




