Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data
收藏资源简介:
The applicability domain of machine learning models trained on structural fingerprints for the prediction of biological endpoints is often limited by the lack of diversity of chemical space of the training data. In this work, we developed “similarity-based merger models” which combined the output of individual models trained on cell morphology (based on Cell Painting) and chemical structure (based on chemical fingerprints) and the structural and morphological similarities of the test compounds to training compounds. We applied these similarity-based merger models using logistic equations to weigh individual features and predicted assay hit calls of 177 assays from ChEMBL, PubChem and the Broad Institute, where the required Cell Painting annotations were available. We found that the similarity-based merger models outperformed other models with an additional 20% assays (79 out of 177 assays) with an AUC>0.70 compared with 65 out of 177 assays using structural models and 50 out of 177 assays using Cell Painting models. Our results demonstrate that similarity-based merger models combining structure and cell morphology models can more accurately predict a wide range of biological assay outcomes and expand the applicability domain by better extrapolating to new structural and morphology spaces.
基于结构指纹(structural fingerprints)训练、用于预测生物学终点(biological endpoints)的机器学习模型,其适用域往往受限于训练数据化学空间的多样性不足。本研究中,我们构建了基于相似性的融合模型(similarity-based merger models),该模型将基于细胞形态(依托细胞绘画(Cell Painting)技术)训练的单个模型、基于化学结构(依托化学指纹)训练的单个模型的输出,与测试化合物相较于训练化合物的结构相似性及形态相似性进行融合。我们采用逻辑回归方程对各特征进行权重赋值,将该基于相似性的融合模型应用于来自ChEMBL、PubChem及布罗德研究所(Broad Institute)的177个生物实验的命中判定结果预测任务,且这些数据集均带有可用的细胞绘画注释信息。结果显示,相较于仅使用结构模型的65/177个实验、仅使用细胞绘画模型的50/177个实验,基于相似性的融合模型可在额外20%的实验中实现性能提升,共177个实验中的79个实现了AUC>0.70的优异表现。本研究结果证实,融合结构模型与细胞形态模型的基于相似性的融合模型,能够更精准地预测多种生物学实验结果,并通过更好地外推至全新的结构空间与形态空间,拓展了模型的适用域。



