Discovering Highly Potent Molecules from an Initial Set of Inactives Using Iterative Screening
收藏资源简介:
The versatility of similarity searching and quantitative structure–activity relationships to model the activity of compound sets within given bioactivity ranges (i.e., interpolation) is well established. However, their relative performance in the common scenario in early stage drug discovery where lots of inactive data but no active data points are available (i.e., extrapolation from the low-activity to the high-activity range) has not been thoroughly examined yet. To this aim, we have designed an iterative virtual screening strategy which was evaluated on 25 diverse bioactivity data sets from ChEMBL. We benchmark the efficiency of random forest (RF), multiple linear regression, ridge regression, similarity searching, and random selection of compounds to identify a highly active molecule in the test set among a large number of low-potency compounds. We use the number of iterations required to find this active molecule to evaluate the performance of each experimental setup. We show that linear and ridge regression often outperform RF and similarity searching, reducing the number of iterations to find an active compound by a factor of 2 or more. Even simple regression methods seem better able to extrapolate to high-bioactivity ranges than RF, which only provides output values in the range covered by the training set. In addition, examination of the scaffold diversity in the data sets used shows that in some cases similarity searching and RF require two times as many iterations as random selection depending on the chemical space covered in the initial training data. Lastly, we show using bioactivity data for COX-1 and COX-2 that our framework can be extended to multitarget drug discovery, where compounds are selected by concomitantly considering their activity against multiple targets. Overall, this study provides an approach for iterative screening where only inactive data are present in early stages of drug discovery in order to discover highly potent compounds and the best experimental set up in which to do so.
相似性搜索与定量构效关系(quantitative structure–activity relationships)在给定生物活性区间内对化合物集活性进行建模(即插值场景)的通用性已得到广泛证实。然而,在药物发现早期阶段常见的场景——即仅存在大量无活性数据却无任何活性数据点的场景(从低活性区间向高活性区间外推)中,二者的相对性能尚未得到彻底研究。为此,我们设计了一种迭代虚拟筛选策略,并基于来自ChEMBL的25个多样化生物活性数据集对该策略进行了评估。我们以在大量低活性化合物中识别测试集内高活性分子为目标,对随机森林(random forest, RF)、多元线性回归、岭回归、相似性搜索以及化合物随机选择的效率开展了基准测试。我们以找到该活性分子所需的迭代次数作为评价指标,对各实验设置的性能进行评估。研究表明,线性回归与岭回归的性能通常优于随机森林与相似性搜索,可将找到活性化合物所需的迭代次数降低2倍及以上。即便是简单的回归方法,也比仅能输出训练集覆盖区间内结果值的随机森林更擅长向高生物活性区间外推。此外,对所用数据集的骨架多样性(scaffold diversity)进行分析后发现,在部分场景下,根据初始训练数据所覆盖的化学空间不同,相似性搜索与随机森林所需的迭代次数可达随机选择的2倍。最后,我们借助环氧合酶-1(COX-1)与环氧合酶-2(COX-2)的生物活性数据证实,我们的框架可拓展至多靶点药物发现场景——即在该场景中,通过同时考量化合物对多个靶点的活性来筛选化合物。综上,本研究为药物发现早期仅存在无活性数据的场景提供了一种迭代筛选方案,可用于发现高活性化合物,并确定实现该目标的最优实验设置。



