Automated Lead Optimization of MMP-12 Inhibitors Using a Genetic Algorithm
收藏资源简介:
Traditional lead optimization projects involve long synthesis and testing cycles, favoring extensive structure−activity relationship (SAR) analysis and molecular design steps, in an attempt to limit the number of cycles that a project must run to optimize a development candidate. Microfluidic-based chemistry and biology platforms, with cycle times of minutes rather than weeks, lend themselves to unattended autonomous operation. The bottleneck in the lead optimization process is therefore shifted from synthesis or test to SAR analysis and design. As such, the way is open to an algorithm-directed process, without the need for detailed user data analysis. Here, we present results of two synthesis and screening experiments, undertaken using traditional methodology, to validate a genetic algorithm optimization process for future application to a microfluidic system. The algorithm has several novel features that are important for the intended application. For example, it is robust to missing data and can suggest compounds for retest to ensure reliability of optimization. The algorithm is first validated on a retrospective analysis of an in-house library embedded in a larger virtual array of presumed inactive compounds. In a second, prospective experiment with MMP-12 as the target protein, 140 compounds are submitted for synthesis over 10 cycles of optimization. Comparison is made to the results from the full combinatorial library that was synthesized manually and tested independently. The results show that compounds selected by the algorithm are heavily biased toward the more active regions of the library, while the algorithm is robust to both missing data (compounds where synthesis failed) and inactive compounds. This publication places the full combinatorial library and biological data into the public domain with the intention of advancing research into algorithm-directed lead optimization methods.
传统先导化合物优化项目通常伴随着漫长的合成与测试周期,其核心流程涵盖大规模构效关系(Structure-Activity Relationship, SAR)分析与分子设计步骤,旨在缩减项目优化开发候选物所需的循环次数。基于微流控技术的化学与生物学平台单循环耗时仅为数分钟而非数周,可支持无人值守的自主运行,因此先导化合物优化流程的瓶颈已从合成与测试环节,转向构效关系分析与分子设计环节。由此,无需用户开展细致数据分析的算法驱动型优化流程便具备了实施条件。本文报道了两项采用传统方法开展的合成与筛选实验结果,旨在验证一款遗传算法优化流程,以供未来应用于微流控系统。该算法具备多项适配目标应用场景的创新特性:例如,其对缺失数据具备良好鲁棒性,还可推荐待复测化合物以保障优化过程的可靠性。首先,研究团队通过对嵌入由假定非活性化合物构成的大型虚拟库的内部自有库开展回顾性分析,完成了该算法的初始验证。在第二项前瞻性实验中,以基质金属蛋白酶-12(MMP-12)为靶蛋白,研究团队在10轮优化循环中共提交140种化合物用于合成,并将该结果与手工合成并独立测试的全组合库实验结果进行对比。结果表明,算法筛选得到的化合物显著偏向于库中活性更高的区域;同时,该算法对合成失败导致的缺失数据以及非活性化合物均具备鲁棒性。本文将全组合库及其生物学数据公开至公共领域,以期推动算法驱动的先导化合物优化方法相关研究的发展。



