Matched Molecular Pair Analysis: Significance and the Impact of Experimental Uncertainty
收藏资源简介:
Matched molecular pair analysis (MMPA) has become a major tool for analyzing large chemistry data sets for promising chemical transformations. However, the dependence of MMPA predictions on data constraints such as the number of pairs involved, experimental uncertainty, source of the experiments, and variability of the true physical effect has not yet been described. In this contribution the statistical basics for judging MMPA are analyzed. We illustrate the connection between overall MMPA statistics and individual pairs with a detailed comparison of average CHEMBL hERG MMPA results versus pairs with extreme transformation effects. Comparing the CHEMBL results to Novartis data, we find that significant transformation effects agree very well if the experimental uncertainty is considered. This indicates that caution must be exercised for predictions from insignificant MMPAs, yet highlights the robustness of statistically validated MMPA and shows that MMPA on public databases can yield results that are very useful for medicinal chemistry.
匹配分子对分析(Matched molecular pair analysis)现已成为分析大型化学数据集以发掘潜在化学转化的核心工具。然而,目前尚未有研究阐明MMPA预测结果对各类数据约束条件的依赖关系,这些约束包括所涉分子对的数量、实验不确定性、实验来源以及真实物理效应的变异性。本文针对用于评估MMPA的统计学基础展开了分析。我们通过对比平均CHEMBL hERG MMPA结果与具有极端转化效应的分子对,阐明了整体MMPA统计特征与单个分子对之间的关联。将CHEMBL的分析结果与诺华(Novartis)的数据集进行对比后,我们发现:若将实验不确定性纳入考量,具有显著转化效应的分子对结果具有高度一致性。这一结果表明,针对无统计学显著性的MMPA预测需保持谨慎,但同时也凸显了经统计学验证的MMPA的稳健性,并证实了基于公共数据库的MMPA分析可为药物化学研究提供极具价值的结果。




