Fortuitous Correlations in Molecular Dynamics Simulations: Their Harmful Influence on the Probability Distributions of the Main Principal Components
收藏资源简介:
Nonsense correlations frequently develop between independent random variables that evolve with time. Therefore, it is not surprising that they appear between the components of vectors carrying out multidimensional random walks, such as those describing the trajectories of biomolecules in molecular dynamics simulations. The existence of these correlations does not imply in itself a problem. Still, it can present a problem when the trajectories are analyzed with an algorithm such as the Principal Component Analysis (PCA) because it seeks to maximize correlations without discriminating whether they have physical origin or not. In this Article, we employ random walks occurring on multidimensional harmonic potentials to evaluate the influence of fortuitous correlations in PCA. We demonstrate that, because of them, this algorithm affords misleading results when applied to a single trajectory. The errors do not only affect the directions of the first eigenvectors and their eigenvalues, but the very definition of the molecule’s “essential space” may be wrong. Additionally, the main principal component’s probability distributions present artificial structures which do not correspond with the shape of the potential energy surface. Finally, we show that the PCA of two realistic protein models, human serum albumin and lysozyme, behave similarly to the simple harmonic models. In all cases, the problems can be mitigated and eventually eliminated by doing PCA on concatenated trajectories formed from a large enough number of individual simulations.
随时间演化的独立随机变量之间,常会出现无意义的伪相关关系。因此,在进行多维随机游走的矢量分量之间出现这类相关关系并不足为奇——例如描述分子动力学模拟中生物分子轨迹的矢量分量。这类相关关系的存在本身并非问题所在。但当使用主成分分析(Principal Component Analysis, PCA)这类算法分析轨迹时,便可能引发问题:该算法会最大化相关关系,却无法区分这些相关关系是否具有物理起源。本文采用多维简谐势场下的随机游走模型,评估伪相关对主成分分析的影响。我们证明,由于这类伪相关的存在,仅应用单条轨迹进行主成分分析时,算法会得到误导性结果。此类误差不仅会影响前几个本征向量的方向及其本征值,甚至可能导致对分子"本质空间"的定义出现偏差。此外,主成分的概率分布会出现人工构造的伪结构,与势能面的真实形状并不匹配。最后,我们通过两个真实蛋白质模型——人血清白蛋白与溶菌酶——的主成分分析结果证明,其表现与简单简谐模型一致。在所有场景中,通过对足够多独立模拟得到的拼接轨迹进行主成分分析,即可缓解并最终消除此类问题。



