遇见数据集

Data from: Modeling site heterogeneity with posterior mean site frequency profiles accelerates accurate phylogenomic estimation

收藏
DataONE2017-08-04 更新2024-06-26 收录
数据链接:
官方服务:

资源简介:

Proteins have distinct structural and functional constraints at different sites that lead to site-specific preferences for particular amino acid residues as the sequences evolve. Heterogeneity in the amino acid substitution process between sites is not modeled by commonly used empirical amino acid exchange matrices. Such model misspecification can lead to artefacts in phylogenetic estimation such as long-branch attraction. Although sophisticated site-heterogeneous mixture models have been developed to address this problem in both Bayesian and maximum likelihood (ML) frameworks, their formidable computational time and memory usage severely limits their use in large phylogenomic analyses. Here we propose a posterior mean site frequency (PMSF) method as a rapid and efficient approximation to full empirical profile mixture models for ML analysis. The PMSF approach assigns a conditional mean amino acid frequency profile to each site calculated based on a mixture model fitted to the data using a preliminary guide tree. These PMSF profiles can then be used for in-depth tree-searching in place of the full mixture model. Compared with widely used empirical mixture models with k classes, our implementation of PMSF in IQ-TREE (http://www.iqtree.org) speeds up the computation by approximately k /1.5-fold and requires a small fraction of the RAM. Furthermore, this speedup allows, for the first time, full nonparametric bootstrap analyses to be conducted under complex site-heterogeneous models on large concatenated data matrices. Our simulations and empirical data analyses demonstrate that PMSF can effectively ameliorate long-branch attraction artefacts. In some empirical and simulation settings PMSF provided more accurate estimates of phylogenies than the mixture models from which they derive.

蛋白质在不同位点存在独特的结构与功能约束,这使得序列演化过程中,不同位点对特定氨基酸残基产生位点特异性偏好。当前常用的经验氨基酸替换矩阵并未对位点间氨基酸替换过程的异质性进行建模。此类模型设定误差可能引发系统发育推断中的人为偏差,例如长枝吸引(long-branch attraction)。尽管针对贝叶斯与最大似然(maximum likelihood, ML)框架下的该问题,已开发出复杂的位点异质性混合模型,但这类模型高昂的计算耗时与内存占用,严重限制了其在大规模系统发育组学分析中的应用。本文提出一种后验均值位点频率(posterior mean site frequency, PMSF)方法,可作为面向最大似然分析的完整经验谱混合模型的快速高效近似方案。PMSF方法会为每个位点分配一个条件平均氨基酸频率谱,该谱基于通过初步指导树拟合至数据集的混合模型计算得到。随后,这些PMSF谱可替代完整混合模型,用于深度树搜索流程。相较于广泛使用的含k个类别的经验混合模型,我们在IQ-TREE(http://www.iqtree.org)中实现的PMSF方法可将计算速度提升约k/1.5倍,且仅需极小部分随机存取内存(RAM)。此外,该速度提升首次使得在大规模串联数据矩阵上,基于复杂位点异质性模型开展完整的非参数自举(bootstrap)分析成为可能。我们的模拟与实证数据分析表明,PMSF可有效缓解长枝吸引人为偏差。在部分实证与模拟场景中,PMSF相较于其衍生的混合模型,能够提供更为准确的系统发育树估计结果。

创建时间:
2017-08-04
二维码
社区交流群
二维码
科研交流群
商业服务