遇见数据集

Data from: A branch-heterogeneous model of protein evolution for efficient inference of ancestral sequences

收藏
DataONE2013-03-04 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

Most models of nucleotide or amino acid substitution used in phylogenetic studies assume that the evolutionary process has been homogeneous across lineages and that composition of nucleotides or amino acids has remained the same throughout the tree. These oversimplified assumptions are refuted by the observation that compositional variability characterizes extant biological sequences. Branch-heterogeneous models of protein evolution that account for compositional variability have been developed, but are not yet in common use because of the large number of parameters required, leading to high computational costs and potential overparameterization. Here, we present a new branch-nonhomogeneous and nonstationary model of protein evolution that captures more accurately the high complexity of sequence evolution. This model, henceforth called Correspondence and likelihood analysis (COaLA), makes use of a correspondence analysis to reduce the number of parameters to be optimized through maximum likelihood, focusing on most of the compositional variation observed in the data. The model was thoroughly tested on both simulated and biological data sets to show its high performance in terms of data fitting and CPU time. COaLA efficiently estimates ancestral amino acid frequencies and sequences, making it relevant for studies aiming at reconstructing and resurrecting ancestral amino acid sequences. Finally, we applied COaLA on a concatenate of universal amino acid sequences to confirm previous results obtained with a nonhomogeneous Bayesian model regarding the early pattern of adaptation to optimal growth temperature, supporting the mesophilic nature of the Last Universal Common Ancestor.

系统发育研究中常用的核苷酸或氨基酸替换模型,大多假设进化过程在各谱系间保持均一,且核苷酸/氨基酸组成在整个系统发育树中始终恒定。现有观测表明,现存生物序列普遍存在组成变异性,这一结论推翻了上述过于简化的假设。为应对组成变异性而开发的蛋白质进化分支异质性模型,目前尚未得到广泛应用——这类模型所需参数体量庞大,会带来极高的计算成本,还可能存在过参数化问题。本文提出一种全新的蛋白质进化分支非均一、非稳态模型,可更精准地捕捉序列进化的高度复杂性。该模型此后命名为对应分析与似然分析(Correspondence and Likelihood Analysis, COaLA),它借助对应分析缩减最大似然估计所需优化的参数数量,重点拟合数据中观测到的绝大部分组成变异。本研究通过模拟数据集与真实生物数据集对该模型开展了全面测试,结果显示其在数据拟合度与CPU运行时间方面均表现优异。COaLA可高效估算祖先氨基酸频率与序列,因此适用于旨在重建、复活祖先氨基酸序列的相关研究。最后,我们将COaLA应用于通用氨基酸序列串联数据集,验证了此前基于非均一贝叶斯模型得到的关于适应性向最适生长温度演化的早期模式的研究结果,支持了通用最后共同祖先(Last Universal Common Ancestor)为中温生物的结论。

创建时间:
2013-03-04
二维码
社区交流群
二维码
科研交流群
商业服务