遇见数据集

Data from: Who let the CAT out of the bag? accurately dealing with substitutional heterogeneity in phylogenomic analyses

收藏
DataONE2016-09-07 更新2024-06-26 收录
数据链接:
官方服务:

资源简介:

As phylogenetic datasets have increased in size, site-heterogeneous substitution models such as CAT-F81 and CAT-GTR have been advocated in favor of other models because they purportedly suppress long-branch attraction (LBA). These models are two of the most commonly used models in phylogenomics, and they have been applied to a variety of taxa ranging from Drosophila to land plants. However, many arguments in favor of CAT models have been based on tenuous assumptions about the true phylogeny rather than rigorous testing with known trees via simulation. Moreover, CAT models have not been compared to other approaches for handling substitutional heterogeneity such as data partitioning with site-homogeneous substitution models. We simulated amino acid sequence datasets with substitutional heterogeneity on a variety of tree shapes including those susceptible to LBA. Data were analyzed with both CAT models and partitioning to explore model performance; in total over 670,000 CPU hours were used, of which over 97% was spent running analyses with CAT models. In many cases, all models recovered branching patterns that were identical to the known tree. However, CAT-F81 consistently performed worse than other models in inferring the correct branching patterns, and both CAT models often overestimated substitutional heterogeneity. Additionally, reanalysis of two empirical metazoan datasets supports the notion that CAT-F81 tends to recover less accurate trees than data partitioning and CAT-GTR. Given these results, we conclude that partitioning and CAT-GTR perform similarly in recovering accurate branching patterns. However, computation time can be orders of magnitude less for data partitioning, with commonly used implementations of CAT-GTR often failing to reach completion in a reasonable time frame (i.e., for Bayesian analyses to converge). Practices such as removing constant sites and parsimony uninformative characters, or using CAT-F81 when CAT-GTR is deemed too computationally expensive, cannot be logically justified. Given clear problems with CAT-F81, phylogenies previously inferred with this model should be reassessed.

随着系统发育数据集规模不断扩大,CAT-F81、CAT-GTR这类位点异质替换模型(site-heterogeneous substitution models)相较于其他模型更受青睐,因为据称它们能够抑制长支吸引(long-branch attraction, LBA)效应。此类模型是系统发育组学中最常用的两类替换模型,已被应用于从果蝇到陆生植物的各类类群研究。然而,支持CAT模型的诸多论证多基于对真实系统发育关系的薄弱假设,而非通过模拟已知树结构开展的严格检验。此外,CAT模型尚未与其他处理替换异质性的方法进行对比,例如采用位点同质替换模型的数据分区策略。 我们在多种树结构(包括易受长支吸引影响的树结构)上模拟了携带替换异质性的氨基酸序列数据集。分别使用CAT模型与数据分区方法对数据进行分析,以探究不同模型的表现;本次分析总计消耗超过67万CPU小时,其中97%以上的算力用于CAT模型的运行。 在多数情形下,所有模型均能恢复与已知树结构完全一致的分支模式。但CAT-F81在推断正确分支模式方面始终劣于其他模型,且两类CAT模型往往会高估替换异质性程度。此外,对两项实测后生动物数据集的重新分析支持了这一结论:相较于数据分区与CAT-GTR,CAT-F81所恢复的系统发育树准确性往往更低。 基于上述结果,我们认为数据分区与CAT-GTR在恢复准确分支模式方面表现相近。但数据分区方法的计算耗时可低数个数量级,而常用的CAT-GTR实现往往无法在合理时限内完成计算(即贝叶斯分析难以收敛)。诸如移除恒定位点与简约法无信息性状,或是在CAT-GTR计算成本过高时改用CAT-F81这类操作,均无法获得逻辑上的支撑。鉴于CAT-F81存在明显缺陷,此前基于该模型推断得到的系统发育树应当重新评估。

创建时间:
2016-09-07
二维码
社区交流群
二维码
科研交流群
商业服务