Data from: Bayes factors unmask highly variable information content, bias, and extreme influence in phylogenomic analyses
收藏资源简介:
As the application of genomic data in phylogenetics has become routine, a number of cases have arisen where alternative datasets strongly support conflicting conclusions. This sensitivity to analytical decisions has prevented firm resolution of some of the most recalcitrant nodes in the tree of life. To better understand the causes and nature of this sensitivity, we analyzed several phylogenomic datasets using an alternative measure of topological support (the Bayes factor) that both demonstrates and averts several limitations of more frequently employed support measures (such as Markov chain Monte Carlo estimates of posterior probabilities). Bayes factors reveal important, previously hidden, differences across six “phylogenomic” datasets collected to resolve the phylogenetic placement of turtles within Amniota. These datasets vary substantially in their support for well-established amniote relationships, particularly in the proportion of genes that contain extreme amounts of information as well as the proportion that strongly reject these uncontroversial relationships. All six datasets contain little information to resolve the phylogenetic placement of turtles relative to other amniotes. Bayes factors also reveal that a very small number of extremely influential genes (less than one percent of genes in a dataset) can fundamentally change significant phylogenetic conclusions. In one example, these genes are shown to contain previously unrecognized paralogs. This study demonstrates both that the resolution of difficult phylogenomic problems remains sensitive to seemingly minor analysis details, and that Bayes factors are a valuable tool for identifying and solving these challenges.
随着基因组数据在系统发育学(phylogenetics)中的应用日趋常规化,诸多案例显示不同数据集会强力支持相互矛盾的研究结论。这种对分析选择的敏感性,阻碍了生命之树中部分最为棘手的分支节点的明确界定。 为更好地理解这种敏感性的成因与本质,我们采用拓扑支持度的替代度量指标——贝叶斯因子(Bayes factor)——对多套系统发育基因组数据集(phylogenomic datasets)展开分析。该指标既能展现现有常用支持度度量方法的诸多局限,又可规避这些局限;而常用支持度度量方法包括后验概率(posterior probabilities)的马尔可夫链蒙特卡洛(Markov chain Monte Carlo)估计值。 贝叶斯因子揭示了6套旨在厘清龟类在羊膜动物(Amniota)中系统发育定位的系统发育基因组数据集之间,存在此前未被发现的重要差异。这些数据集对已被广泛认可的羊膜类演化关系的支持程度存在显著差异,具体体现在携带极多信息的基因占比,以及强力反对这些无争议关系的基因占比两方面。 上述6套数据集均几乎未提供足够信息,以厘清龟类相较于其他羊膜类的系统发育定位。贝叶斯因子还显示,极少数极具影响力的基因(占单套数据集基因总数的1%以下)即可从根本上改变重要的系统发育研究结论。在一则案例中,这类基因被发现携带此前未被识别的旁系同源基因(paralog)。 本研究既证实了棘手的系统发育基因组问题的解决仍对看似细微的分析细节极为敏感,也证明了贝叶斯因子是识别并解决此类挑战的宝贵工具。



