遇见数据集

Data from: Bayes factors unmask highly variable information content, bias, and extreme influence in phylogenomic analyses

收藏
DataONE2016-11-09 更新2024-06-26 收录
数据链接:
官方服务:

资源简介:

As the application of genomic data in phylogenetics has become routine, a number of cases have arisen where alternative datasets strongly support conflicting conclusions. This sensitivity to analytical decisions has prevented firm resolution of some of the most recalcitrant nodes in the tree of life. To better understand the causes and nature of this sensitivity, we analyzed several phylogenomic datasets using an alternative measure of topological support (the Bayes factor) that both demonstrates and averts several limitations of more frequently employed support measures (such as Markov chain Monte Carlo estimates of posterior probabilities). Bayes factors reveal important, previously hidden, differences across six “phylogenomic” datasets collected to resolve the phylogenetic placement of turtles within Amniota. These datasets vary substantially in their support for well-established amniote relationships, particularly in the proportion of genes that contain extreme amounts of information as well as the proportion that strongly reject these uncontroversial relationships. All six datasets contain little information to resolve the phylogenetic placement of turtles relative to other amniotes. Bayes factors also reveal that a very small number of extremely influential genes (less than one percent of genes in a dataset) can fundamentally change significant phylogenetic conclusions. In one example, these genes are shown to contain previously unrecognized paralogs. This study demonstrates both that the resolution of difficult phylogenomic problems remains sensitive to seemingly minor analysis details, and that Bayes factors are a valuable tool for identifying and solving these challenges.

随着基因组数据在系统发育学(phylogenetics)中的应用日趋常规化,诸多案例显示不同数据集会强力支持相互矛盾的结论。这种对分析决策的敏感性,阻碍了生命之树中部分最为顽固的节点的确定性解析。为更好地理解该敏感性的成因与本质,我们采用拓扑结构支持度的替代度量指标——贝叶斯因子(Bayes factor),对多套系统发育基因组数据集展开分析。该指标既能够展现,同时又规避了更为常用的支持度度量方式(如后验概率的马尔可夫链蒙特卡洛(Markov chain Monte Carlo)估计)所存在的若干局限。贝叶斯因子揭示了此前未被发现的重要差异:在为解析龟类在羊膜动物(Amniota)中的系统发育位置而收集的六套“系统发育基因组”数据集中,各数据集对已确立的羊膜动物类群关系的支持度存在显著差异,尤其体现在携带极端信息量的基因比例,以及强烈反对这些无争议类群关系的基因比例两方面。六套数据集均几乎未携带足够信息以解析龟类相较于其他羊膜动物的系统发育位置。贝叶斯因子还显示,数量极少的极具影响力的基因(占数据集基因总数不足1%)可从根本上改变重要的系统发育结论。在一个案例中,这类基因被发现存在此前未被识别的旁系同源基因(paralogs)。本研究表明,即便看似细微的分析细节,仍会对棘手的系统发育基因组问题的解析产生显著影响;同时也证实贝叶斯因子是识别并解决此类挑战的宝贵工具。

创建时间:
2016-11-09
二维码
社区交流群
二维码
科研交流群
商业服务