遇见数据集

Data from: Pruning rogue taxa improves phylogenetic accuracy: an efficient algorithm and webservice

收藏
DataONE2012-10-12 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

The presence of rogue taxa (rogues) in a set of trees can frequently have a negative impact on the results of a bootstrap analysis (e.g., the overall support in consensus trees). We introduce an efficient graph-based algorithm for rogue taxon identification as well as an interactive web-service implementing this algorithm. Compared to our previous method, the new algorithm is up to four orders of magnitude faster, while returning qualitatively identical results. Because of this significant improvement in scalability, the new algorithm can now identify substantially more complex and compute-intensive rogue taxon constellations. On a large and diverse collection of real-world datasets, we show that, our method yields better supported reduced/pruned consensus trees than any competing rogue taxon identification method. Using the parallel version of our open-source code, we successfully identified rogue taxa in a set of 100 trees with 116,334 taxa each. Using simulated datasets we show that, when removing/pruning rogue taxa with our method from a tree set, we consistently obtain bootstrap consensus trees as well as maximum likelihood trees that are topologically closer to the respective true trees.

在系统发育树集合中,异常分类群(rogue taxa,简称rogues)的存在往往会对自举分析(bootstrap analysis)的结果产生负面影响,例如共识树(consensus trees)的整体支持度。本研究提出了一种用于异常分类群识别的高效图基算法(graph-based algorithm),以及实现该算法的交互式网页服务工具。相较于本团队此前的方法,新算法的运行速度最高可提升四个数量级,且所得结果在质量上保持一致。得益于可扩展性的显著提升,新算法如今可识别更为复杂、计算量更大的异常分类群组合。在大规模且多样的真实世界数据集集合上,本研究证明,相较于其他同类异常分类群识别方法,本方法所得到的精简/修剪后共识树拥有更高的支持度。借助本团队开源代码(open-source code)的并行版本,我们成功在每组包含116334个分类群、共100棵系统发育树的集合中识别出了异常分类群。通过模拟数据集实验,本研究证明,若利用本方法从树集合中移除/修剪异常分类群,最终得到的自举共识树以及最大似然树(maximum likelihood trees)在拓扑结构上均更接近对应的真实系统发育树。

创建时间:
2012-10-12
二维码
社区交流群
二维码
科研交流群
商业服务