遇见数据集

Data from: Phylogenetic tree estimation with and without alignment: new distance methods and benchmarking

收藏
DataONE2016-08-24 更新2024-06-26 收录
数据链接:
官方服务:

资源简介:

Phylogenetic tree inference is a critical component of many systematic and evolutionary studies. The majority of these studies are based on the two-step process of multiple sequence alignment followed by tree inference, despite persistent evidence that the alignment step can lead to biased results. Here we present a two-part study that first presents PaHMM-Tree, a novel neighbour joining-based method that estimates pairwise distances without assuming a single alignment. We then use simulations to benchmark its performance against a wide-range of other phylogenetic tree inference methods, including the first comparison of alignment-free distance-based methods against more conventional tree estimation methods. Our new method for calculating pairwise distances based on statistical alignment provides distance estimates that are as accurate as those obtained using standard methods based on the true alignment. Pairwise distance estimates based on the two-step process tend to be substantially less accurate. This improved performance carries through to tree inference, where PaHMM-Tree provides more accurate tree estimates than all of the pairwise distance methods assessed. For close to moderately divergent sequence data we find that the two-step methods using statistical inference, where information from all sequences is included in the estimation procedure, tend to perform better than PaHMM-Tree, particularly full statistical alignment, which simultaneously estimates both the tree and the alignment. For deep divergences we find the alignment step becomes so prone to error that our distance-based PaHMM-Tree outperforms all other methods of tree inference. Finally, we find that the accuracy of alignment-free methods tends to decline faster than standard two-step methods in the presence of alignment uncertainty, and identify no conditions where alignment-free methods are equal to or more accurate than standard phylogenetic methods even in the presence of substantial alignment error.

系统发育树推断(Phylogenetic tree inference)是诸多分类学与进化研究的核心组成部分。尽管已有持续证据表明比对步骤可能引入结果偏差,但此类研究大多采用先多序列比对(multiple sequence alignment)再进行树推断的两步流程。本研究分为两个部分:首先提出PaHMM-Tree——一种基于邻接法(Neighbour Joining)的新型方法,无需预设单一比对即可估算两两距离。随后我们通过模拟实验,将该方法与多款主流系统发育树推断方法开展性能基准测试,其中首次实现了无比对距离法与传统树推断方法的对比分析。我们基于统计比对(statistical alignment)提出的两两距离计算新方法,其距离估算精度与基于真实比对的标准方法相当;而基于两步流程的两两距离估算精度则显著偏低。这一性能优势同样延伸至树推断环节:在所评估的所有两两距离法中,PaHMM-Tree可生成精度更高的树结构推断结果。针对近中等分化程度的序列数据,采用统计推断(将所有序列信息纳入估算流程)的两步法性能优于PaHMM-Tree,其中可同时估算树结构与比对结果的全统计比对(full statistical alignment)表现尤为突出。而对于深度分化的序列数据,比对步骤极易产生误差,此时基于距离法的PaHMM-Tree在所有树推断方法中性能最优。最后我们发现,当存在比对不确定性时,无比对法(alignment-free methods)的精度下降速度快于标准两步法;且即便存在显著的比对误差,我们也未发现任何可使无比对法精度达到或超越标准系统发育学方法的应用场景。

创建时间:
2016-08-24
二维码
社区交流群
二维码
科研交流群
商业服务