Data from: The impact of paralogy on phylogenomic studies - a case study on annelid relationships
收藏资源简介:
Phylogenomic studies based on hundreds of genes derived from expressed sequence tags libraries are increasingly used to reveal the phylogeny of taxa. A prerequisite for these studies is the assignment of genes into clusters of orthologous sequences. Sophisticated methods of orthology prediction are used in such analyses, but it is rarely assessed whether paralogous sequences have been erroneously grouped together as orthologous sequences after the prediction, and whether this had an impact on the phylogenetic reconstruction using a super-matrix approach. Herein, I tested the impact of paralogous sequences on the reconstruction of annelid relationships based on phylogenomic datasets. Using single-partition analyses, screening for bootstrap support, blast searches and pruning of sequences in the supermatrix, wrongly assigned paralogous sequences were found in eight partitions and the placement of five taxa (the annelids Owenia, Scoloplos, Sthenelais and Eurythoe and the nemertean Cerebratulus) including the robust bootstrap support could be attributed to the presence of paralogous sequences in two partitions. Excluding these sequences resulted in a different, weaklier supported placement for these taxa. Moreover, the analyses revealed that paralogous sequences impacted the reconstruction when only a single taxon represented a previously supported higher taxon such as a polychaete family. One possibility of a priori detection of wrongly assigned paralogous sequences could to combine 1) a screening of single-partition analyses based on criteria such as nodal support or internal branch length with 2) blast searches of suspicious cases as presented herein. Also possible are a posteriori approaches in which support for specific clades is investigated by comparing alternative hypotheses based on differences in per-site likelihoods. Increasing the sizes of EST libraries will also decrease the likelihood of wrongly assigned paralogous sequences, and in the case of orthology prediction methods like HaMStR it is likewise decreased by using more than one reference taxon.
基于源自表达序列标签(expressed sequence tags, EST)文库的数百个基因开展的系统发育基因组学研究,正日益广泛地用于揭示类群的系统发育关系。此类研究的核心前提是将基因分配至直系同源序列(orthologous sequences)簇。此类分析中会采用成熟的直系同源预测方法,但现有研究极少评估:在预测流程完成后,是否存在旁系同源序列(paralogous sequences)被错误归类为直系同源序列的情况,以及这一现象是否会对采用超级矩阵法(super-matrix approach)开展的系统发育重建造成影响。 本文中,笔者针对旁系同源序列对基于系统发育基因组学数据集的环节动物(annelid)亲缘关系重建的影响开展了测试。通过单分区分析、自举支持度(bootstrap support)筛查、BLAST搜索(blast searches)以及超级矩阵的序列修剪操作,研究团队在8个分区中发现了被错误归类的旁系同源序列;而5个类群(环节动物Owenia、Scoloplos、Sthenelais、Eurythoe以及纽形动物nemertean的Cerebratulus)的系统发育位置,包括其较高的自举支持度,均可归因于两个分区中存在旁系同源序列。剔除这些序列后,这些类群的系统发育位置发生改变,且其支持度显著降低。 此外,分析结果显示,当某一高级分类单元(如多毛类科)仅由单个类群代表时,旁系同源序列会对系统发育重建造成显著影响。 一种可用于预先(a priori)检测错误归类旁系同源序列的方案,可结合以下两种方法:1)基于节点支持度或内部分支长度等标准开展的单分区分析筛查;2)针对可疑案例开展如本文所述的BLAST搜索。另外也可采用事后(a posteriori)分析方法,即通过基于每位点似然值(per-site likelihoods)差异的不同假设进行比较,以探究特定支系(clades)的支持情况。 扩大EST文库的规模同样可降低旁系同源序列被错误归类的概率;而对于HaMStR这类直系同源预测工具而言,使用多个参考类群也可降低此类错误的发生概率。



