Data from: Why do phylogenomic data sets yield conflicting trees? Data type influences the avian tree of life more than taxon sampling
收藏资源简介:
Phylogenomics, the use of large-scale data matrices in phylogenetic analyses, has been viewed as the ultimate solution to the problem of resolving difficult nodes in the tree of life. However, it has become clear that analyses of these large genomic data sets can also result in conflicting estimates of phylogeny. Here, we use the early divergences in Neoaves, the largest clade of extant birds, as a “model system” to understand the basis for incongruence among phylogenomic trees. We were motivated by the observation that trees from two recent avian phylogenomic studies exhibit conflicts. Those studies used different strategies: 1) collecting many characters [42 mega base pairs (Mbp) of sequence data] from 48 birds, sometimes including only one taxon for each major clade; and 2) collecting fewer characters (0.4 Mbp) from 198 birds, selected to subdivide long branches. However, the studies also used different data types: the taxon-poor data matrix comprised 68% non-coding sequences whereas coding exons dominated the taxon-rich data matrix. This difference raises the question of whether the primary reason for incongruence is the number of sites, the number of taxa, or the data type. To test among these alternative hypotheses we assembled a novel, large-scale data matrix comprising 90% non-coding sequences from 235 bird species. Although increased taxon sampling appeared to have a positive impact on phylogenetic analyses the most important variable was data type. Indeed, by analyzing different subsets of the taxa in our data matrix we found that increased taxon sampling actually resulted in increased congruence with the tree from the previous taxon-poor study (which had a majority of non-coding data) instead of the taxon-rich study (which largely used coding data). We suggest that the observed differences in the estimates of topology for these studies reflect data-type effects due to violations of the models used in phylogenetic analyses, some of which may be difficult to detect. If incongruence among trees estimated using phylogenomic methods largely reflects problems with model fit developing more “biologically-realistic” models is likely to be critical for efforts to reconstruct the tree of life.
系统发育基因组学(Phylogenomics),即在系统发育分析中应用大规模数据矩阵,曾被视作解决生命之树疑难节点解析难题的终极方案。然而现有研究表明,对这类大型基因组数据集的分析同样可能得到相互冲突的系统发育估计结果。本研究以现存鸟类最大演化支——新鸟类(Neoaves)的早期分化事件作为模式系统,旨在探究系统发育基因组学树形间不一致性的成因。我们的研究动机源于两项近期鸟类系统发育基因组学研究所得树形存在冲突的观测结果。这两项研究采用了不同的研究策略:其一,对48种鸟类采集了大量性状数据[42兆碱基对(Mbp)的序列数据],部分类群仅选取每个主要演化支的单一代表;其二,对198种鸟类采集了较少的性状数据(0.4 Mbp),所选类群用于拆分长分支。此外,两项研究使用的数据类型也存在差异:类群稀疏的数据矩阵中68%为非编码序列,而类群密集的数据矩阵则以编码外显子为主。这一差异引出了核心问题:导致不一致性的主要原因究竟是位点数量、类群数量,还是数据类型?为检验这三种备选假说,我们构建了一个全新的大规模数据矩阵,该矩阵包含来自235种鸟类的90%非编码序列。尽管增加类群采样量似乎对系统发育分析产生了积极影响,但最为关键的变量仍是数据类型。通过对数据矩阵中不同类群子集进行分析,我们发现,增加类群采样量实际上使得树形与此前类群稀疏研究(该研究以非编码数据为主)的结果更为一致,而非与类群密集研究(该研究主要使用编码数据)的结果相符。我们认为,两项研究得到的拓扑结构估计差异,反映了因违反系统发育分析所用模型而产生的数据类型效应,其中部分效应可能难以被检测到。如果基于系统发育基因组学方法得到的树形之间的不一致性,主要反映的是模型适配性问题,那么开发更具生物学真实性的模型,对于重建生命之树的研究而言或将至关重要。



