Data from: Analyzing contentious relationships and outlier genes in phylogenomics
收藏资源简介:
Recent studies have demonstrated that conflict is common among gene trees in phylogenomic studies, and that less than one percent of genes may ultimately drive species tree inference in supermatrix analyses. Here, we examined two datasets where supermatrix and coalescent-based species trees conflict. We identified two highly influential “outlier” genes in each dataset. When removed from each dataset, the inferred supermatrix trees matched the topologies obtained from coalescent analyses. We also demonstrate that, while the outlier genes in the vertebrate dataset have been shown in a previous study to be the result of errors in orthology detection, the outlier genes from a plant dataset did not exhibit any obvious systematic error and therefore may be the result of some biological process yet to be determined. While topological comparisons among a small set of alternate topologies can be helpful in discovering outlier genes, they can be limited in several ways, such as assuming all genes share the same topology. Coalescent species tree methods relax this assumption but do not explicitly facilitate the examination of specific edges. Coalescent methods often also assume that conflict is the result of incomplete lineage sorting (ILS). Here we explored a framework that allows for quickly examining alternative edges and support for large phylogenomic datasets that does not assume a single topology for all genes. For both datasets, these analyses provided detailed results confirming the support for coalescent-based topologies. This framework suggests that we can improve our understanding of the underlying signal in phylogenomic datasets by asking more targeted edge-based questions.
已有研究表明,系统基因组学(phylogenomics)研究中基因树间的冲突现象极为普遍;在超矩阵(supermatrix)分析中,最终推动物种树推断的基因占比可能不足1%。本研究针对两类存在超矩阵物种树与溯祖物种树冲突的数据集展开分析,在每个数据集中均鉴定出两个极具影响力的“异常”基因。将这两个异常基因从数据集中移除后,推断得到的超矩阵物种树拓扑结构与溯祖分析得到的拓扑结构完全一致。我们还证实,尽管既往研究显示脊椎动物数据集的异常基因源于直系同源检测失误,但植物数据集的异常基因未表现出任何明显的系统性误差,因此其成因可能来自尚未明确的某种生物学过程。尽管针对少量备选拓扑结构开展拓扑比较有助于发现异常基因,但这类方法存在多方面局限,例如其默认所有基因共享同一套拓扑结构。溯祖物种树推断方法放宽了这一假设,但无法直接支持对特定分支的针对性检视。此外,溯祖方法通常默认冲突仅由不完全谱系分选(incomplete lineage sorting, ILS)引发。本研究探索了一套分析框架,可在无需假设所有基因共享同一拓扑结构的前提下,快速对大型系统基因组数据集的备选分支及其支持度展开检视。针对两类数据集的分析均得到了详实结果,证实了溯祖物种树拓扑结构的支持度。该框架表明,通过提出更具针对性的分支相关问题,我们能够加深对系统基因组数据集内在信号的理解。



