Data from: Impacts of inference method and dataset filtering on phylogenomic resolution in a rapid radiation of ground squirrels (Xerinae: Marmotini)
收藏资源简介:
Phylogenomic datasets are illuminating many areas of the Tree of Life. However, the large size of these datasets alone may be insufficient to resolve problematic nodes in the most rapid evolutionary radiations, because inferences in zones of extraordinarily low phylogenetic signal can be sensitive to the model and method of inference, as well as the information content of loci employed. We used a dataset of >3,950 ultraconserved element (UCE) loci from a classic mammalian radiation, ground-dwelling squirrels of the tribe Marmotini (Sciuridae: Xerinae), to assess sensitivity of phylogenetic estimates to varying per-locus information content across 4 different inference methods (RAxML, ASTRAL, NJst, SVDquartets). Persistent discordance was found in topology and bootstrap support between concatenation- and coalescent-based inferences; among methods within the coalescent framework; and within all methods in response to different filtering scenarios. Contrary to some recent empirical UCE-based studies, filtering by information content did not promote complete among-method concordance. Nevertheless, filtering did improve concordance relative to randomly selected locus sets, largely via improved consistency of two-step summary methods (particularly NJst) under conditions of higher average per-locus variation (and thus increasing gene tree precision). The benefits of dataset filtering are notably variable among classes of inference methods and across different evolutionary scenarios, reiterating the complexities of resolving rapid radiations, even with robust taxon and character sampling.
系统发育基因组学数据集(phylogenomic datasets)正为阐明生命之树的诸多研究领域带来关键洞见。然而,仅凭这类数据集的庞大规模,或许仍不足以解决演化快速辐射事件中的疑难系统发育节点:在系统发育信号极度匮乏的区域开展推断时,结果极易受到所采用的推断模型、推断方法以及基因座信息含量的影响。本研究以经典哺乳动物辐射演化类群——松鼠科(Sciuridae)旱獭族(Marmotini)的陆栖松鼠为研究材料,利用包含3950余个超保守元件(ultraconserved element, UCE)基因座的数据集,针对4种不同的系统发育推断方法(RAxML、ASTRAL、NJst、SVDquartets),评估了系统发育估计结果对不同单位基因座信息含量的敏感性。研究发现,基于串联的推断与基于溯祖的推断之间、溯祖框架内不同方法之间,以及所有方法在不同过滤方案下的结果中,均存在持续存在的拓扑结构与自举支持度(bootstrap support)差异。与近期部分基于UCE的实证研究结论相悖,仅依据信息含量进行基因座过滤并未实现不同方法间的完全一致。不过相较于随机选取的基因座集,过滤处理确实提升了方法间的一致性,这主要得益于在更高的平均单位基因座变异水平(即更高的基因树精度)条件下,两步汇总方法(尤其是NJst)的推断一致性得到了改善。值得注意的是,数据集过滤的益处因推断方法类别与不同演化场景而异,这再次强调了即便在分类群与性状采样都较为充分的情况下,解析快速辐射事件的演化关系仍存在诸多复杂性。



