Data from: Comparison of methods for molecular species delimitation across a range of speciation scenarios
收藏资源简介:
Species are fundamental units in biological research and can be defined on the basis of various operational criteria. There has been growing use of molecular approaches for species delimitation. Among the most widely used methods, the generalized mixed Yule-coalescent (GMYC) and Poisson tree processes (PTP) were designed for the analysis of single-locus data but are often applied to concatenations of multilocus data. In contrast, the Bayesian multispecies coalescent approach in the software BPP explicitly models the evolution of multilocus data. In this study, we compare the performance of GMYC, PTP, and BPP using synthetic data generated by simulation under various speciation scenarios. We show that in the absence of gene flow, the main factor influencing the performance of these methods is the ratio of population size to divergence time, while number of loci and sample size per species have smaller effects. Given appropriate priors and correct guide trees, BPP shows lower rates of species overestimation and underestimation, and is generally robust to various potential confounding factors except high levels of gene flow. The single-threshold GMYC and the best strategy that we identified in PTP generally perform well for scenarios involving more than a single putative species when gene flow is absent, but PTP outperforms GMYC when fewer species are involved. Both methods are more sensitive than BPP to the effects of gene flow and potential confounding factors. Case studies of bears and bees further validate some of the findings from our simulation study, and reveal the importance of using an informed starting point for molecular species delimitation. Our results highlight the key factors affecting the performance of molecular species delimitation, with potential benefits for using these methods within an integrative taxonomic framework.
物种是生物学研究的基本单元,可依据多种操作标准进行界定。当前分子物种界定方法的应用愈发广泛。在目前应用最广泛的方法中,广义混合耶尔-合并模型(generalized mixed Yule-coalescent, GMYC)与泊松树过程模型(Poisson tree processes, PTP)最初专为单基因座数据分析设计,但常被用于多基因座数据的拼接分析。与之相对,软件BPP中的贝叶斯多物种合并模型则显式地对多基因座数据的演化过程进行建模。本研究借助不同物种形成场景下通过模拟生成的合成数据集,对比了GMYC、PTP与BPP的性能表现。研究结果显示,在无基因交流的情况下,影响这些方法性能的核心因素为种群规模与分化时间的比值,而基因座数量及每个物种的样本量仅产生较小影响。在设置合理先验分布且使用正确引导树(guide tree)的前提下,BPP的物种高估与低估率均更低,且除高程度基因交流外,通常对各类潜在混杂因素均具有稳健性。单阈值GMYC以及我们在PTP中确定的最优策略,在无基因交流且涉及多个推定物种的场景下通常表现优异,但当物种数量较少时,PTP的性能优于GMYC。相较于BPP,这两种方法对基因交流与潜在混杂因素的影响更为敏感。针对熊类与蜂类的案例研究进一步验证了本模拟研究的部分结论,并揭示了为分子物种界定选择合理起始点的重要性。本研究结果明确了影响分子物种界定方法性能的关键因素,可为在整合分类学框架下应用此类方法提供参考价值。



