遇见数据集

Data from: Comparison of methods for molecular species delimitation across a range of speciation scenarios

收藏
DataONE2018-02-13 更新2024-06-25 收录
数据链接:
官方服务:

资源简介:

Species are fundamental units in biological research and can be defined on the basis of various operational criteria. There has been growing use of molecular approaches for species delimitation. Among the most widely used methods, the generalized mixed Yule-coalescent (GMYC) and Poisson tree processes (PTP) were designed for the analysis of single-locus data but are often applied to concatenations of multilocus data. In contrast, the Bayesian multispecies coalescent approach in the software BPP explicitly models the evolution of multilocus data. In this study, we compare the performance of GMYC, PTP, and BPP using synthetic data generated by simulation under various speciation scenarios. We show that in the absence of gene flow, the main factor influencing the performance of these methods is the ratio of population size to divergence time, while number of loci and sample size per species have smaller effects. Given appropriate priors and correct guide trees, BPP shows lower rates of species overestimation and underestimation, and is generally robust to various potential confounding factors except high levels of gene flow. The single-threshold GMYC and the best strategy that we identified in PTP generally perform well for scenarios involving more than a single putative species when gene flow is absent, but PTP outperforms GMYC when fewer species are involved. Both methods are more sensitive than BPP to the effects of gene flow and potential confounding factors. Case studies of bears and bees further validate some of the findings from our simulation study, and reveal the importance of using an informed starting point for molecular species delimitation. Our results highlight the key factors affecting the performance of molecular species delimitation, with potential benefits for using these methods within an integrative taxonomic framework.

物种是生物学研究的基本单元,可依据多种操作标准进行界定。当前,用于物种界定的分子生物学方法应用愈发广泛。其中,应用最为广泛的广义混合尤尔-凝聚模型(generalized mixed Yule-coalescent, GMYC)与泊松树过程模型(Poisson tree processes, PTP),最初均设计用于单基因座数据分析,但实际应用中常被直接用于多基因座序列的联合数据集。与之形成对比的是,软件BPP所搭载的贝叶斯多物种凝聚分析方法,可显式建模多基因座数据的演化进程。 本研究通过在不同物种形成场景下模拟生成的数据集,对比了GMYC、PTP与BPP三种方法的分析性能。结果显示,在无基因流的前提下,影响这些方法性能的核心因素为种群规模与分化时间的比值,而基因座数量及每个物种的样本量对其影响相对较小。当先验设置合理且指导树(guide tree)正确时,BPP的物种过界定与欠界定率均更低,且除高水平基因流外,对各类潜在混杂因素总体表现稳健。 我们的研究发现,单阈值GMYC模型与本研究在PTP中确定的最优策略,在无基因流且推定物种数量多于单个类群的场景下通常表现良好;但当推定物种数量较少时,PTP的性能优于GMYC。相较于BPP,这两种方法对基因流及潜在混杂因素的影响更为敏感。 针对熊类与蜂类的案例研究进一步验证了本模拟研究的部分结论,并揭示了为分子物种界定设置合理起始点的重要性。本研究结果明确了影响分子物种界定方法性能的关键因素,为在整合分类学框架下应用此类方法提供了潜在的应用价值。

创建时间:
2018-02-13
二维码
社区交流群
二维码
科研交流群
商业服务