Data from: The impact of the tree prior on molecular dating of data sets containing a mixture of inter- and intraspecies sampling
收藏资源简介:
In Bayesian phylogenetic analyses of genetic data, prior probability distributions need to be specified for the model parameters, including the tree. When Bayesian methods are used for molecular dating, available tree priors include those designed for species-level data, such as the pure-birth and birth-death priors, and coalescent-based priors designed for population-level data. However, molecular dating methods are frequently applied to data sets that include multiple individuals across multiple species. Such data sets violate the assumptions of both the speciation and coalescent-based tree priors, making it unclear which should be chosen and whether this choice can affect the estimation of node times. To investigate this problem, we used a simulation approach to produce data sets with different proportions of within- and between-species sampling under the multispecies coalescent model. These data sets were then analysed under pure-birth, birth-death, constant-size coalescent, and skyline coalescent tree priors. We also explored the ability of Bayesian model testing to select the best-performing priors. We confirmed the applicability of our results to empirical data sets from cetaceans, phocids, and coregonid whitefish. Estimates of node times were generally robust to the choice of tree prior, but some combinations of tree priors and sampling schemes led to large differences in the age estimates. In particular, the pure-birth tree prior frequently led to inaccurate estimates for data sets containing a mixture of inter- and intraspecific sampling, whereas the birth-death and skyline coalescent priors produced stable results across all scenarios. Model testing provided an adequate means of rejecting inappropriate tree priors. Our results suggest that tree priors do not strongly affect Bayesian molecular dating results in most cases, even when severely misspecified. However, the choice of tree prior can be significant for the accuracy of dating results in the case of data sets with mixed inter- and intraspecies sampling.
在遗传数据的贝叶斯系统发育分析中,需为包括系统发育树在内的模型参数指定先验概率分布。当使用贝叶斯方法开展分子定年时,可用的树先验(tree prior)包括针对物种级数据设计的先验(如纯生先验与生灭先验),以及针对种群级数据设计的基于溯祖的先验(coalescent-based prior)。然而,分子定年方法常被应用于包含跨多个物种的多个个体的数据集。此类数据集违背了物种形成类与基于溯祖的两类树先验的假设,因此难以确定应选用何种先验,以及该选择是否会影响节点时间的估计。为探究该问题,本研究采用模拟方法,在多物种溯祖模型(multispecies coalescent model)下生成了具有不同种内与种间采样比例的数据集。随后,分别采用纯生先验、生灭先验、恒定种群大小溯祖先验与溯祖天际线先验对这些数据集开展分析。本研究还探究了贝叶斯模型检验筛选表现最优先验的能力。本研究证实了研究结果适用于鲸类、海豹科动物以及白鲑属白鱼的实证数据集。节点时间估计总体上对树先验的选择具有稳健性,但部分树先验与采样方案的组合会导致年代估计出现显著差异。具体而言,纯生树先验常会导致包含种间与种内混合采样的数据集的估计结果出现偏差,而生灭先验与溯祖天际线先验在所有情景下均能生成稳定的结果。贝叶斯模型检验可有效识别并排除不合适的树先验。本研究结果表明,在大多数情况下,即使树先验存在严重误设,其对贝叶斯分子定年结果的影响也并不显著。但对于包含种间与种内混合采样的数据集而言,树先验的选择对定年结果的准确性具有显著影响。




