遇见数据集

Data from: A new method for handling missing species in diversification analysis applicable to randomly or non-randomly sampled phylogenies

收藏
DataONE2012-01-06 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

Chronograms from molecular dating are increasingly being used to infer rates of diversification and their change over time. A major limitation in such analyses is incomplete species sampling that moreover is usually non-random. While the widely used γ statistic with the MCCR test or the birth-death likelihood analysis with the ∆AICrc test statistic are appropriate for comparing the fit of different diversification models in phylogenies with random species sampling, no objective, automated method has been developed for fitting diversification models to non-randomly sampled phylogenies. Here we introduce a novel approach, CorSiM, which involves simulating missing splits under a constant-rate birth-death model and allows the user to specify whether species sampling in the phylogeny being analyzed is random or non-random. The completed trees can be used in subsequent model-fitting analyses. This is fundamentally different from previous diversification rate estimation methods, which were based on null distributions derived from the incomplete trees. CorSiM is automated in an R package and can easily be applied to large data sets. We illustrate the approach in two Araceae clades, one with a random species sampling of 52% and one with a non-random sampling of 55%. In the latter clade, the CorSiM approach detects and quantifies an increase in diversification rate while classic approaches prefer a constant rate model, whereas in the former clade, results do not differ among methods (as indeed expected since the classic approaches are valid only for randomly sampled phylogenies). The CorSiM method greatly reduces the type I error in diversification analysis, but type II error remains a methodological problem.

基于分子定年的时间树(chronograms from molecular dating)正日益被用于推断物种多样化速率及其随时间的变化趋势。此类分析的核心局限之一为物种采样不完整,且采样往往并非随机。尽管结合MCCR检验(MCCR test)的γ统计量(γ statistic),或是结合∆AICrc检验统计量(∆AICrc test statistic)的生灭似然分析,均可适用于在随机物种采样的系统发育树(phylogeny)中比较不同多样化模型的拟合效果,但目前尚未开发出可将多样化模型拟合至非随机采样系统发育树的客观自动化方法。 在此我们提出一种全新方法CorSiM,该方法基于恒定速率生灭模型模拟缺失的分支事件,并允许使用者指定待分析系统发育树中的物种采样是否为随机。补全后的系统发育树可直接应用于后续的模型拟合分析,这与此前的多样化速率估算方法有着本质区别:后者的空分布均基于未完成采样的系统发育树推导而来。 CorSiM已封装为R语言工具包并实现全流程自动化,可便捷应用于大型数据集。我们以天南星科(Araceae)的两个演化支为例展示该方法:其中一个演化支的随机物种采样率为52%,另一个演化支的非随机物种采样率为55%。在后者演化支中,CorSiM方法能够检测并量化多样化速率的提升趋势,而经典方法则倾向于恒定速率模型;而在前者演化支中,不同方法得到的结果并无差异,这完全符合预期——因为经典方法仅适用于随机采样的系统发育树。 CorSiM方法可大幅降低多样化分析中的一类错误(type I error),但二类错误(type II error)仍是一个亟待解决的方法论难题。

创建时间:
2012-01-06
二维码
社区交流群
二维码
科研交流群
商业服务