遇见数据集

Data from: Detecting adaptive evolution in phylogenetic comparative analysis using the Ornstein-Uhlenbeck model

收藏
DataONE2015-07-07 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

Phylogenetic comparative analysis is an approach to inferring evolutionary process from a combination of phylogenetic and phenotypic data. The last few years have seen increasingly sophisticated models employed in the evaluation of more and more detailed evolutionary hypotheses, including adaptive hypotheses with multiple selective optima and hypotheses with rate variation within and across lineages. The statistical performance of these sophisticated models has received relatively little systematic attention, however. We conducted an extensive simulation study to quantify the statistical properties of a class of models toward the simpler end of the spectrum that model phenotypic evolution using Ornstein–Uhlenbeck processes. We focused on identifying where, how, and why these methods break down so that users can apply them with greater understanding of their strengths and weaknesses. Our analysis identifies three key determinants of performance: a discriminability ratio, a signal-to-noise ratio, and the number of taxa sampled. Interestingly, we find that model-selection power can be high even in regions that were previously thought to be difficult, such as when tree size is small. On the other hand, we find that model parameters are in many circumstances difficult to estimate accurately, indicating a relative paucity of information in the data relative to these parameters. Nevertheless, we note that accurate model selection is often possible when parameters are only weakly identified. Our results have implications for more sophisticated methods inasmuch as the latter are generalizations of the case we study.

系统发育比较分析(Phylogenetic comparative analysis)是一种结合系统发育数据与表型数据推断进化过程的研究范式。近年间,演化假说检验领域涌现出愈发精密复杂的模型,可用于检验多重选择最适值的适应性假说,以及类群内部与跨类群的速率变异假说等诸多精细演化假设。然而,此类高阶模型的统计性能却鲜有系统性的研究关注。本研究开展了一项大规模模拟研究,旨在量化一类偏向简化的模型的统计特性——这类模型借助奥恩斯坦-乌伦贝克过程(Ornstein–Uhlenbeck processes)模拟表型演化过程。我们的核心目标在于探明这些方法失效的场景、内在机制与根本原因,从而帮助使用者更清晰地认知模型的优势与局限,进而更合理地应用此类方法。本分析明确了影响模型性能的三项关键决定因素:可判别比、信噪比,以及采样类群的数量。值得注意的是,我们发现即便在以往被认为极具挑战性的场景中,模型选择效能仍可维持在较高水平——例如当系统发育树的规模较小时。另一方面,我们发现在诸多实际场景中,模型参数难以被精准估计,这反映出相较于待估参数而言,数据集所能提供的信息相对不足。尽管如此,我们仍观察到,即便参数仅被弱识别,精准的模型选择仍往往具备可行性。我们的研究结果对更精密复杂的演化模型同样具有参考意义,因为后者正是本研究所分析场景的推广形式。

创建时间:
2015-07-07
二维码
社区交流群
二维码
科研交流群
商业服务