Data from: Detecting adaptive evolution in phylogenetic comparative analysis using the Ornstein-Uhlenbeck model
收藏资源简介:
Phylogenetic comparative analysis is an approach to inferring evolutionary process from a combination of phylogenetic and phenotypic data. The last few years have seen increasingly sophisticated models employed in the evaluation of more and more detailed evolutionary hypotheses, including adaptive hypotheses with multiple selective optima and hypotheses with rate variation within and across lineages. The statistical performance of these sophisticated models has received relatively little systematic attention, however. We conducted an extensive simulation study to quantify the statistical properties of a class of models toward the simpler end of the spectrum that model phenotypic evolution using Ornstein–Uhlenbeck processes. We focused on identifying where, how, and why these methods break down so that users can apply them with greater understanding of their strengths and weaknesses. Our analysis identifies three key determinants of performance: a discriminability ratio, a signal-to-noise ratio, and the number of taxa sampled. Interestingly, we find that model-selection power can be high even in regions that were previously thought to be difficult, such as when tree size is small. On the other hand, we find that model parameters are in many circumstances difficult to estimate accurately, indicating a relative paucity of information in the data relative to these parameters. Nevertheless, we note that accurate model selection is often possible when parameters are only weakly identified. Our results have implications for more sophisticated methods inasmuch as the latter are generalizations of the case we study.
系统发育比较分析(Phylogenetic Comparative Analysis)是一种结合系统发育数据与表型数据推断进化过程的研究方法。近十余年来,研究者在评估愈发精细的进化假说时,所采用的模型日趋复杂——这类假说既包含多选择最适点的适应性假说,也涵盖支系内部及跨支系存在速率变异的假说。然而,此类复杂模型的统计性能尚未得到充分的系统性研究。 为此,我们开展了一项大规模模拟研究,旨在量化一类偏向简化的模型的统计特性:这类模型采用奥恩斯坦-乌伦贝克过程(Ornstein–Uhlenbeck Process)模拟表型进化。我们的研究重点在于厘清这些方法失效的场景、机制与原因,以便使用者能够更清晰地理解其优劣并合理应用。 本研究明确了影响模型性能的三个关键因素:可区分度比、信噪比与采样类群数。有趣的是,我们发现即便在过往认为难度较高的场景中,模型选择效能仍可处于较高水平——例如当系统发育树规模较小时。 另一方面,在多数场景下,模型参数难以被精准估计,这表明相较于这些参数,现有数据所能提供的信息相对匮乏。尽管如此,我们仍注意到:即便参数仅能被弱识别,准确的模型选择仍往往可行。 鉴于更复杂的模型均是我们所研究案例的推广形式,本研究结果对这类高级方法同样具有参考价值。



