nRCFV: A sequence, taxon and character state-normalised metric for the pre-reconstruction evaluation of compositional heterogeneity
收藏资源简介:
Motivation Compositional heterogeneity â when the proportions of nucleotides and amino acids are not broadly similar across the dataset â is a cause of a great number of phylogenetic artefacts. Whilst a variety of methods can identify it post-hoc, few metrics exist to quantify compositional heterogeneity prior to the computationally intensive task of phylogenetic tree reconstruction. Here we assess the efficacy of one such existing, widely used, metric: Relative Composition Frequency Variability (RCFV), using both real and simulated data. Results Our results show that RCFV can be biased by sequence length, the number of taxa, and the number of possible character states within the dataset. However, we also find that missing data does not appear to have an appreciable value on RCFV. We discuss the theory behind this and the consequences of this for the future of the usage of the RCFV value and propose a new metric, nRCFV, which accounts for these biases. Alongside this, we present a new s..., ,
研究动因 组成异质性(Compositional heterogeneity)指数据集内核苷酸与氨基酸的占比整体并不均一,这是引发大量系统发育伪影(phylogenetic artefacts)的重要诱因之一。尽管已有多种方法可在事后识别该异质性,但在开展计算成本高昂的系统发育树重建工作之前,能够量化组成异质性的指标却寥寥无几。本研究基于真实与模拟数据集,评估了现有广泛使用的该类指标之一——相对组成频率变异性(Relative Composition Frequency Variability,简称RCFV)的有效性。 研究结果 本研究结果显示,序列长度、分类单元(taxa)数量以及数据集内可能存在的性状状态数目,均会对RCFV产生偏倚。但本研究同时发现,缺失数据似乎并不会对RCFV产生显著影响。本研究就此现象背后的理论机制以及该结论对RCFV应用前景的影响展开了讨论,并提出了一种可校正上述偏倚的新型指标nRCFV。与此同时,本研究还提出了一项全新的……



