Data from: The relationship between dN/dS and scaled selection coefficients
收藏资源简介:
Numerous computational methods exist to assess the mode and strength of natural selection in protein-coding sequences, yet how distinct methods relate to one another remains largely unknown. Here, we elucidate the relationship between two widely-used phylogenetic modeling frameworks: dN/dS models and mutation-selection (MutSel) models. We derive a mathematical relationship between dN/dS and scaled selection coefficients, the focal parameters of MutSel models, and use this relationship to gain deeper insight into the behaviors, limitations, and applicabilities of these two modeling frameworks. We prove that, if all synonymous changes are neutral, standard MutSel models correspond to dN/dS < 1. However, if synonymous codons differ in fitness, dN/dS can take on arbitrarily high values even if all selection is purifying. Thus, the MutSel modeling framework cannot necessarily accommodate positive, diversifying selection, while dN/dS cannot distinguish between purifying selection on synonymous codons and positive selection on amino acids. We further propose a new benchmarking strategy of dN/dS inferences against MutSel simulations and demonstrate that the widely-used Goldman-Yang-style dN/dS models yield substantially biased dN/dS estimates on realistic sequence data. By contrast, the less frequently used Muse-Gaut-style models display much less bias. Strikingly, the least-biased and most-precise dN/dS estimates are never found in the models with the best fit to the data, measured through both AIC and BIC scores. Thus, selecting models based on goodness-of-fit criteria can yield poor parameter estimates if the models considered do not precisely correspond to the underlying mechanism that generated the data. In conclusion, establishing mathematical links among modeling frameworks represents a novel, powerful strategy to pinpoint previously unrecognized model limitations and strengths.
目前已有诸多计算方法可用于评估蛋白编码序列中自然选择的模式与强度,但不同方法间的关联机制仍大多尚未明确。本研究阐明了两类广泛使用的系统发育建模框架——dN/dS模型与突变选择(mutation-selection, MutSel)模型之间的关系。我们推导了dN/dS与突变选择模型的核心参数缩放选择系数(scaled selection coefficients)之间的数学关联,并借助该关联深入剖析了这两类建模框架的运行特性、局限性与适用范围。我们证明,若所有同义突变均为中性,则标准的突变选择模型对应的dN/dS值小于1。然而,若同义密码子间存在适合度差异,即便所有选择均为净化选择,dN/dS值也可达到任意高的水平。由此可见,突变选择建模框架未必能适配正向选择与歧化选择,而dN/dS则无法区分针对同义密码子的净化选择与针对氨基酸的正向选择。我们进一步提出了一种基于突变选择模拟的dN/dS推断基准测试新策略,并证实,在真实序列数据上,广泛使用的戈德曼-杨(Goldman-Yang)型dN/dS模型会生成存在显著偏倚的dN/dS估计值。相较而言,使用频次较低的缪斯-戈特(Muse-Gaut)型模型的偏倚则要小得多。值得注意的是,通过赤池信息准则(Akaike Information Criterion, AIC)与贝叶斯信息准则(Bayesian Information Criterion, BIC)衡量,拟合效果最优的模型,从未给出偏倚最低且精度最高的dN/dS估计值。因此,若所考量的模型未能精准匹配生成数据的底层机制,仅以拟合优度标准选择模型,可能会得到精度不佳的参数估计结果。综上,建立不同建模框架间的数学关联,是一种全新且有效的策略,可用于精准识别此前未被发现的模型局限性与优势。



