Selective Constraints on Amino Acids Estimated by a Mechanistic Codon Substitution Model with Multiple Nucleotide Changes
收藏资源简介:
BackgroundEmpirical substitution matrices represent the average tendencies of substitutions over various protein families by sacrificing gene-level resolution. We develop a codon-based model, in which mutational tendencies of codon, a genetic code, and the strength of selective constraints against amino acid replacements can be tailored to a given gene. First, selective constraints averaged over proteins are estimated by maximizing the likelihood of each 1-PAM matrix of empirical amino acid (JTT, WAG, and LG) and codon (KHG) substitution matrices. Then, selective constraints specific to given proteins are approximated as a linear function of those estimated from the empirical substitution matrices. ResultsAkaike information criterion (AIC) values indicate that a model allowing multiple nucleotide changes fits the empirical substitution matrices significantly better. Also, the ML estimates of transition-transversion bias obtained from these empirical matrices are not so large as previously estimated. The selective constraints are characteristic of proteins rather than species. However, their relative strengths among amino acid pairs can be approximated not to depend very much on protein families but amino acid pairs, because the present model, in which selective constraints are approximated to be a linear function of those estimated from the JTT/WAG/LG/KHG matrices, can provide a good fit to other empirical substitution matrices including cpREV for chloroplast proteins and mtREV for vertebrate mitochondrial proteins. Conclusions/SignificanceThe present codon-based model with the ML estimates of selective constraints and with adjustable mutation rates of nucleotide would be useful as a simple substitution model in ML and Bayesian inferences of molecular phylogenetic trees, and enables us to obtain biologically meaningful information at both nucleotide and amino acid levels from codon and protein sequences.
研究背景 经验替换矩阵(empirical substitution matrices)通过牺牲基因层面的分辨率,来表征各类蛋白质家族中氨基酸替换的平均趋势。我们开发了一种基于密码子的模型(codon-based model),该模型可针对特定基因,定制密码子的突变倾向、遗传密码,以及针对氨基酸替换的选择约束强度。首先,我们通过最大化经验氨基酸(JTT、WAG与LG)及密码子(KHG)替换矩阵的各1-PAM矩阵的似然值,估算得到全蛋白质的平均选择约束。随后,将特定靶标蛋白质的选择约束近似为基于经验替换矩阵估算所得约束值的线性函数。 研究结果 赤池信息准则(Akaike Information Criterion, AIC)值表明,允许多核苷酸替换的模型对经验替换矩阵的拟合效果显著更优。此外,从这些经验矩阵中得到的转换-颠换偏倚的最大似然(Maximum Likelihood, ML)估计值,并未如既往研究所估算的那般显著。选择约束具有蛋白质特异性而非物种特异性。不过,氨基酸对间的相对约束强度似乎并不太依赖于蛋白质家族,而是更多取决于氨基酸对本身——这是由于本模型将选择约束近似为基于JTT/WAG/LG/KHG矩阵估算的约束值的线性函数,能够很好地拟合其他经验替换矩阵,包括针对叶绿体蛋白质的cpREV与针对脊椎动物线粒体蛋白质的mtREV。 结论与意义 本研究提出的结合选择约束最大似然估计值与可调节核苷酸突变率的基于密码子的模型,可作为分子系统发育树最大似然与贝叶斯推断中的简便替换模型,同时能够从密码子序列与蛋白质序列中获取兼具生物学意义的核苷酸与氨基酸层面信息。



