scnRCA: A Novel Method to Detect Consistent Patterns of Translational Selection in Mutationally-Biased Genomes
收藏资源简介:
Codon usage bias (CUB) results from the complex interplay between translational selection and mutational biases. Current methods for CUB analysis apply heuristics to integrate both components, limiting the depth and scope of CUB analysis as a technique to probe into the evolution and optimization of protein-coding genes. Here we introduce a self-consistent CUB index (scnRCA) that incorporates implicit correction for mutational biases, facilitating exploration of the translational selection component of CUB. We validate this technique using gene expression data and we apply it to a detailed analysis of CUB in the Pseudomonadales. Our results illustrate how the selective enrichment of specific codons among highly expressed genes is preserved in the context of genome-wide shifts in codon frequencies, and how the balance between mutational and translational biases leads to varying definitions of codon optimality. We extend this analysis to other moderate and fast growing bacteria and we provide unified support for the hypothesis that C- and A-ending codons of two-box amino acids, and the U-ending codons of four-box amino acids, are systematically enriched among highly expressed genes across bacteria. The use of an unbiased estimator of CUB allows us to report for the first time that the signature of translational selection is strongly conserved in the Pseudomonadales in spite of drastic changes in genome composition, and extends well beyond the core set of highly optimized genes in each genome. We generalize these results to other moderate and fast growing bacteria, hinting at selection for a universal pattern of gene expression that is conserved and detectable in conserved patterns of codon usage bias.
密码子使用偏好(Codon usage bias, CUB)源于翻译选择与突变偏倚间的复杂相互作用。当前用于CUB分析的方法多通过启发式策略整合两类因素,限制了CUB作为探究蛋白质编码基因演化与优化工具的分析深度与应用范围。本研究提出一种自洽的CUB指数(scnRCA),该指数可对突变偏倚进行隐性校正,从而便于解析CUB中的翻译选择成分。我们利用基因表达数据对该方法进行了验证,并将其应用于假单胞菌目(Pseudomonadales)物种的CUB详细分析。研究结果阐明了两点:一是在全基因组密码子频率发生整体偏移的背景下,高表达基因中特定密码子的选择性富集如何得以保留;二是突变偏倚与翻译选择间的平衡如何催生了密码子最优性的不同定义。我们将该分析拓展至其他中等生长速率与快速生长的细菌类群,并为以下假说提供了统一的支持证据:在各类细菌的高表达基因中,双密码子家族氨基酸的以C或A结尾的密码子,以及四密码子家族氨基酸的以U结尾的密码子,均呈现系统性富集现象。借助无偏的CUB估计量,我们首次报道:尽管假单胞菌目的基因组组成发生了剧烈变化,但其翻译选择的特征仍高度保守,且该特征的覆盖范围远超出各基因组中高度优化的核心基因集合。我们将上述结果推广至其他中等生长速率与快速生长的细菌类群,这暗示自然界存在一种受选择作用的通用基因表达模式,该模式可通过密码子使用偏好的保守特征得以保留并被检测。



