Designing gene libraries from protein profiles for combinatorial protein experiments
收藏资源简介:
Protein combinatorial libraries provide new ways to probe the determinants of folding and to discover novel proteins. Such libraries are often constructed by expressing an ensemble of partially random gene sequences. Given the intractably large number of possible sequences, some limitation on diversity must be imposed. A non-uniform distribution of nucleotides can be used to reduce the number of possible sequences and encode peptide sequences having a predetermined set of amino acid probabilities at each residue position, i.e., the amino acid sequence profile. Such profiles can be determined by inspection, multiple sequence alignment or physically-based computational methods. Here we present a computational method that takes as input a desired sequence profile and calculates the individual nucleotide probabilities among partially random genes. The calculated gene library can be readily used in the context of standard DNA synthesis to generate a protein library with essentially the desired profile. The fidelity between the desired profile and the calculated one coded by these partially random genes is quantitatively evaluated using the linear correlation coefficient and a relative entropy, each of which provides a measure of profile agreement at each position of the sequence. On average, this method of identifying such codon frequencies performs as well or better than other methods with regard to fidelity to the original profile. Importantly, the method presented here provides much better yields of complete sequences that do not contain stop codons, a feature that is particularly important when all or large fractions of a gene are subject to combinatorial mutation.
蛋白质组合文库 (Protein combinatorial libraries) 为探究蛋白质折叠的决定因素以及发现新型蛋白质提供了全新途径。此类文库通常通过表达一组部分随机化的基因序列来构建。鉴于潜在序列的数量多得难以遍历,必须对文库的多样性施加一定限制。采用非均匀分布的核苷酸可减少潜在序列的数量,并编码出在每个残基位置具备预设氨基酸概率分布的肽序列,即氨基酸序列谱 (amino acid sequence profile)。此类序列谱可通过人工检视、多序列比对或是基于物理原理的计算方法确定。本研究提出一种计算方法,以目标序列谱作为输入,计算部分随机化基因中各核苷酸的概率分布。经计算得到的基因文库可直接应用于标准DNA合成流程,以生成具备目标序列谱的蛋白质文库。我们采用线性相关系数 (linear correlation coefficient) 与相对熵 (relative entropy),对目标序列谱与该部分随机化基因编码得到的计算序列谱之间的保真度进行定量评估,二者均可衡量序列各位置上的谱图匹配程度。平均而言,该确定密码子频率的方法在还原原始序列谱的保真度方面,表现不逊于甚至优于其他同类方法。尤为关键的是,本方法所获得的不含终止密码子 (stop codons) 的完整序列产率显著更高,这一特性在基因全部或大部分区域发生组合式突变的场景下尤为重要。



