NullSeq: A Tool for Generating Random Coding Sequences with Desired Amino Acid and GC Contents
收藏资源简介:
The existence of over- and under-represented sequence motifs in genomes provides evidence of selective evolutionary pressures on biological mechanisms such as transcription, translation, ligand-substrate binding, and host immunity. In order to accurately identify motifs and other genome-scale patterns of interest, it is essential to be able to generate accurate null models that are appropriate for the sequences under study. While many tools have been developed to create random nucleotide sequences, protein coding sequences are subject to a unique set of constraints that complicates the process of generating appropriate null models. There are currently no tools available that allow users to create random coding sequences with specified amino acid composition and GC content for the purpose of hypothesis testing. Using the principle of maximum entropy, we developed a method that generates unbiased random sequences with pre-specified amino acid and GC content, which we have developed into a python package. Our method is the simplest way to obtain maximally unbiased random sequences that are subject to GC usage and primary amino acid sequence constraints. Furthermore, this approach can easily be expanded to create unbiased random sequences that incorporate more complicated constraints such as individual nucleotide usage or even di-nucleotide frequencies. The ability to generate correctly specified null models will allow researchers to accurately identify sequence motifs which will lead to a better understanding of biological processes as well as more effective engineering of biological systems.
基因组中存在的过丰度与低丰度序列基序(sequence motif),为转录、翻译、配体-底物结合以及宿主免疫等生物学机制受到的选择性进化压力提供了佐证。为精准识别目标基序及其他全基因组尺度的感兴趣模式,生成适配研究序列的准确零模型(null model)是至关重要的前提。尽管现有诸多工具可用于生成随机核苷酸序列,但蛋白质编码序列受到一系列独特约束,这使得适配性零模型的构建过程变得复杂。目前尚无工具可支持用户为假设检验目的,生成具备指定氨基酸组成与GC含量(GC content)的随机编码序列。基于最大熵原理(maximum entropy principle),我们开发了一种可生成具备预设氨基酸组成与GC含量的无偏随机序列的方法,并将其封装为Python包。本方法是获取受限于GC使用偏好与一级氨基酸序列约束的最大无偏随机序列的最简途径。此外,该方法可轻松拓展至构建包含更复杂约束的无偏随机序列,例如单个核苷酸使用偏好乃至二核苷酸频率(di-nucleotide frequency)。生成精准适配的零模型的能力,将助力研究人员精准识别序列基序,进而深化对生物学过程的理解,并实现生物系统更为高效的工程化改造。



