Embedding Permanent Watermarks in Synthetic Genes
收藏资源简介:
As synthetic biology advances, labeling of genes or organisms, like other high-value products, will become important not only to pinpoint their identity, origin, or spread, but also for intellectual property, classification, bio-security or legal reasons. Ideally information should be inseparably interlaced into expressed genes. We describe a method for embedding messages within open reading frames of synthetic genes by adapting steganographic algorithms typically used for watermarking digital media files. Text messages are first translated into a binary string, and then represented in the reading frame by synonymous codon choice. To aim for good expression of the labeled gene in its host as well as retain a high degree of codon assignment flexibility for gene optimization, codon usage tables of the target organism are taken into account. Preferably amino acids with 4 or 6 synonymous codons are used to comprise binary digits. Several different messages were embedded into open reading frames of T7 RNA polymerase, GFP, human EMG1 and HIV gag, variously optimized for bacterial, yeast, mammalian or plant expression, without affecting their protein expression or function. We also introduced Vigenère polyalphabetic substitution to cipher text messages, and developed an identifier as a key to deciphering codon usage ranking stored for a specific organism within a sequence of 35 nucleotides.
随着合成生物学(synthetic biology)的发展,对基因或生物体进行标记,如同其他高价值产品一般,其重要性不仅体现在可精准确定其身份、来源与传播范围,同时也可满足知识产权管理、物种分类、生物安全保障及法律合规等多方面的需求。理想情况下,信息应与表达基因不可分割地整合在一起。我们提出了一种将信息嵌入合成基因开放阅读框(open reading frame, ORF)的方法:该方法通过适配通常用于数字媒体文件水印的隐写算法(steganographic algorithm)实现。首先将文本信息转换为二进制字符串,再通过选择同义密码子(synonymous codon)将其编码至阅读框内。为确保标记基因在宿主细胞中实现良好表达,同时保留基因优化所需的密码子分配灵活性,本方法纳入了目标生物体的密码子使用偏好表(codon usage table)。优选选取拥有4个或6个同义密码子的氨基酸,以对应二进制数位。我们将多条不同的信息嵌入了T7 RNA聚合酶(T7 RNA polymerase)、绿色荧光蛋白(green fluorescent protein, GFP)、人类EMG1以及HIV gag的开放阅读框中:这些基因均针对细菌、酵母、哺乳动物或植物表达系统进行了优化,且未影响其蛋白质表达与功能。此外,我们还引入了维吉尼亚多字母替换密码(Vigenère polyalphabetic substitution)对文本信息进行加密,并开发了一种标识符作为解密密钥:该密钥可对存储于35个核苷酸(nucleotide)序列中的特定生物体密码子使用优先级信息进行解密。



