Exploration of Noncoding Sequences in Metagenomes
收藏资源简介:
Environment-dependent genomic features have been defined for different metagenomes, whose genes and their associated processes are related to specific environments. Identification of ORFs and their functional categories are the most common methods for association between functional and environmental features. However, this analysis based on finding ORFs misses noncoding sequences and, therefore, some metagenome regulatory or structural information could be discarded. In this work we analyzed 23 whole metagenomes, including coding and noncoding sequences using the following sequence patterns: (G+C) content, Codon Usage (Cd), Trinucleotide Usage (Tn), and functional assignments for ORF prediction. Herein, we present evidence of a high proportion of noncoding sequences discarded in common similarity-based methods in metagenomics, and the kind of relevant information present in those. We found a high density of trinucleotide repeat sequences (TRS) in noncoding sequences, with a regulatory and adaptive function for metagenome communities. We present associations between trinucleotide values and gene function, where metagenome clustering correlate with microorganism adaptations and kinds of metagenomes. We propose here that noncoding sequences have relevant information to describe metagenomes that could be considered in a whole metagenome analysis in order to improve their organization, classification protocols, and their relation with the environment.
针对不同宏基因组,学界已定义了依赖环境的基因组特征,这类特征的基因及其关联过程均与特定环境紧密相关。开放阅读框(Open Reading Frame, ORF)的识别及其功能分类,是关联功能特征与环境特征的最常用方法。然而,这类基于开放阅读框搜寻的分析方法会遗漏非编码序列,进而可能丢失部分宏基因组的调控或结构信息。本研究共分析了23套完整宏基因组,针对编码序列与非编码序列,采用以下四类序列模式展开分析:(G+C)含量、密码子使用情况(Codon Usage, Cd)、三核苷酸使用情况(Trinucleotide Usage, Tn),以及用于开放阅读框预测的功能注释。本研究证实,宏基因组学领域常用的基于相似性的分析方法会丢弃大量非编码序列,并揭示了这类序列中蕴含的核心信息类型。我们发现,非编码序列中存在高密度的三核苷酸重复序列(Trinucleotide Repeat Sequence, TRS),这类序列对宏基因组群落兼具调控与适应功能。此外,我们还揭示了三核苷酸特征值与基因功能之间的关联:宏基因组的聚类结果与微生物适应性及宏基因组类型高度相关。本研究提出,非编码序列蕴含可用于描述宏基因组的关键信息,若将其纳入完整宏基因组分析流程,可优化宏基因组的组织形式、分类方案及其与环境的关联关系。




