Integrated Sequence-Structure Motifs Suffice to Identify microRNA Precursors
收藏资源简介:
BackgroundUpwards of 1200 miRNA loci have hitherto been annotated in the human genome. The specific features defining a miRNA precursor and deciding its recognition and subsequent processing are not yet exhaustively described and miRNA loci can thus not be computationally identified with sufficient confidence. ResultsWe rendered pre-miRNA and non-pre-miRNA hairpins as strings of integrated sequence-structure information, and used the software Teiresias to identify sequence-structure motifs (ss-motifs) of variable length in these data sets. Using only ss-motifs as features in a Support Vector Machine (SVM) algorithm for pre-miRNA identification achieved 99.2% specificity and 97.6% sensitivity on a human test data set, which is comparable to previously published algorithms employing combinations of sequence-structure and additional features. Further analysis of the ss-motif information contents revealed strongly significant deviations from those of the respective training sets, revealing important potential clues as to how the sequence and structural information of RNA hairpins are utilized by the miRNA processing apparatus. ConclusionIntegrated sequence-structure motifs of variable length apparently capture nearly all information required to distinguish miRNA precursors from other stem-loop structures.
背景 迄今已在人类基因组中注释了超过1200个微小RNA(miRNA)基因座。目前尚未详尽阐明定义微小RNA前体、决定其识别与后续加工过程的特异性特征,因此尚无法通过计算方法以足够置信度精准识别微小RNA基因座。 结果 我们将前体微小RNA(pre-miRNA)与非前体微小RNA发夹结构编码为整合序列与结构信息的字符串,并借助Teiresias软件在上述数据集中识别可变长度的序列-结构基序(ss-motifs)。仅将序列-结构基序作为特征输入支持向量机(Support Vector Machine, SVM)算法以识别前体微小RNA,在人类测试数据集上取得了99.2%的特异性与97.6%的灵敏度,该性能与此前已发表的、结合序列-结构特征与额外特征的算法相当。对序列-结构基序的信息含量展开进一步分析后发现,其与对应训练集的信息含量存在极显著差异,这为解析RNA发夹结构的序列与结构信息如何被微小RNA加工机制所利用提供了关键潜在线索。 结论 可变长度的整合序列-结构基序显然几乎涵盖了区分微小RNA前体与其他茎环结构所需的全部信息。



