遇见数据集

Evaluation of Sequence Features from Intrinsically Disordered Regions for the Estimation of Protein Function

收藏
Figshare2016-01-18 更新2026-04-29 收录
官方服务:

资源简介:

With the exponential increase in the number of sequenced organisms, automated annotation of proteins is becoming increasingly important. Intrinsically disordered regions are known to play a significant role in protein function. Despite their abundance, especially in eukaryotes, they are rarely used to inform function prediction systems. In this study, we extracted seven sequence features in intrinsically disordered regions and developed a scheme to use them to predict Gene Ontology Slim terms associated with proteins. We evaluated the function prediction performance of each feature. Our results indicate that the residue composition based features have the highest precision while bigram probabilities, based on sequence profiles of intrinsically disordered regions obtained from PSIBlast, have the highest recall. Amino acid bigrams and features based on secondary structure show an intermediate level of precision and recall. Almost all features showed a high prediction performance for GO Slim terms related to extracellular matrix, nucleus, RNA and DNA binding. However, feature performance varied significantly for different GO Slim terms emphasizing the need for a unique classifier optimized for the prediction of each functional term. These findings provide a first comprehensive and quantitative evaluation of sequence features in intrinsically disordered regions and will help in the development of a more informative protein function predictor.

随着测序生物的数量呈指数级增长,蛋白质的自动化注释正变得愈发重要。固有无序区域(intrinsically disordered regions)已被证实对蛋白质功能发挥着关键作用。尽管这类区域广泛存在,尤其在真核生物中,但极少被用于为功能预测系统提供参考依据。本研究提取了固有无序区域的7种序列特征,并构建了一套利用这些特征预测与蛋白质相关的基因本体精简版(Gene Ontology Slim,以下简称GO Slim)术语的方案。我们评估了每种特征的功能预测性能。研究结果显示,基于残基组成的特征拥有最高的精确率,而基于PSIBlast获取的固有无序区域序列谱所计算得到的二元组(bigram)概率特征则拥有最高的召回率。氨基酸二元组与基于二级结构的特征则展现出中等水平的精确率与召回率。几乎所有特征在预测与细胞外基质、细胞核、RNA及DNA结合相关的GO Slim术语时,均表现出优异的预测性能。但针对不同的GO Slim术语,特征的性能差异显著,这凸显出为每个功能术语的预测定制专属优化分类器的必要性。本研究首次对固有无序区域的序列特征开展了全面且定量的评估,将助力开发出信息更为丰富的蛋白质功能预测工具。

创建时间:
2016-01-18
二维码
社区交流群
二维码
科研交流群
商业服务