遇见数据集

Exploiting Publicly Available Biological and Biochemical Information for the Discovery of Novel Short Linear Motifs

收藏
Figshare2016-01-18 更新2026-04-29 收录
官方服务:

资源简介:

The function of proteins is often mediated by short linear segments of their amino acid sequence, called Short Linear Motifs or SLiMs, the identification of which can provide important information about a protein function. However, the short length of the motifs and their variable degree of conservation makes their identification hard since it is difficult to correctly estimate the statistical significance of their occurrence. Consequently, only a small fraction of them have been discovered so far. We describe here an approach for the discovery of SLiMs based on their occurrence in evolutionarily unrelated proteins belonging to the same biological, signalling or metabolic pathway and give specific examples of its effectiveness in both rediscovering known motifs and in discovering novel ones. An automatic implementation of the procedure, available for download, allows significant motifs to be identified, automatically annotated with functional, evolutionary and structural information and organized in a database that can be inspected and queried. An instance of the database populated with pre-computed data on seven organisms is accessible through a publicly available server and we believe it constitutes by itself a useful resource for the life sciences (http://www.biocomputing.it/modipath).

蛋白质的功能通常由其氨基酸序列中的短线性片段介导,这类片段被称为短线性基序(Short Linear Motifs,SLiMs),对其进行识别可为蛋白质功能研究提供重要信息。然而,由于基序长度较短且保守程度各异,其识别颇具难度——难以准确估算其出现的统计学显著性,因此目前已被发现的SLiMs仅占极小一部分。本文描述了一种用于发现SLiMs的方法,该方法基于进化上无关联的蛋白质在同一生物学、信号转导或代谢通路中的出现特征,并给出了其在重新发现已知基序与发现新型基序两方面均有效的具体实例。该流程的自动化实现版本可供下载,可识别出具有统计学意义的基序,自动为其注释功能、进化与结构信息,并整合至可被检索与查询的数据库中。现有一个针对7种生物的预计算数据填充后的数据库实例,可通过公开服务器访问。我们认为该数据库本身即可成为生命科学领域的一项实用资源(http://www.biocomputing.it/modipath)。

创建时间:
2016-01-18
二维码
社区交流群
二维码
科研交流群
商业服务