遇见数据集
数据链接:
官方服务:

资源简介:

RepeatsDB (https://repeatsdb.org/) is a database of annotated tandem repeat protein structures. Tandem repeats pose a difficult problem for the analysis of protein structures, as the underlying sequence can be highly degenerate. Several repeat types haven been studied over the years, but their annotation was done in a case-by-case basis, thus making large-scale analysis difficult. We developed RepeatsDB to fill this gap. Using state-of-the-art repeat detection methods and manual curation, we systematically annotated the Protein Data Bank, predicting 10 745 repeat structures. In all, 2797 structures were classified according to a recently proposed classification schema, which was expanded to accommodate new findings. In addition, detailed annotations were performed in a subset of 321 proteins. These annotations feature information on start and end positions for the repeat regions and units. RepeatsDB is an ongoing effort to systematically classify and annotate structural protein repeats in a consistent way. It provides users with the possibility to access and download high-quality datasets either interactively or programmatically through web services.

RepeatsDB(https://repeatsdb.org/)是一个收录经注释的串联重复(tandem repeat)蛋白质结构的数据库。串联重复的分析是蛋白质结构研究中的一大难点,因其对应的底层序列可呈现高度简并性。多年来,学界已对多种重复蛋白类型展开研究,但相关注释均采用逐例分析的方式开展,这极大阻碍了大规模分析的推进。为此我们开发了RepeatsDB以填补这一空白。借助当前最先进的重复序列检测方法与人工精修流程,我们对蛋白质数据银行(Protein Data Bank)中的数据进行了系统性注释,共预测得到10745个重复蛋白质结构。其中2797个结构依据最新提出的分类框架完成了分类,该框架后续还进行了扩展以适配新的研究发现。此外,我们还针对321个蛋白质子集开展了详细注释,这些注释包含重复区域与重复单元的起始、终止位置信息。RepeatsDB是一项持续性项目,旨在以统一标准对蛋白质结构重复单元开展系统性分类与注释工作。该数据库支持用户通过网页服务,以交互式或编程式两种方式访问并下载高质量数据集。

提供机构:
比利时列日大学
搜集汇总
数据集介绍
RepeatsDB 数据集图片
背景与挑战
背景概述
RepeatsDB是一个专门用于注释和分类结构串联重复蛋白质(STRPs)的数据库,提供了重复区域的起止位置、重复单元以及基于类、拓扑、折叠和家族的四级分类体系。该数据库通过系统化方法对蛋白质结构进行大规模标注,旨在支持对串联重复蛋白质的持续分析和研究。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务