遇见数据集

A Large-Scale Assessment of Nucleic Acids Binding Site Prediction Programs

收藏
Figshare2016-01-15 更新2026-04-29 收录
官方服务:

资源简介:

Computational prediction of nucleic acid binding sites in proteins are necessary to disentangle functional mechanisms in most biological processes and to explore the binding mechanisms. Several strategies have been proposed, but the state-of-the-art approaches display a great diversity in i) the definition of nucleic acid binding sites; ii) the training and test datasets; iii) the algorithmic methods for the prediction strategies; iv) the performance measures and v) the distribution and availability of the prediction programs. Here we report a large-scale assessment of 19 web servers and 3 stand-alone programs on 41 datasets including more than 5000 proteins derived from 3D structures of protein-nucleic acid complexes. Well-defined binary assessment criteria (specificity, sensitivity, precision, accuracy…) are applied. We found that i) the tools have been greatly improved over the years; ii) some of the approaches suffer from theoretical defects and there is still room for sorting out the essential mechanisms of binding; iii) RNA binding and DNA binding appear to follow similar driving forces and iv) dataset bias may exist in some methods.

在蛋白质中开展核酸结合位点的计算预测,是解析多数生物过程的功能机制、探究分子结合机制的必要前提。目前已有诸多预测策略被提出,但现有最先进的方法在以下五个方面存在显著差异:i) 核酸结合位点的定义标准;ii) 训练与测试数据集;iii) 预测策略所采用的算法方法;iv) 性能评估指标;v) 预测工具的分发与可获取性。本研究针对源自蛋白质-核酸复合物三维结构的41个数据集(涵盖5000余个蛋白质),对19个网页服务器工具与3个独立运行程序开展了大规模评估,并采用了定义明确的二元分类评估准则(包括特异性、敏感性、精确率、准确率等)。研究结果显示:i) 各类预测工具在多年间已有显著提升;ii) 部分方法存在理论缺陷,针对结合核心机制的阐释仍有优化空间;iii) RNA结合与DNA结合似乎遵循相似的驱动机制;iv) 部分方法中可能存在数据集偏倚问题。

创建时间:
2016-01-15
二维码
社区交流群
二维码
科研交流群
商业服务