遇见数据集

<i>De Novo</i> Structure Prediction of Globular Proteins Aided by Sequence Variation-Derived Contacts

收藏
NIAID Data Ecosystem2026-03-08 收录
官方服务:

资源简介:

The advent of high accuracy residue-residue intra-protein contact prediction methods enabled a significant boost in the quality of de novo structure predictions. Here, we investigate the potential benefits of combining a well-established fragment-based folding algorithm – FRAGFOLD, with PSICOV, a contact prediction method which uses sparse inverse covariance estimation to identify co-varying sites in multiple sequence alignments. Using a comprehensive set of 150 diverse globular target proteins, up to 266 amino acids in length, we are able to address the effectiveness and some limitations of such approaches to globular proteins in practice. Overall we find that using fragment assembly with both statistical potentials and predicted contacts is significantly better than either statistical potentials or contacts alone. Results show up to nearly 80% of correct predictions (TM-score ≥0.5) within analysed dataset and a mean TM-score of 0.54. Unsuccessful modelling cases emerged either from conformational sampling problems, or insufficient contact prediction accuracy. Nevertheless, a strong dependency of the quality of final models on the fraction of satisfied predicted long-range contacts was observed. This not only highlights the importance of these contacts on determining the protein fold, but also (combined with other ensemble-derived qualities) provides a powerful guide as to the choice of correct models and the global quality of the selected model. A proposed quality assessment scoring function achieves 0.93 precision and 0.77 recall for the discrimination of correct folds on our dataset of decoys. These findings suggest the approach is well-suited for blind predictions on a variety of globular proteins of unknown 3D structure, provided that enough homologous sequences are available to construct a large and accurate multiple sequence alignment for the initial contact prediction step.

高精度残基-残基蛋白质内接触预测方法的问世,显著推动了从头蛋白质结构预测质量的提升。本研究探讨了将成熟的基于片段折叠算法FRAGFOLD与PSICOV相结合的潜在优势:PSICOV是一种借助稀疏逆协方差估计,从多序列比对中识别共变位点的接触预测方法。本研究采用包含150个不同长度(最长达266个氨基酸)的多样化球状靶蛋白的综合数据集,分析了此类方法在实际应用中针对球状蛋白的有效性与部分局限性。整体而言,我们发现同时结合统计势能与预测接触信息的片段组装方法,其性能显著优于仅使用统计势能或仅使用预测接触信息的单一方法。分析数据集的结果显示,正确预测(TM评分(TM-score)≥0.5)占比最高接近80%,平均TM评分为0.54。建模失败的案例主要源于构象采样问题,或是接触预测精度不足。尽管如此,我们观察到最终模型的质量与满足预测要求的长程接触占比存在显著相关性。这不仅凸显了此类长程接触在决定蛋白质折叠模式中的关键作用,还(结合其他集成衍生的质量指标)为正确模型的筛选以及所选模型的整体质量评估提供了强有力的指导依据。在我们的诱饵模型(decoys)数据集上,所提出的质量评估评分函数在区分正确折叠模式时,可达到0.93的精确率与0.77的召回率。上述研究结果表明,该方法非常适用于对多种未知三维结构的球状蛋白进行盲预测,前提是能够获取足够的同源序列,以构建大型且准确的多序列比对,用于初始接触预测步骤。

创建时间:
2014-03-17
二维码
社区交流群
二维码
科研交流群
商业服务