BCL::Fold - <em>De Novo</em> Prediction of Complex and Large Protein Topologies by Assembly of Secondary Structure Elements
收藏资源简介:
Computational de novo protein structure prediction is limited to small proteins of simple topology. The present work explores an approach to extend beyond the current limitations through assembling protein topologies from idealized α-helices and β-strands. The algorithm performs a Monte Carlo Metropolis simulated annealing folding simulation. It optimizes a knowledge-based potential that analyzes radius of gyration, β-strand pairing, secondary structure element (SSE) packing, amino acid pair distance, amino acid environment, contact order, secondary structure prediction agreement and loop closure. Discontinuation of the protein chain favors sampling of non-local contacts and thereby creation of complex protein topologies. The folding simulation is accelerated through exclusion of flexible loop regions further reducing the size of the conformational search space. The algorithm is benchmarked on 66 proteins with lengths between 83 and 293 amino acids. For 61 out of these proteins, the best SSE-only models obtained have an RMSD100 below 8.0 Å and recover more than 20% of the native contacts. The algorithm assembles protein topologies with up to 215 residues and a relative contact order of 0.46. The method is tailored to be used in conjunction with low-resolution or sparse experimental data sets which often provide restraints for regions of defined secondary structure.
从头计算蛋白质结构预测目前仅能处理拓扑结构简单的小型蛋白质。本研究探索了一种突破现有局限的方法:通过将理想化α螺旋与β折叠链组装来构建蛋白质拓扑结构。该算法采用蒙特卡洛-梅特罗波利斯(Monte Carlo Metropolis)模拟退火折叠模拟策略,优化了一项基于知识的势能函数,该函数会分析回转半径、β链配对、二级结构元件(SSE)排布、氨基酸残基对间距、氨基酸残基所处环境、接触序、二级结构预测一致性以及环区闭合情况。打断蛋白质链的连续性有助于采样非局部接触,进而构建复杂的蛋白质拓扑结构。通过排除柔性环区,折叠模拟的速度得以提升,同时进一步缩小了构象搜索空间的规模。本算法在66条长度介于83至293个氨基酸残基的蛋白质上进行了基准测试,其中61个蛋白质的最优纯二级结构元件模型的RMSD100低于8.0 Å,且恢复了超过20%的天然接触。该算法可组装拥有多达215个残基、相对接触序为0.46的蛋白质拓扑结构。本方法专为结合低分辨率或稀疏实验数据集而设计,这类数据集通常可为已确定二级结构的区域提供约束条件。



