TOUCHSTONE II: A new approach to ab initio protein structure prediction

TOUCHSTONE II: A new approach to ab initio protein structure prediction
复制标题

DOI:
10.1016/s0006-3495(03)74551-2
复制
发表时间:
2003-08-01
影响因子:
3.4
通讯作者:
Skolnick, J
Skolnick, J
中科院分区:
生物学3区
文献类型:
--
作者:
Zhang, Y;Kolinski, A;Skolnick, J

文献摘要

被引文献

相似文献

我们开发了一种新的从头预测蛋白质结构的综合方法。蛋白质构象被描述为连接C -α原子的晶格链,带有连接的C -β原子和侧链质心。模型力场包括从蛋白质结构规律的统计分析中得出的各种基于知识的短程和长程势能。通过使30×60,000个诱饵与天然结构的均方根偏差(RMSD)和能量之间以及天然结构与诱饵集合之间的能隙的相关性最大化,对这些能量项的组合进行了优化。为了加速构象搜索,在蒙特卡罗模拟过程中使用了一种新开发的具有复合移动集的并行双曲线采样算法。我们利用这种策略成功折叠了41/100个小蛋白质(36 - 120个残基),在前五个聚类质心中,预测结构与天然结构的RMSD低于6.5埃。为了折叠更大尺寸的蛋白质以及提高小蛋白质的折叠产率,我们将来自我们的穿线程序PROSPECTOR的侧链接触预测纳入基本力场,在该程序中,同源蛋白质被排除在数据库之外。有了这些基于穿线的约束,该程序能够折叠83/125个测试蛋白质(36 - 174个残基),在前五个聚类质心中,结构与天然结构的RMSD低于6.5埃。这表明通过使用预测的三级约束,折叠有了显著改善,特别是当侧链接触预测的准确性>20%时。对于天然折叠选择,我们引入了取决于聚类密度以及能量和自由能组合的量,它们在选择天然结构方面比先前使用的聚类能量或聚类大小具有更高的判别能力,并且可用于盲模拟中的天然结构识别。这些程序很容易自动化,并正在基因组规模上实施。
We have developed a new combined approach for ab initio protein structure prediction. The protein conformation is described as a lattice chain connecting C-alpha atoms, with attached C-beta atoms and side-chain centers of mass. The model force field includes various short-range and long-range knowledge-based potentials derived from a statistical analysis of the regularities of protein structures. The combination of these energy terms is optimized through the maximization of correlation for 30 x 60,000 decoys between the root mean square deviation (RMSD) to native and energies, as well as the energy gap between native and the decoy ensemble. To accelerate the conformational search, a newly developed parallel hyperbolic sampling algorithm with a composite movement set is used in the Monte Carlo simulation processes. We exploit this strategy to successfully fold 41/100 small proteins (36 similar to 120 residues) with predicted structures having a RMSD from native below 6.5 Angstrom in the top five cluster centroids. To fold larger-size proteins as well as to improve the folding yield of small proteins, we incorporate into the basic force field side-chain contact predictions from our threading program PROSPECTOR where homologous proteins were excluded from the data base. With these threading-based restraints, the program can fold 83/125 test proteins (36 similar to 174 residues) with structures having a RMSD to native below 6.5 Angstrom in the top five cluster centroids. This shows the significant improvement of folding by using predicted tertiary restraints, especially when the accuracy of side-chain contact prediction is >20%. For native fold selection, we introduce quantities dependent on the cluster density and the combination of energy and free energy, which show a higher discriminative power to select the native structure than the previously used cluster energy or cluster size, and which can be used in native structure identification in blind simulations. These procedures are readily automated and are being implemented on a genomic scale.