A hierarchical approach to all-atom protein loop prediction

A hierarchical approach to all-atom protein loop prediction
复制标题

DOI:
10.1002/prot.10613
复制
发表时间:
2004-05-01
影响因子:
2.9
通讯作者:
Friesner, RA
Friesner, RA
中科院分区:
生物学4区
文献类型:
--
作者:
Jacobson, MP;Pincus, DL;Friesner, RA

文献摘要

被引文献

相似文献

将全原子力场(以及显式或隐式溶剂模型)应用于蛋白质同源建模任务,如侧链和环预测,仍然具有挑战性,这既是因为单个能量计算的费用,也是因为对崎岖的全原子能量表面进行采样的困难。在这里,我们通过开发许多新的算法来解决循环预测问题的这一挑战,重点是多尺度和分层技术。作为评估我们的循环预测算法性能的第一步,我们已经将其应用于重建自然结构中的循环的问题;我们还明确包括晶体填充,以提供与晶体结构的公平比较。简而言之,通过使用基于二面角的积聚过程来生成大量的环路,然后对所选择的环路结构进行聚类、侧链优化和完全能量最小化的迭代循环。我们使用迄今用于验证循环预测方法的最大测试集来评估该方法,总共有833个循环,长度从4到12个残基不等。5个残基环的平均/中位主干均方根偏差(RMSD)为0.42/0.24埃,8个残基环的RMSD为1.00/0.44埃,11个残基环的RMSD为2.47/1.83埃。由于少量的异常值,中位数的RMSD大大低于平均值;对这些失败的原因进行了一些详细的研究,许多可以归因于可滴定残基的质子化状态的指定错误,模拟中的配体遗漏,以及在少数情况下,实验确定的结构中可能的错误。当滤除数据集中这些明显的问题时,5个残基环的天然结构的平均RMSD提高到0.43埃,8个残基环的平均RMSD为0.84埃,11个残基环的平均RMSD为1.63埃。在绝大多数情况下,该方法定位的能量最小值低于或等于最小化本机循环的能量最小值,从而表明采样很少限制预测精度。据我们所知,总体结果是迄今为止报道的最好的,我们将这一成功归功于准确的全原子能量函数、环路建立和侧链优化的有效方法,特别是对于较长的环路,分层精化协议。(C)2004年Wiley-Liss公司
The application of all-atom force fields (and explicit or implicit solvent models) to protein homology-modeling tasks such as side-chain and loop prediction remains challenging both because of the expense of the individual energy calculations and because of the difficulty of sampling the rugged all-atom energy surface. Here we address this challenge for the problem of loop prediction through the development of numerous new algorithms, with an emphasis on multiscale and hierarchical techniques. As a first step in evaluating the performance of our loop prediction algorithm, we have applied it to the problem of reconstructing loops in native structures; we also explicitly include crystal packing to provide a fair comparison with crystal structures. In brief, large numbers of loops are generated by using a dihedral angle-based buildup procedure followed by iterative cycles of clustering, side-chain optimization, and complete energy minimization of selected loop structures. We evaluate this method by using the largest test set yet used for validation of a loop prediction method, with a total of 833 loops ranging from 4 to 12 residues in length. Average/median backbone root-mean-square deviations (RMSDs) to the native structures (superimposing the body of the protein, not the loop itself) are 0.42/0.24 Angstrom for 5 residue loops, 1.00/0.44 Angstrom for 8 residue loops, and 2.47/1.83 Angstrom for 11 residue loops. Median RMSDs are substantially lower than the averages because of a small number of outliers; the causes of these failures are examined in some detail, and many can be attributed to errors in assignment of protonation states of titratable residues, omission of ligands from the simulation, and, in a few cases, probable errors in the experimentally determined structures. When these obvious problems in the data sets are filtered out, average RMSDs to the native structures improve to 0.43 Angstrom for 5 residue loops, 0.84 Angstrom for 8 residue loops, and 1.63 Angstrom for 11 residue loops. In the vast majority of cases, the method locates energy minima that are lower than or equal to that of the minimized native loop, thus indicating that sampling rarely limits prediction accuracy. The overall results are, to our knowledge, the best reported to date, and we attribute this success to the combination of an accurate all-atom energy function, efficient methods for loop buildup and side-chain optimization, and, especially for the longer loops, the hierarchical refinement protocol. (C) 2004 Wiley-Liss, Inc.