SimFold energy function for de novo protein structure prediction: Consensus with Rosetta

SimFold energy function for de novo protein structure prediction: Consensus with Rosetta
复制标题

DOI:
10.1002/prot.20748
复制
发表时间:
2006-02-01
影响因子:
2.9
通讯作者:
Takada, S
Takada, S
中科院分区:
生物学4区
文献类型:
--
作者:
Fujitsuka, Y;Chikenji, G;Takada, S

文献摘要

被引文献

相似文献

对于有新折叠的蛋白质来说,通过计算机模拟折叠来预测蛋白质的三级结构仍然是非常困难的。在这里,我们开发了一个粗粒度的能量函数SimFold,用于从头结构预测,对38个测试蛋白质进行了片段组装模拟预测的基准测试,并提出了Rosetta的共识预测。SimFold能量由许多项组成,这些项在物理化学考虑的基础上考虑了溶剂诱导的效应。在基准测试中,SimFold成功地预测了38种蛋白质中12种蛋白质的天然结构在6.5埃范围内;这一成功率与使用默认参数运行的公开版本Rosetta(ab initio version 1.2)相同。我们研究了SimFold中哪些能量项对结构预测性能有贡献,发现疏水相互作用对预测最关键,而其他序列特异性项则具有微弱但积极的作用。在基准测试中,SimFold和Rosetta预测的蛋白质在12种蛋白质中有5种是不同的,这使得我们引入了共识预测。通过组合诱饵,我们成功预测了16种蛋白质,比SimFold或Rosetta分别多4种。对于38种蛋白质中的每一种,通过将采样的结构空间映射到两个维度上来定性地比较SimFold和Rosetta生成的结构集合。对于两种方法中的一种预测成功而另一种预测失败的蛋白质,前者具有位于天然蛋白质周围的较不分散的系综。对于两种方法都能成功预测的蛋白质,通常会混淆两个集合。
Predicting protein tertiary structures by in silico folding is still very difficult for proteins that have new folds. Here, we developed a coarse-grained energy function, SimFold, for de novo structure prediction, performed a benchmark test of prediction with fragment assembly simulations for 38 test proteins, and proposed consensus prediction with Rosetta. The SimFold energy consists of many terms that take into account solvent-induced effects on the basis of physicochemical consideration. In the benchmark test, SimFold succeeded in predicting native structures within 6.5 angstrom for 12 of 38 proteins; this success rate was the same as that by the publicly available version of Rosetta (ab initio version 1.2) run with default parameters. We investigated which energy terms in SimFold contribute to structure prediction performance, finding that the hydrophobic interaction is the most crucial for the prediction, whereas other sequence-specific terms have weak but positive roles. In the benchmark, well-predicted proteins by SimFold and by Rosetta were not the same for 5 of 12 proteins, which led us to introduce consensus prediction. With combined decoys, we succeeded in prediction for 16 proteins, four more than SimFold or Rosetta separately. For each of 38 proteins, structural ensembles generated by SimFold and by Rosetta were qualitatively compared by mapping sampled structural space onto two dimensions. For proteins of which one of the two methods succeeded and the other failed in prediction, the former had a less scattered ensemble located around the native. For proteins of which both methods succeeded in prediction, often two ensembles were mixed up.