Improving consensus contact prediction via server correlation reduction

Improving consensus contact prediction via server correlation reduction
复制标题

DOI:
10.1186/1472-6807-9-28
复制
发表时间:
2009-05-06
影响因子:
--
通讯作者:
Li, Ming
Li, Ming
中科院分区:
生物4区
文献类型:
--
作者:
Gao, Xin;Bu, Dongbo;Li, Ming

文献摘要

被引文献

相似文献

背景:蛋白质残基间接触在蛋白质结构的测定和预测中起着至关重要的作用。以往的接触预测研究表明,尽管基于模板的一致性方法在典型模板目标上优于基于序列的方法,但在新褶皱目标上的一致性方法表现不佳。然而,我们发现,即使对于新的折叠目标,由线程程序生成的模型也可以包含许多真实的接触。挑战在于如何识别它们。结果:本文建立了一个用于一致接触预测的整数线性规划模型。与简单多数投票法假设所有独立的服务器同等重要不同,该方法利用最大似然估计来评估它们之间的相关性,并利用主成分分析来从中提取独立的潜在服务器。然后应用整数线性规划方法为每个潜在服务器分配权重,以最大限度地提高真接触和假接触之间的差异。在CASP7数据集上对该方法进行了验证。如果对前L/5预测接触进行评估,其中L为蛋白质大小,平均准确率为73%,远远高于之前报道的任何研究。此外,如果只考虑15个新的折叠CASP7目标,我们的方法平均准确率为37%,远远优于多数投票方法、SVM-LOMETS、SVM-SEQ和SAM-T06。这些方法的平均准确率分别为13.0%、10.8%、25.8%和21.2%。结论:降低服务器相关性,优化组合独立潜在服务器,较传统的共识方法有显著的改进。该方法有望为蛋白质结构的精化和预测提供有力的工具。
Background: Protein inter-residue contacts play a crucial role in the determination and prediction of protein structures. Previous studies on contact prediction indicate that although template-based consensus methods outperform sequence-based methods on targets with typical templates, such consensus methods perform poorly on new fold targets. However, we find out that even for new fold targets, the models generated by threading programs can contain many true contacts. The challenge is how to identify them.Results: In this paper, we develop an integer linear programming model for consensus contact prediction. In contrast to the simple majority voting method assuming that all the individual servers are equally important and independent, the newly developed method evaluates their correlation by using maximum likelihood estimation and extracts independent latent servers from them by using principal component analysis. An integer linear programming method is then applied to assign a weight to each latent server to maximize the difference between true contacts and false ones. The proposed method is tested on the CASP7 data set. If the top L/5 predicted contacts are evaluated where L is the protein size, the average accuracy is 73%, which is much higher than that of any previously reported study. Moreover, if only the 15 new fold CASP7 targets are considered, our method achieves an average accuracy of 37%, which is much better than that of the majority voting method, SVM-LOMETS, SVM-SEQ, and SAM-T06. These methods demonstrate an average accuracy of 13.0%, 10.8%, 25.8% and 21.2%, respectively.Conclusion: Reducing server correlation and optimally combining independent latent servers show a significant improvement over the traditional consensus methods. This approach can hopefully provide a powerful tool for protein structure refinement and prediction use.