Comparison of probabilistic combination methods for protein secondary structure prediction

Comparison of probabilistic combination methods for protein secondary structure prediction
复制标题

DOI:
10.1093/bioinformatics/bth370
复制
发表时间:
2004-11-22
期刊:
影响因子:
5.8
通讯作者:
Gopalakrishnan, V
Gopalakrishnan, V
中科院分区:
生物学3区
文献类型:
--
作者:
Liu, Y;Carbonell, J;Gopalakrishnan, V

文献摘要

被引文献

相似文献

动机:蛋白质二级结构预测是理解蛋白质如何在三维空间折叠的重要一步。最近的信息论分析表明,相邻二级结构之间的相关性比相邻氨基酸之间的相关性强得多。在这篇文章中,我们专注于序列的组合问题,即在整个序列的约束下,将单个或多个预测系统的得分或分配组合起来,作为改进蛋白质二级结构预测的目标。结果:我们应用了几种图形链模型来解决组合问题,并表明它们始终比传统的基于窗口的方法更有效。特别是,条件随机场(CRF)适度提高螺旋的预测,更重要的是,β片,这是蛋白质二级结构预测的主要瓶颈。
Motivation: Protein secondary structure prediction is an important step towards understanding how proteins fold in three dimensions. Recent analysis by information theory indicates that the correlation between neighboring secondary structures are much stronger than that of neighboring amino acids. In this article, we focus on the combination problem for sequences, i.e. combining the scores or assignments from single or multiple prediction systems under the constraint of a whole sequence, as a target for improvement in protein secondary structure prediction.Results: We apply several graphical chain models to solve the combination problem and show that they are consistently more effective than the traditional window-based methods. In particular, conditional random fields (CRFs) moderately improve the predictions for helices and, more importantly, for beta sheets, which are the major bottleneck for protein secondary structure prediction.