Critical assessment of high-throughput standalone methods for secondary structure prediction

Critical assessment of high-throughput standalone methods for secondary structure prediction
复制标题

用于二级结构预测的高通量独立方法的批判性评估

DOI:
10.1093/bib/bbq088
复制
发表时间:
2011-01
影响因子:
9.5
通讯作者:
Marcin J. Mizianty
Marcin J. Mizianty
中科院分区:
生物学2区
文献类型:
--
作者:
Qingbo Bao;Lukasz Kurgan;Tuo Zhang;Wojciech Stach;Ke Chen;Kanaka Durga Kedarisetti;Hua Zhang;Marcin J. Mizianty

文献摘要

参考文献

被引文献

相似文献

基于序列的蛋白质二级结构预测在分析和预测蛋白质的众多结构和功能特征方面得到了广泛的应用和越来越多的应用。由于最近缺乏对众多预测方法的全面和大规模的比较,导致对SS预测器的选择往往是随意的。为了解决这一空白,我们比较和分析了一大组1975年蛋白质上的12个流行的、独立的和高通量的预测器,以提供深入、新颖和实用的见解。我们表明,不存在通用的最佳预测值,因此需要详细的比较研究来支持针对特定应用的SS预测值的知情选择。我们的研究表明,目前SS预测的三态准确率(Q3)和分段重叠(SOV3)分别达到82%和81%。我们证明,精心设计的基于共识的预测者将第三季度的预测提高了2%,基于同源建模的方法比从头计算方法的第三季度显著提高了1.5%。我们的经验分析表明,溶剂暴露和柔性线圈的预测质量比埋入式和刚性线圈高,而股和螺旋的预测质量则相反。我们还表明,更长的螺旋更容易预测,这与更难发现的更长的链形成对比。目前的方法将1-6%的链残基与螺旋残基混淆,反之亦然,而且它们对β-桥和3(10)-螺旋构象中的残基表现不佳。最后,我们将四种性能良好的方法的独立实现的预测与它们对应的Web服务器进行了比较。
Sequence-based prediction of protein secondary structure (SS) enjoys wide-spread and increasing use for the analysis and prediction of numerous structural and functional characteristics of proteins. The lack of a recent comprehensive and large-scale comparison of the numerous prediction methods results in an often arbitrary selection of a SS predictor. To address this void, we compare and analyze 12 popular, standalone and high-throughput predictors on a large set of 1975 proteins to provide in-depth, novel and practical insights. We show that there is no universally best predictor and thus detailed comparative studies are needed to support informed selection of SS predictors for a given application. Our study shows that the three-state accuracy (Q3) and segment overlap (SOV3) of the SS prediction currently reach 82% and 81%, respectively. We demonstrate that carefully designed consensus-based predictors improve the Q3 by additional 2% and that homology modeling-based methods are significantly better by 1.5% Q3 than ab initio approaches. Our empirical analysis reveals that solvent exposed and flexible coils are predicted with a higher quality than the buried and rigid coils, while inverse is true for the strands and helices. We also show that longer helices are easier to predict, which is in contrast to longer strands that are harder to find. The current methods confuse 1-6% of strand residues with helical residues and vice versa and they perform poorly for residues in the β- bridge and 3(10)-helix conformations. Finally, we compare predictions of the standalone implementations of four well-performing methods with their corresponding web servers.
DOI: 10.1371/journal.pone.0002399
发表时间: 2008-06-11
期刊: PloS one
影响因子: 3.7
作者:
Shen H;Chou JJ
通讯作者: Chou JJ
分析用于二级结构预测的最佳隐藏马尔可夫模型。
DOI: 10.1186/1472-6807-6-25
发表时间: 2006-12-13
影响因子: --
作者:
Martin, Juliette;Gibrat, Jean-Francois;Rodolphe, Francois
通讯作者: Rodolphe, Francois
DOI: 10.1093/nar/gkn238
发表时间: 2008-07-01
影响因子: 14.9
作者:
Cole C;Barber JD;Barton GJ
通讯作者: Barton GJ
DOI: 10.1093/nar/gki396
发表时间: 2005-07-01
影响因子: 14.9
作者:
Cheng, J;Randall, AZ;Sweredoski, MJ;Baldi, P
通讯作者: Baldi, P
DOI: 10.1093/bioinformatics/bti203
发表时间: 2005-04-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Pollastri, G;McLysaght, A
通讯作者: McLysaght, A