A new representation for protein secondary structure prediction based on frequent patterns

A new representation for protein secondary structure prediction based on frequent patterns
复制标题

DOI:
10.1093/bioinformatics/btl453
复制
发表时间:
2006-11-01
期刊:
影响因子:
5.8
通讯作者:
Kramer, Stefan
Kramer, Stefan
中科院分区:
生物学3区
文献类型:
--
作者:
Birzele, Fabian;Kramer, Stefan

文献摘要

被引文献

相似文献

动机:一个新的表示蛋白质二级结构预测的基础上频繁的氨基酸模式进行了描述和评估。我们详细讨论了如何使用逐层搜索技术识别蛋白质序列数据库中的频繁模式,如何从这些模式中定义一组特征,以及如何使用这些特征使用支持向量机(SVM)预测蛋白质序列的二级结构。三组不同的功能的基础上频繁的模式进行了评估,在一个盲测试设置使用150个目标从伊娃比赛和PSI-PRED,PHD和PROFsec的预测相比。尽管只对940种蛋白质进行了训练,但基于这种新表示的简单SVM分类器产生的结果与PSI-PRED和PROFsec相当。最后,我们表明,该方法有助于共识预测的重要信息。
Motivation: A new representation for protein secondary structure prediction based on frequent amino acid patterns is described and evaluated. We discuss in detail how to identify frequent patterns in a protein sequence database using a level-wise search technique, how to define a set of features from those patterns and how to use those features in the prediction of the secondary structure of a protein sequence using support vector machines (SVMs).Results: Three different sets of features based on frequent patterns are evaluated in a blind testing setup using 150 targets from the EVA contest and compared to predictions of PSI-PRED, PHD and PROFsec. Despite being trained on only 940 proteins, a simple SVM classifier based on this new representation yields results comparable to PSI-PRED and PROFsec. Finally, we show that the method contributes significant information to consensus predictions.