Secondary structure prediction with support vector machines

Secondary structure prediction with support vector machines
复制标题

DOI:
10.1093/bioinformatics/btg223
复制
发表时间:
2003-09-01
期刊:
影响因子:
5.8
通讯作者:
Jones, DT
Jones, DT
中科院分区:
生物学3区
文献类型:
--
作者:
Ward, JJ;McGuffin, LJ;Jones, DT

文献摘要

被引文献

相似文献

动机:描述并评估了一种使用支持​​向量机 (SVM) 预测蛋白质二级结构的新方法。该研究旨在使用替代技术开发一种可靠的预测方法,并研究 SVM 对此类生物信息学问题的适用性。方法:训练二元 SVM 来区分两个结构类别。二元分类器以多种方式组合来预测多类二级结构。结果:通过交叉验证估计每个蛋白质的平均三态预测精度 (Q(3)) 为 77.07+/-0.26%,段重叠 (Sov) 得分为 73.32+/-0.39%。 SVM 在 121 种蛋白质的非同源测试集上的表现与“最先进的”PSIPRED 预测方法类似,尽管训练的例子要少得多。 SVM、PSIPRED 和 PROFsec 的简单共识可实现比单独方法显着更高的预测精度。
Motivation: A new method that uses support vector machines (SVMs) to predict protein secondary structure is described and evaluated. The study is designed to develop a reliable prediction method using an alternative technique and to investigate the applicability of SVMs to this type of bioinformatics problem.Methods: Binary SVMs are trained to discriminate between two structural classes. The binary classifiers are combined in several ways to predict multi-class secondary structure.Results: The average three-state prediction accuracy per protein (Q(3)) is estimated by cross-validation to be 77.07+/-0.26% with a segment overlap (Sov) score of 73.32+/-0.39%. The SVM performs similarly to the 'state-of-the-art' PSIPRED prediction method on a non-homologous test set of 121 proteins despite being trained on substantially fewer examples. A simple consensus of the SVM, PSIPRED and PROFsec achieves significantly higher prediction accuracy than the individual methods.