Non-Alignment Features Based Enzyme/Non-Enzyme Classification Using an Ensemble Method.

Non-Alignment Features Based Enzyme/Non-Enzyme Classification Using an Ensemble Method.
复制标题

DOI:
10.1109/icmla.2010.167
复制
发表时间:
2010-12-12
期刊:
Proceedings of the ... International Conference on Machine Learning and Applications. International Conference on Machine Learning and Applications
影响因子:
--
通讯作者:
Wang X
Wang X
中科院分区:
其他
文献类型:
--
作者:
Davidson NJ;Wang X

文献摘要

被引文献

相似文献

随着越来越多的蛋白质结构在功能未知的情况下被解析,使用计算方法来帮助从结构预测蛋白质功能变得越来越重要。一些计算方法通过与具有已知功能的同源蛋白质比对来预测蛋白质功能,但是如果不能识别这种同源性,则它们无法工作。在本文中,我们分类酶/非酶使用非对齐功能。我们提出了一种新的集成方法,包括三个支持向量机(SVM)和两个k-近邻算法(k-NN),并使用一个简单的多数表决规则。对改编自多布森和多伊格的697种酶和480种非酶的数据集的测试显示,在10倍交叉验证中的准确度为85.59%,在留一法验证中的准确度为86.49%。预测精度比其他基于非对齐特征的方法好得多,甚至比基于对齐特征的方法略好。据我们所知,我们的方法是第一次使用集成方法来分类酶/非酶,是上级比一个单一的分类器。
As a growing number of protein structures are resolved without known functions, using computational methods to help predict protein functions from the structures becomes more and more important. Some computational methods predict protein functions by aligning to homologous proteins with known functions, but they fail to work if such homology cannot be identified. In this paper we classify enzymes/non-enzymes using non-alignment features. We propose a new ensemble method that includes three support vector machines (SVM) and two k-nearest neighbor algorithms (k-NN) and uses a simple majority voting rule. The test on a data set of 697 enzymes and 480 non-enzymes adapted from Dobson and Doig shows 85.59% accuracy in a 10-fold cross validation and 86.49% accuracy in a leave-one-out validation. The prediction accuracy is much better than other non-alignment features based methods and even slightly better than alignment features based methods. To our knowledge, our method is the first time to use ensemble methods to classify enzymes/non-enzymes and is superior over a single classifier.