Support Vector Machine-based classification of protein folds using the structural properties of amino acid residues and amino acid residue pairs

Support Vector Machine-based classification of protein folds using the structural properties of amino acid residues and amino acid residue pairs
复制标题

DOI:
10.1093/bioinformatics/btm527
复制
发表时间:
2007-12-15
期刊:
影响因子:
5.8
通讯作者:
Nagarajaram, H. A.
Nagarajaram, H. A.
中科院分区:
生物学3区
文献类型:
--
作者:
Shamim, Mohammad Tabrez Anwar;Anwaruddin, Mohammad;Nagarajaram, H. A.

文献摘要

被引文献

相似文献

动机:折叠识别是蛋白质结构发现过程中的关键步骤,特别是当传统的序列比较方法无法产生令人信服的结构同源性时。虽然已经开发了许多方法用于蛋白质折叠识别,但它们的准确性仍然很低。这可以归因于折叠歧视feature.Results开发不足:我们已经开发了一种新的方法,蛋白质折叠识别使用的氨基酸残基和氨基酸残基对的结构信息。由于蛋白质折叠识别可以被视为蛋白质折叠分类问题,我们已经开发了一种基于支持向量机(SVM)的分类方法,使用二级结构状态和溶剂可及性状态频率的氨基酸和氨基酸对作为特征向量。在检查的单个属性中,氨基酸的二级结构状态频率给出了65.2%的总体准确度,用于倍数辨别,这优于文献中迄今为止报道的任何方法的准确度。二级结构状态频率与氨基酸和氨基酸对的溶剂可及性状态频率的组合进一步将倍数判别准确度提高到70%以上,这比最佳可用方法高出8%。在这项研究中,我们还测试了,第一次,一个所有在一起的多类方法被称为Crammer和Singer方法的蛋白质折叠分类。我们的研究表明,三个多类分类方法,即一对所有,一对一和Crammer和Singer方法,产生类似的预测。
Motivation: Fold recognition is a key step in the protein structure discovery process, especially when traditional sequence comparison methods fail to yield convincing structural homologies. Although many methods have been developed for protein fold recognition, their accuracies remain low. This can be attributed to insufficient exploitation of fold discriminatory features.Results: We have developed a new method for protein fold recognition using structural information of amino acid residues and amino acid residue pairs. Since protein fold recognition can be treated as a protein fold classification problem, we have developed a Support Vector Machine (SVM) based classifier approach that uses secondary structural state and solvent accessibility state frequencies of amino acids and amino acid pairs as feature vectors. Among the individual properties examined secondary structural state frequencies of amino acids gave an overall accuracy of 65.2% for fold discrimination, which is better than the accuracy by any method reported so far in the literature. Combination of secondary structural state frequencies with solvent accessibility state frequencies of amino acids and amino acid pairs further improved the fold discrimination accuracy to more than 70%, which is similar to 8% higher than the best available method. In this study we have also tested, for the first time, an all-together multi-class method known as Crammer and Singer method for protein fold classification. Our studies reveal that the three multi-class classification methods, namely one versus all, one versus one and Crammer and Singer method, yield similar predictions.