Predicting membrane protein types by incorporating protein topology, domains, signal peptides, and physicochemical properties into the general form of Chou's pseudo amino acid composition

Predicting membrane protein types by incorporating protein topology, domains, signal peptides, and physicochemical properties into the general form of Chou's pseudo amino acid composition
复制标题

DOI:
10.1016/j.jtbi.2012.10.033
复制
发表时间:
2013-02-07
影响因子:
2
通讯作者:
Li, Kuo-Bin
Li, Kuo-Bin
中科院分区:
生物学4区
文献类型:
--
作者:
Chen, Yen-Kuang;Li, Kuo-Bin

文献摘要

被引文献

相似文献

未注销的膜蛋白的类型信息为其生物学功能提供了重要的提示。尽管昂贵的实验室程序,但并非总是可行的膜蛋白类型的实验性测定,从而产生了生物信息学方法的发展。本文介绍了使用蛋白质序列预测膜蛋白类型的新型计算分类器。分类器包含一个单一的支持向量机的集合,使用以下序列属性:(1)阳离子贴片大小,跨膜段的方向和拓扑; (2)氨基酸物理化学特性; (3)存在信号肽或锚固剂; (4)特定蛋白基序。实施了一项新的投票计划,以应对多级预测。训练和测试序列均从瑞士普罗特(Swissprot)收集。除去同源蛋白,以使数据集中没有序列身份高于40%的序列。分类器的性能通过折刀交叉验证和独立的测试实验评估。结果表明,在八种膜蛋白类型中的七种中,所提出的分类器在预测准确性方面的表现优于早期预测因子。总体准确性从78.3%增加到88.2%。与早期的方法不同,这些方法在很大程度上取决于特定于位置的替代矩阵和氨基酸组成,而拟议的分类器中实现的大多数序列属性具有支持的文献证据。分类器已被部署为Web服务器,可以通过http://bsaltools.ym.edu.tw/predmpt访问。 (c)2012 Elsevier Ltd.保留所有权利。
The type information of un-annotated membrane proteins provides an important hint for their biological functions. The experimental determination of membrane protein types, despite being more accurate and reliable, is not always feasible due to the costly laboratory procedures, thereby creating a need for the development of bioinformatics methods. This article describes a novel computational classifier for the prediction of membrane protein types using proteins' sequences. The classifier, comprising a collection of one-versus-one support vector machines, makes use of the following sequence attributes: (1) the cationic patch sizes, the orientation, and the topology of transmembrane segments; (2) the amino acid physicochemical properties; (3) the presence of signal peptides or anchors; and (4) the specific protein motifs. A new voting scheme was implemented to cope with the multi-class prediction. Both the training and the testing sequences were collected from SwissProt. Homologous proteins were removed such that there is no pair of sequences left in the datasets with a sequence identity higher than 40%. The performance of the classifier was evaluated by a Jackknife cross-validation and an independent testing experiments. Results show that the proposed classifier outperforms earlier predictors in prediction accuracy in seven of the eight membrane protein types. The overall accuracy was increased from 78.3% to 88.2%. Unlike earlier approaches which largely depend on position-specific substitution matrices and amino acid compositions, most of the sequence attributes implemented in the proposed classifier have supported literature evidences. The classifier has been deployed as a web server and can be accessed at http://bsaltools.ym.edu.tw/predmpt. (C) 2012 Elsevier Ltd. All rights reserved.