Hum-PLoc: A novel ensemble classifier for predicting human protein subcellular localization

Hum-PLoc: A novel ensemble classifier for predicting human protein subcellular localization
复制标题

DOI:
10.1016/j.bbrc.2006.06.059
复制
发表时间:
2006-08-18
影响因子:
3.1
通讯作者:
Shen, Hong-Bin
Shen, Hong-Bin
中科院分区:
生物学4区
文献类型:
--
作者:
Chou, Kuo-Chen;Shen, Hong-Bin

文献摘要

被引文献

相似文献

预测人类蛋白质的亚细胞定位是一个具有挑战性的问题,特别是当未知的查询蛋白质与已知亚细胞位置的蛋白质没有显著的同源性时,以及当需要覆盖更多的位置时。为了解决这个问题,蛋白质样品通过基因本体(GO)数据库和两亲性伪氨基酸组成(PseAA)杂交来表达。基于这样的表示框架,通过投票系统融合了许多基本的个体分类器,开发了一种新的集成分类器,称为“Hum-PLoc”。这些基本分类器的“引擎”由KNN (k -最近邻)规则操作。作为示范,用集合分类器对人类蛋白质在以下12个位置进行了测试:(1)中心粒;(2)胞浆;(3)细胞骨架;(4)内质网;(5) extracell;(六)高尔基体;(7)溶酶体;(8)微粒体;(9)线粒体;(10)核;(11)过氧物酶体;(12)质膜。为了消除冗余和同源性偏差,本文研究的蛋白质在同一亚细胞位置上与任何其他蛋白质都没有>= 25%的序列同一性。通过叠刀交叉验证检验和独立数据集检验获得的总体成功率分别为81.1%和85.0%,在同样严格的数据集上,比现有方法的成功率高出50%以上。此外,给出了一个深刻而令人信服的分析,以阐明新预测器获得的压倒性的高成功率绝不是由于GO注释的微不足道的利用。这是因为,对于Swiss-Prot数据库中标注为“亚细胞位置未知”的蛋白,其在GO数据库中对应的GO号,大部分(超过99%)也标注为“细胞成分未知”。预测蛋白质亚细胞位置的信息和线索实际上被埋在一系列乏味的氧化石墨烯数字中,就像它们被埋在一堆复杂的氨基酸序列中一样,只是方式和“深度”不同。为了挖掘出它们的位置信息,需要一个复杂的操作引擎。目前的预测器就是其中一种,而且被证明是非常强大的。Hum-PLoc分类器可以在http://202.120.37.186/bioinf/hum上作为web服务器获得。(c) 2006爱思唯尔公司版权所有。
Predicting subcellular localization of human proteins is a challenging problem, especially when unknown query proteins do not have significant homology to proteins of known subcellular locations and when more locations need to be covered. To tackle the challenge, protein samples are expressed by hybridizing the gene ontology (GO) database and amphiphilic pseudo amino acid composition (PseAA). Based on such a representation frame, a novel ensemble classifier, called "Hum-PLoc", was developed by fusing many basic individual classifiers through a voting system. The "engine" of these basic classifiers was operated by the KNN (K-nearest neighbor) rule. As a demonstration, tests were performed with the ensemble classifier for human proteins among the following 12 locations: (1) centriole; (2) cytoplasm; (3) cytoskeleton; (4) endoplasmic reticulum; (5) extracell; (6) Golgi apparatus; (7) lysosome; (8) microsome; (9) mitochondrion; (10) nucleus; (11) peroxisome; (12) plasma membrane. To get rid of redundancy and homology bias, none of the proteins investigated here had >= 25% sequence identity to any other in a same subcellular location. The overall success rates thus obtained via the jackknife cross-validation test and independent dataset test were 81.1% and 85.0%, respectively, which are more than 50% higher than those obtained by the other existing methods on the same stringent datasets. Furthermore, an incisive and compelling analysis was given to elucidate that the overwhelmingly high success rate obtained by the new predictor is by no means due to a trivial utilization of the GO annotations. This is because, for those proteins with "subcellular location unknown" annotation in Swiss-Prot database, most (more than 99%) of their corresponding GO numbers in GO database are also annotated with "cellular component unknown". The information and clues for predicting subcellular locations of proteins are actually buried into a series of tedious GO numbers, just like they are buried into a pile of complicated amino acid sequences although with a different manner and "depth". To dig out the knowledge about their locations, a sophisticated operation engine is needed. And the current predictor is one of these kinds, and has proved to be a very powerful one. The Hum-PLoc classifier is available as a web-server at http://202.120.37.186/bioinf/hum. (c) 2006 Elsevier Inc. All rights reserved.