A New Method for Predicting the Subcellular Localization of Eukaryotic Proteins with Both Single and Multiple Sites: Euk-mPLoc 2.0

A New Method for Predicting the Subcellular Localization of Eukaryotic Proteins with Both Single and Multiple Sites: Euk-mPLoc 2.0
复制标题

预测单位点和多位点真核蛋白亚细胞定位的新方法:Euk-mPLoc 2.0

DOI:
10.1371/journal.pone.0009931
复制
发表时间:
2010-04-01
期刊:
影响因子:
3.7
通讯作者:
Shen, Hong-Bin
Shen, Hong-Bin
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Chou, Kuo-Chen;Shen, Hong-Bin

文献摘要

被引文献

相似文献

蛋白质的亚细胞位置信息对于细胞生物学的深入研究非常重要。它对于蛋白质组学、系统生物学和药物开发也非常有用。然而,大多数现有的预测蛋白质亚细胞位置的方法只能覆盖5到12个位置位点。此外,它们仅限于处理单位置蛋白质,因此无法处理多重蛋白质,因为多重蛋白质可以同时存在于两个或多个位置位点或在两个或多个位置位点之间移动。事实上,这类多重蛋白通常具有一些值得我们特别关注的重要生物学功能。一种名为“Euk-mPLoc 2.0”的新预测器是通过三种不同的伪氨基酸组成模式混合基因本体信息、功能域信息和顺序进化信息而开发的。它可用于鉴定以下 22 个位置的真核蛋白:(1) 顶体、(2) 细胞壁、(3) 中心粒、(4) 叶绿体、(5) 蓝藻、(6) 细胞质、(7) 细胞骨架、(8) 内质网、(9) 内体、(10) 细胞外、(11) 高尔基体、(12) 氢酶体、(13) 溶酶体、 (14)黑素体、(15)微粒体、(16)线粒体、(17)细胞核、(18)过氧化物酶体、(19)质膜、(20)质体、(21)纺锤体和(22)液泡。与现有的预测真核蛋白质亚细胞定位的方法相比,新的预测器更加强大和灵活,特别是在处理具有多个位置的蛋白质和没有可用登录号的蛋白质时。对于新构建的严格基准数据集,其中包含单位置和多位置蛋白质,并且其中没有任何蛋白质与同一位置的任何其他蛋白质具有成对序列同一性,Euk-mPLoc 2.0 实现的整体折刀成功率比任何现有预测器高出 24% 以上。作为一个用户友好的网络服务器,Euk-mPLoc 2.0可以免费访问http://www.csbio.sjtu.edu.cn/bioinf/euk-multi-2/。对于400个氨基酸的查询蛋白质序列,网络服务器大约需要15秒才能得出预测结果;序列越长,通常需要的时间就越多。预计本文提出的新颖方法和强大的预测将对分子细胞生物学、系统生物学、蛋白质组学、生物信息学和药物开发产生重大影响。
Information of subcellular locations of proteins is important for in-depth studies of cell biology. It is very useful for proteomics, system biology and drug development as well. However, most existing methods for predicting protein subcellular location can only cover 5 to 12 location sites. Also, they are limited to deal with single-location proteins and hence failed to work for multiplex proteins, which can simultaneously exist at, or move between, two or more location sites. Actually, multiplex proteins of this kind usually posses some important biological functions worthy of our special notice. A new predictor called “Euk-mPLoc 2.0” is developed by hybridizing the gene ontology information, functional domain information, and sequential evolutionary information through three different modes of pseudo amino acid composition. It can be used to identify eukaryotic proteins among the following 22 locations: (1) acrosome, (2) cell wall, (3) centriole, (4) chloroplast, (5) cyanelle, (6) cytoplasm, (7) cytoskeleton, (8) endoplasmic reticulum, (9) endosome, (10) extracell, (11) Golgi apparatus, (12) hydrogenosome, (13) lysosome, (14) melanosome, (15) microsome (16) mitochondria, (17) nucleus, (18) peroxisome, (19) plasma membrane, (20) plastid, (21) spindle pole body, and (22) vacuole. Compared with the existing methods for predicting eukaryotic protein subcellular localization, the new predictor is much more powerful and flexible, particularly in dealing with proteins with multiple locations and proteins without available accession numbers. For a newly-constructed stringent benchmark dataset which contains both single- and multiple-location proteins and in which none of proteins has pairwise sequence identity to any other in a same location, the overall jackknife success rate achieved by Euk-mPLoc 2.0 is more than 24% higher than those by any of the existing predictors. As a user-friendly web-server, Euk-mPLoc 2.0 is freely accessible at http://www.csbio.sjtu.edu.cn/bioinf/euk-multi-2/. For a query protein sequence of 400 amino acids, it will take about 15 seconds for the web-server to yield the predicted result; the longer the sequence is, the more time it may usually need. It is anticipated that the novel approach and the powerful predictor as presented in this paper will have a significant impact to Molecular Cell Biology, System Biology, Proteomics, Bioinformatics, and Drug Development.