ngLOC: software and web server for predicting protein subcellular localization in prokaryotes and eukaryotes.

ngLOC: software and web server for predicting protein subcellular localization in prokaryotes and eukaryotes.
复制标题

DOI:
10.1186/1756-0500-5-351
复制
发表时间:
2012-07-10
期刊:
影响因子:
1.8
通讯作者:
Guda C
Guda C
中科院分区:
其他
文献类型:
--
作者:
King BR;Vural S;Pandey S;Barteau A;Guda C

文献摘要

被引文献

相似文献

了解蛋白质的亚细胞定位是了解蛋白质整体功能的必要组成部分。在过去的十年中,已经发表了许多计算方法,取得了不同程度的成功。尽管在这一领域发表了大量的方法,但只有一小部分可供研究人员在他们自己的研究中使用。在这些可用的方法中,许多都是有限的,只能预测细胞中的少量细胞器。此外,大多数方法只能预测序列的单个位置,尽管已知真核生物物种中的大部分蛋白质在位置之间穿梭以执行其功能。我们提出了一个基于ngLOC方法预测蛋白质序列亚细胞定位的软件包和web服务器。ngLOC是一种基于n-gram的贝叶斯分类器,用于预测原核生物和真核生物中蛋白质的亚细胞定位。不同物种的总体预测准确率在89.8% ~ 91.4%之间。这个程序可以预测植物和动物物种的11个不同的位置。ngLOC还分别预测了革兰氏阳性和革兰氏阴性细菌数据集上的4个和5个不同的位置。ngLOC是一种通用的方法,可以通过来自各种物种或类别的数据进行训练,用于预测蛋白质亚细胞定位。在GNU GPL下,独立软件可以免费用于学术用途,ngLOC web服务器也可以在http://ngloc.unmc.edu上访问。
Understanding protein subcellular localization is a necessary component toward understanding the overall function of a protein. Numerous computational methods have been published over the past decade, with varying degrees of success. Despite the large number of published methods in this area, only a small fraction of them are available for researchers to use in their own studies. Of those that are available, many are limited by predicting only a small number of organelles in the cell. Additionally, the majority of methods predict only a single location for a sequence, even though it is known that a large fraction of the proteins in eukaryotic species shuttle between locations to carry out their function. We present a software package and a web server for predicting the subcellular localization of protein sequences based on the ngLOC method. ngLOC is an n-gram-based Bayesian classifier that predicts subcellular localization of proteins both in prokaryotes and eukaryotes. The overall prediction accuracy varies from 89.8% to 91.4% across species. This program can predict 11 distinct locations each in plant and animal species. ngLOC also predicts 4 and 5 distinct locations on gram-positive and gram-negative bacterial datasets, respectively. ngLOC is a generic method that can be trained by data from a variety of species or classes for predicting protein subcellular localization. The standalone software is freely available for academic use under GNU GPL, and the ngLOC web server is also accessible at http://ngloc.unmc.edu.