pLoc_bal-mEuk: Predict Subcellular Localization of Eukaryotic Proteins by General PseAAC and Quasi-balancing Training Dataset

pLoc_bal-mEuk: Predict Subcellular Localization of Eukaryotic Proteins by General PseAAC and Quasi-balancing Training Dataset
复制标题

pLoc_bal-mEuk:通过通用 PseAAC 和准平衡训练数据集预测真核蛋白质的亚细胞定位

DOI:
10.2174/1573406415666181218102517
复制
发表时间:
2019-01-01
影响因子:
2.3
通讯作者:
Xiao, Xuan
Xiao, Xuan
中科院分区:
医学4区
文献类型:
--
作者:
Chou, Kuo-Chen;Cheng, Xiang;Xiao, Xuan

文献摘要

被引文献

相似文献

背景/目的:蛋白质亚细胞定位信息对于基础研究和药物开发都至关重要。随着后基因组时代蛋白质序列的爆炸式增长,迫切需要开发出功能强大的生物信息学工具,仅凭序列信息就能及时、有效地确定蛋白质的亚细胞定位。最近,一种称为“pLoc-mEuk”的预测器被开发用于鉴定真核蛋白质的亚细胞定位。它的性能是压倒性的优于其他预测器用于相同的目的,特别是在处理多标记系统,其中许多蛋白质,称为“多重蛋白质”,可能同时发生在两个或更多个亚细胞位置。虽然它确实是一个非常强大的预测器,但肯定需要更多的努力来进一步改进它。这是因为pLoc-mEuk是由一个极度倾斜的数据集训练的,其中一些子集的大小是其他子集的200倍。因此,它不能避免由这样一个不均匀的训练datasets.Methods造成的偏差后果:为了减轻这种偏差,我们已经开发了一个新的预测称为pLoc_bal-mEuk准平衡的训练datasets. Methods。交叉验证测试完全相同的实验证实的数据集表明,所提出的新的预测是显着上级的pLoc-mEuk,现有的国家的最先进的预测在确定真核蛋白的亚细胞定位。结果:为了最大限度地方便大多数实验科学家,一个用户友好的网络服务器已经在www.example.com上建立了新的预测http://www.jci-bioinfo.cn/pLoc_bal-mEuk/Conclusion:预计pLoc_bal-1000 - 10000 - 1000 Euk预测因子在真核生物蛋白质的亚细胞定位研究中具有很高的潜力,特别是寻找多靶点药物是目前药物开发的一个非常热门的趋势。
Background/Objective: Information of protein subcellular localization is crucially important for both basic research and drug development. With the explosive growth of protein sequences discovered in the post-genomic age, it is highly demanded to develop powerful bioinformatics tools for timely and effectively identifying their subcellular localization purely based on the sequence information alone. Recently, a predictor called "pLoc-mEuk" was developed for identifying the subcellular localization of eukaryotic proteins. Its performance is overwhelmingly better than that of the other predictors for the same purpose, particularly in dealing with multi-label systems where many proteins, called "multiplex proteins", may simultaneously occur in two or more subcellular locations. Although it is indeed a very powerful predictor, more efforts are definitely needed to further improve it. This is because pLoc-mEuk was trained by an extremely skewed dataset where some subset was about 200 times the size of the other subsets. Accordingly, it cannot avoid the biased consequence caused by such an uneven training dataset.Methods: To alleviate such bias, we have developed a new predictor called pLoc_bal-mEuk by quasi-balancing the training dataset. Cross-validation tests on exactly the same experiment-confirmed dataset have indicated that the proposed new predictor is remarkably superior to pLoc-mEuk, the existing state-of-the-art predictor in identifying the subcellular localization of eukaryotic proteins. It has not escaped our notice that the quasi-balancing treatment can also be used to deal with many other biological systems.Results: To maximize the convenience for most experimental scientists, a user-friendly web-server for the new predictor has been established at http://www.jci-bioinfo.cn/pLoc_bal-mEuk/Conclusion: It is anticipated that the pLoc_bal-Euk predictor holds very high potential to become a useful high throughput tool in identifying the subcellular localization of eukaryotic proteins, particularly for finding multi-target drugs that is currently a very hot trend trend in drug development.