PLPD: reliable protein localization prediction from imbalanced and overlapped datasets.

PLPD: reliable protein localization prediction from imbalanced and overlapped datasets.
复制标题

DOI:
10.1093/nar/gkl638
复制
发表时间:
2006
影响因子:
14.9
通讯作者:
Lee, Doheon
Lee, Doheon
中科院分区:
生物学2区
文献类型:
--
作者:
Lee, KiYoung;Kim, Dae-Won;Na, DoKyun;Lee, Kwang H.;Lee, Doheon

文献摘要

参考文献

被引文献

相似文献

蛋白质的亚细胞定位是其重要的功能特征之一。由于大规模基因组分析的需要,迫切需要一种自动高效的蛋白质亚细胞定位预测方法。从机器学习的角度来看,蛋白质定位的数据集有几个特点:数据集有太多的类(一个细胞中有10个以上的定位),它是一个多标签数据集(一个蛋白质可能出现在几个不同的亚细胞位置),它太不平衡(每个定位中的蛋白质数量显着不同)。尽管以前已经做了许多工作来预测蛋白质的亚细胞定位,但没有一个能同时有效地处理这些特征。因此,最终需要一种新的蛋白质定位计算方法来获得更可靠的结果。针对这一问题,本文提出了一种基于D-SVDD的蛋白质定位预测器(PLPD),该预测器可以更容易、更准确地预测蛋白质的特定定位。此外,我们引入了三个测量更精确的蛋白质定位预测的评估。作为Huh等人(2003)的实验所得到的各种数据集的结果,所提出的PLPD方法代表了一种不同的方法,其可能对现有方法(例如最近邻方法和判别协变方法)起到补充作用。最后,在使用5184个分类的蛋白质作为训练数据为每个定位找到一个好的边界后,我们预测了138个蛋白质,其亚细胞定位不能通过Huh等人的实验清楚地观察到。
Subcellular localization is one of the key functional characteristics of proteins. An automatic and efficient prediction method for the protein subcellular localization is highly required owing to the need for large-scale genome analysis. From a machine learning point of view, a dataset of protein localization has several characteristics: the dataset has too many classes (there are more than 10 localizations in a cell), it is a multi-label dataset (a protein may occur in several different subcellular locations), and it is too imbalanced (the number of proteins in each localization is remarkably different). Even though many previous works have been done for the prediction of protein subcellular localization, none of them tackles effectively these characteristics at the same time. Thus, a new computational method for protein localization is eventually needed for more reliable outcomes. To address the issue, we present a protein localization predictor based on D-SVDD (PLPD) for the prediction of protein localization, which can find the likelihood of a specific localization of a protein more easily and more correctly. Moreover, we introduce three measurements for the more precise evaluation of a protein localization predictor. As the results of various datasets which are made from the experiments of Huh et al. (2003), the proposed PLPD method represents a different approach that might play a complimentary role to the existing methods, such as Nearest Neighbor method and discriminate covariant method. Finally, after finding a good boundary for each localization using the 5184 classified proteins as training data, we predicted 138 proteins whose subcellular localizations could not be clearly observed by the experiments of Huh et al. (2003).
DOI: 10.1016/j.patcog.2005.03.020
发表时间: 2005-10-01
影响因子: 8
作者:
Lee, K;Kim, DW;Lee, KH
通讯作者: Lee, KH
DOI: 10.1128/ec.4.1.36-45.2005
发表时间: 2005-01-01
期刊: EUKARYOTIC CELL
影响因子: --
作者:
Denis, V;Cyert, MS
通讯作者: Cyert, MS
DOI: 10.1093/protein/12.2.107
发表时间: 1999-02-01
期刊: PROTEIN ENGINEERING
影响因子: --
作者:
Chou, KC;Elrod, DW
通讯作者: Elrod, DW
DOI: 10.1006/mcbr.2001.0285
发表时间: 2000-10-01
期刊: Molecular Cell Biology Research Communications
影响因子: --
作者:
Cai, Yu-Dong;Liu, Xiao-Jun;Chou, Kuo-Chen
通讯作者: Chou, Kuo-Chen
DOI: 10.1093/nar/28.1.45
发表时间: 2000-01-01
影响因子: 14.9
作者:
Bairoch, A;Apweiler, R
通讯作者: Apweiler, R