Discriminative motif finding for predicting protein subcellular localization.

Discriminative motif finding for predicting protein subcellular localization.
复制标题

DOI:
10.1109/tcbb.2009.82
复制
发表时间:
2011-03
期刊:
IEEE/ACM transactions on computational biology and bioinformatics
影响因子:
--
通讯作者:
Bar-Joseph Z
Bar-Joseph Z
中科院分区:
其他
文献类型:
--
作者:
Lin TH;Murphy RF;Bar-Joseph Z

文献摘要

被引文献

相似文献

从序列信息预测蛋白质亚细胞位置的方法有很多。然而,这些方法大多依赖于全局序列特性或使用一组已知的蛋白质靶向基序来预测蛋白质的定位。在这里,我们开发并测试了一种新的方法,该方法使用基于隐马尔可夫模型(discriminative hmm)的判别方法来识别潜在的目标基序。这些模型通过利用模仿蛋白质分选机制的分层结构来搜索存在于一个隔室中但在其他邻近隔室中不存在的基序。我们表明,判别基序发现和层次结构都提高了酵母蛋白基准数据集的定位预测。所鉴定的基序可以映射到已知的靶向基序,它们比一般的蛋白质序列更保守。利用我们基于基序的预测,我们可以在公共数据库中识别出一些蛋白质位置的潜在注释错误。本文中描述的软件实现和数据集可从http://murphylab.web.cmu.edu/software/2009_TCBB_motif/获得
Many methods have been described to predict the subcellular location of proteins from sequence information. However, most of these methods either rely on global sequence properties or use a set of known protein targeting motifs to predict protein localization. Here we develop and test a novel method that identifies potential targeting motifs using a discriminative approach based on hidden Markov models (discriminative HMMs). These models search for motifs that are present in a compartment but absent in other, nearby, compartments by utilizing an hierarchical structure that mimics the protein sorting mechanism. We show that both discriminative motif finding and the hierarchical structure improves localization prediction on a benchmark dataset of yeast proteins. The motifs identified can be mapped to known targeting motifs and they are more conserved than the average protein sequence. Using our motif-based predictions we can identify potential annotation errors in public databases for the location of some of the proteins. A software implementation and the dataset described in this paper are available from http://murphylab.web.cmu.edu/software/2009_TCBB_motif/