SeqNLS: nuclear localization signal prediction based on frequent pattern mining and linear motif scoring.

SeqNLS: nuclear localization signal prediction based on frequent pattern mining and linear motif scoring.
复制标题

SeqNLS:基于频繁模式挖掘和线性基序评分的核定位信号预测。

DOI:
10.1371/journal.pone.0076864
复制
发表时间:
2013
期刊:
影响因子:
3.7
通讯作者:
Hu J
Hu J
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Lin JR;Hu J

文献摘要

参考文献

被引文献

相似文献

核定位信号(Nuclear localization signals,NLS)是蛋白质中的一段残基,介导蛋白质进入细胞核。已知NLS具有不同的模式,其中只有有限数量的模式被目前已知的NLS基序所覆盖。在这里,我们提出了一个序列模式挖掘算法SeqNLS有效地识别潜在的NLS模式,而不受限制的NLS的现有知识。提取的频繁序列模式用于预测NLS候选项,然后通过基于预测的序列混乱的线性基序评分方案和基于相对局部保守性(IRLC)的掩蔽来过滤NLS候选项。在新策划的Yeast和Hybrid数据集上的实验结果表明,SeqNLS在检测潜在的NLS方面是有效的。SeqNLS与线性基序评分之间的性能比较表明,线性基序特征与识别NLS的序列特征高度互补。对于两个独立的数据集,我们的SeqNLS不仅可以始终找到超过50%的NLS,预测精度至少为0.7,而且在F1得分或预测精度方面优于其他最先进的NLS预测方法,具有相似或更高的召回率。SeqNLS算法的网络服务器可在http://mleg.cse.sc.edu/seqNLS获得。
Nuclear localization signals (NLSs) are stretches of residues in proteins mediating their importing into the nucleus. NLSs are known to have diverse patterns, of which only a limited number are covered by currently known NLS motifs. Here we propose a sequential pattern mining algorithm SeqNLS to effectively identify potential NLS patterns without being constrained by the limitation of current knowledge of NLSs. The extracted frequent sequential patterns are used to predict NLS candidates which are then filtered by a linear motif-scoring scheme based on predicted sequence disorder and by the relatively local conservation (IRLC) based masking. The experiment results on the newly curated Yeast and Hybrid datasets show that SeqNLS is effective in detecting potential NLSs. The performance comparison between SeqNLS with and without the linear motif scoring shows that linear motif features are highly complementary to sequence features in discerning NLSs. For the two independent datasets, our SeqNLS not only can consistently find over 50% of NLSs with prediction precision of at least 0.7, but also outperforms other state-of-the-art NLS prediction methods in terms of F1 score or prediction precision with similar or higher recall rates. The web server of the SeqNLS algorithm is available at http://mleg.cse.sc.edu/seqNLS.
DOI: 10.1074/jbc.m303275200
发表时间: 2003-07-25
影响因子: 4.8
作者:
Fontes, MRM;Teh, T;Kobe, B
通讯作者: Kobe, B
Slimfinder:一种概率方法,用于识别蛋白质中占代表性过多的短线线性基序。
DOI: 10.1371/journal.pone.0000967
发表时间: 2007-10-03
期刊: PLOS ONE
影响因子: 3.7
作者:
Edwards, Richard J.;Davey, Norman E.;Shields, Denis C.
通讯作者: Shields, Denis C.
DOI: 10.1111/j.1600-0854.2009.01028.x
发表时间: 2010-03
期刊: Traffic (Copenhagen, Denmark)
影响因子: --
作者:
Lange A;McLane LM;Mills RE;Devine SE;Corbett AH
通讯作者: Corbett AH
DOI: 10.1093/nar/gkm363
发表时间: 2007-07
影响因子: 14.9
作者:
Ishida T;Kinoshita K
通讯作者: Kinoshita K
DOI: 10.1039/c1mb05231d
发表时间: 2012-01-01
影响因子: --
作者:
Davey, Norman E.;Van Roey, Kim;Gibson, Toby J.
通讯作者: Gibson, Toby J.