PSSM-based prediction of DNA binding sites in proteins.

PSSM-based prediction of DNA binding sites in proteins.
复制标题

DOI:
10.1186/1471-2105-6-33
复制
发表时间:
2005-02-19
期刊:
影响因子:
3
通讯作者:
Sarai A
Sarai A
中科院分区:
生物学4区
文献类型:
--
作者:
Ahmad S;Sarai A

文献摘要

参考文献

被引文献

相似文献

蛋白质中DNA结合位点的检测对于靶向基因调控和操纵的技术具有巨大的意义。我们以前已经表明,一个残基和它的序列邻居的信息可以用来预测DNA结合候选蛋白质序列。即使没有观察到与先前已知的DNA结合蛋白的序列同源性,这种基于序列的预测方法也是适用的。在这里,我们实现了一个基于神经网络的算法,利用其位置特异性评分矩阵(PSSMs)的氨基酸序列的进化信息,以更好地预测DNA结合位点。使用PSSM的平均灵敏度和特异性比仅使用序列信息的预测高出8.7%。可以使用小得多的数据集来生成PSSM,而预测精度的损失最小。使用PSSM衍生预测的一个问题是获得针对大型序列数据库的冗长且耗时的比对。为了加速生成PSSM的过程,我们尝试使用不同的参考数据集(序列空间),针对这些参考数据集扫描靶蛋白以进行PSI-BLAST迭代。我们发现,一个非常小的蛋白质集实际上可以用作这样的参考数据,而不会损失太多的预测值。这使得产生PSSM的过程非常快速,甚至可以在基因组水平上使用。已经开发了一个网络服务器,以提供任何新蛋白质的氨基酸序列的DNA结合位点的预测。基于此方法的在线预测可在
Detection of DNA-binding sites in proteins is of enormous interest for technologies targeting gene regulation and manipulation. We have previously shown that a residue and its sequence neighbor information can be used to predict DNA-binding candidates in a protein sequence. This sequence-based prediction method is applicable even if no sequence homology with a previously known DNA-binding protein is observed. Here we implement a neural network based algorithm to utilize evolutionary information of amino acid sequences in terms of their position specific scoring matrices (PSSMs) for a better prediction of DNA-binding sites. An average of sensitivity and specificity using PSSMs is up to 8.7% better than the prediction with sequence information only. Much smaller data sets could be used to generate PSSM with minimal loss of prediction accuracy. One problem in using PSSM-derived prediction is obtaining lengthy and time-consuming alignments against large sequence databases. In order to speed up the process of generating PSSMs, we tried to use different reference data sets (sequence space) against which a target protein is scanned for PSI-BLAST iterations. We find that a very small set of proteins can actually be used as such a reference data without losing much of the prediction value. This makes the process of generating PSSMs very rapid and even amenable to be used at a genome level. A web server has been developed to provide these predictions of DNA-binding sites for any new protein from its amino acid sequence. Online predictions based on this method are available at
DOI: 10.1016/j.jmb.2004.05.058
发表时间: 2004-07-30
影响因子: 5.6
作者:
Ahmad, S;Sarai, A
通讯作者: Sarai, A
DOI: 10.1016/s0022-2836(02)00846-x
发表时间: 2002-10-04
影响因子: 5.6
作者:
Selvaraj, S;Kono, H;Sarai, A
通讯作者: Sarai, A
DOI: 10.1006/jmbi.1999.3091
发表时间: 1999-09-17
影响因子: 5.6
作者:
Jones, DT
通讯作者: Jones, DT
DOI: 10.1073/pnas.90.16.7558
发表时间: 1993-08-15
影响因子: 11.1
作者:
ROST, B;SANDER, C
通讯作者: SANDER, C
DOI: 10.1016/s0022-2836(03)00031-7
发表时间: 2003-02-28
影响因子: 5.6
作者:
Stawiski, EW;Gregoret, LM;Mandel-Gutfreund, Y
通讯作者: Mandel-Gutfreund, Y