Predicting DNA-binding sites of proteins from amino acid sequence.

Predicting DNA-binding sites of proteins from amino acid sequence.
复制标题

DOI:
10.1186/1471-2105-7-262
复制
发表时间:
2006-05-19
期刊:
影响因子:
3
通讯作者:
--
中科院分区:
生物学4区
文献类型:
--
作者:

文献摘要

被引文献

相似文献

了解蛋白质-DNA相互作用的分子细节对于破译基因调控机制至关重要。我们提出了一种机器学习方法来识别参与蛋白质-DNA相互作用的氨基酸残基。我们从训练朴素贝叶斯分类器开始,根据给定氨基酸残基的同一性及其序列邻居的同一性来预测该氨基酸残基是否为DNA结合残基。分类器的输入由目标残基和目标残基两侧的4个序列邻居的身份组成。分类器在171个蛋白质的非冗余集合上进行训练和评估(使用留一交叉验证)。结果表明,基于局部序列信息识别界面残基是可行的。通过留一法交叉验证,该分类器在识别界面残基方面的总体准确率为71%,相关系数为0.24,特异度为35%,敏感度为53%。我们表明,通过使用目标残基的序列熵(通过将目标序列与其序列同源序列进行比对而获得的多次比对中对应列的熵)作为额外输入,分类器的性能得到了改善。该分类器识别界面残留物的总体准确率为78%,相关系数为0.28,特异度为44%,灵敏度为41%。在蛋白质三维结构背景下的预测检验表明,这种方法在从序列信息中识别DNA结合位点方面是有效的。在33%(171个蛋白质中的56个)蛋白质中,分类器通过正确识别至少一半的界面残基来识别相互作用位点。在87%(171个蛋白质中的149个)蛋白质中,分类器正确识别至少20%的界面残基。这表明有可能使用这种分类器来识别潜在的DNA结合基序,并获得对蛋白质-DNA相互作用的序列相关性的潜在有用的见解。利用序列信息识别DNA结合残基的朴素贝叶斯分类器为识别DNA结合蛋白中可能的DNA结合位点和识别潜在的DNA结合基序提供了一种计算高效的方法。
Understanding the molecular details of protein-DNA interactions is critical for deciphering the mechanisms of gene regulation. We present a machine learning approach for the identification of amino acid residues involved in protein-DNA interactions. We start with a Naïve Bayes classifier trained to predict whether a given amino acid residue is a DNA-binding residue based on its identity and the identities of its sequence neighbors. The input to the classifier consists of the identities of the target residue and 4 sequence neighbors on each side of the target residue. The classifier is trained and evaluated (using leave-one-out cross-validation) on a non-redundant set of 171 proteins. Our results indicate the feasibility of identifying interface residues based on local sequence information. The classifier achieves 71% overall accuracy with a correlation coefficient of 0.24, 35% specificity and 53% sensitivity in identifying interface residues as evaluated by leave-one-out cross-validation. We show that the performance of the classifier is improved by using sequence entropy of the target residue (the entropy of the corresponding column in multiple alignment obtained by aligning the target sequence with its sequence homologs) as additional input. The classifier achieves 78% overall accuracy with a correlation coefficient of 0.28, 44% specificity and 41% sensitivity in identifying interface residues. Examination of the predictions in the context of 3-dimensional structures of proteins demonstrates the effectiveness of this method in identifying DNA-binding sites from sequence information. In 33% (56 out of 171) of the proteins, the classifier identifies the interaction sites by correctly recognizing at least half of the interface residues. In 87% (149 out of 171) of the proteins, the classifier correctly identifies at least 20% of the interface residues. This suggests the possibility of using such classifiers to identify potential DNA-binding motifs and to gain potentially useful insights into sequence correlates of protein-DNA interactions. Naïve Bayes classifiers trained to identify DNA-binding residues using sequence information offer a computationally efficient approach to identifying putative DNA-binding sites in DNA-binding proteins and recognizing potential DNA-binding motifs.