Using structural motif descriptors for sequence-based binding site prediction.

Using structural motif descriptors for sequence-based binding site prediction.
复制标题

使用结构基序描述符进行基于序列的结合位点预测。

DOI:
10.1186/1471-2105-8-s4-s5
复制
发表时间:
2007-05-22
期刊:
影响因子:
3
通讯作者:
Schroeder, Michael
Schroeder, Michael
中科院分区:
生物学4区
文献类型:
--
作者:
Henschel, Andreas;Winter, Christof;Kim, Wan Kyu;Schroeder, Michael

文献摘要

被引文献

相似文献

许多蛋白质序列仍然没有得到很好的注释。蛋白质的功能特征通常通过识别其相互作用伙伴来改进。在这里,我们的目标是利用3D信息在序列水平上预测蛋白质-蛋白质相互作用(PPI)和蛋白质-配体相互作用(PLI)。为此,我们使用机器学习将构成交互站点结构特征的连续片段编译成一个简档隐马尔可夫模型描述符。所得到的描述符集合可用于筛选序列数据库以预测功能位点。我们为740种分类的蛋白质-蛋白质结合位点和3000多个蛋白质-配体结合位点生成描述符。交叉验证表明,三分之二的PPI描述符足够保守和重要,足以用于结合位点识别。我们进一步验证了从文献中提取的230个PPI,其中我们还识别了界面残基。最后,我们测试了ATP情况下的配体结合描述符。从带有Swiss-Prot注释的序列中,我们实现了25%的召回率,准确率为%,而ProSite的P-loop基序识别了相同数量的命中,但代价是错误阳性的数量要高得多(精度:57%)。我们的方法产生了771个命中,准确率为96%,这是以前任何ProSite模式都没有获得的。自动生成的描述符是对已知ProSite/InterPro主题的有用补充。它们用于预测蛋白质-蛋白质以及蛋白质-配体的相互作用以及它们与蛋白质的结合部位残基,在这些蛋白质中,只有序列信息可用。
Many protein sequences are still poorly annotated. Functional characterization of a protein is often improved by the identification of its interaction partners. Here, we aim to predict protein-protein interactions (PPI) and protein-ligand interactions (PLI) on sequence level using 3D information. To this end, we use machine learning to compile sequential segments that constitute structural features of an interaction site into one profile Hidden Markov Model descriptor. The resulting collection of descriptors can be used to screen sequence databases in order to predict functional sites. We generate descriptors for 740 classified types of protein-protein binding sites and for more than 3,000 protein-ligand binding sites. Cross validation reveals that two thirds of the PPI descriptors are sufficiently conserved and significant enough to be used for binding site recognition. We further validate 230 PPIs that were extracted from the literature, where we additionally identify the interface residues. Finally we test ligand-binding descriptors for the case of ATP. From sequences with Swiss-Prot annotation "ATP-binding", we achieve a recall of 25% with a precision of 89%, whereas Prosite's P-loop motif recognizes an equal amount of hits at the expense of a much higher number of false positives (precision: 57%). Our method yields 771 hits with a precision of 96% that were not previously picked up by any Prosite-pattern. The automatically generated descriptors are a useful complement to known Prosite/InterPro motifs. They serve to predict protein-protein as well as protein-ligand interactions along with their binding site residues for proteins where merely sequence information is available.