Graphlet kernels for prediction of functional residues in protein structures.

Graphlet kernels for prediction of functional residues in protein structures.
复制标题

DOI:
10.1089/cmb.2009.0029
复制
发表时间:
2010-01
期刊:
Journal of computational biology : a journal of computational molecular cell biology
影响因子:
--
通讯作者:
Radivojac P
Radivojac P
中科院分区:
其他
文献类型:
--
作者:
Vacic V;Iakoucheva LM;Lonardi S;Radivojac P

文献摘要

参考文献

被引文献

相似文献

我们介绍了一种新的基于图的核方法注释蛋白质结构中的功能残基。首先将结构建模为蛋白质接触图,其中节点对应于残基,边缘连接空间相邻的残基。然后,图中的每个顶点被表示为以感兴趣的顶点为中心的标记的非同构子图(小图)的计数的向量。两个顶点之间的相似性度量表示为它们各自的计数向量的内积,并且在监督学习框架中用于对蛋白质残基进行分类。我们在两个功能预测问题上评估了我们的方法:蛋白质中催化残基的识别,这是一个经过充分研究的适合基准测试的问题,以及一个较少探索的预测蛋白质结构中磷酸化位点的问题。然后将graphlet内核方法的性能与两种替代方法进行比较,基于序列的预测器和我们的FEATURE框架的实现。在这两项任务上,graphlet内核表现良好;然而,在磷酸化位点预测问题上的差异幅度要高得多。虽然有数据表明磷酸化位点优先位于内在无序区域,但我们提供的证据表明,对于位于结构化区域的位点,无论是单独的表面可及性还是从FEATURE使用的残留物微环境计算的平均测量值都不足以实现高精度。图表示的主要优点是它能够通过枚举相应标记图中的局部连接模式来捕获蛋白质结构中的邻域相似性。
We introduce a novel graph-based kernel method for annotating functional residues in protein structures. A structure is first modeled as a protein contact graph, where nodes correspond to residues and edges connect spatially neighboring residues. Each vertex in the graph is then represented as a vector of counts of labeled non-isomorphic subgraphs (graphlets), centered on the vertex of interest. A similarity measure between two vertices is expressed as the inner product of their respective count vectors and is used in a supervised learning framework to classify protein residues. We evaluated our method on two function prediction problems: identification of catalytic residues in proteins, which is a well-studied problem suitable for benchmarking, and a much less explored problem of predicting phosphorylation sites in protein structures. The performance of the graphlet kernel approach was then compared against two alternative methods, a sequence-based predictor and our implementation of the FEATURE framework. On both tasks the graphlet kernel performed favorably; however, the margin of difference was considerably higher on the problem of phosphorylation site prediction. While there is data that phosphorylation sites are preferentially positioned in intrinsically disordered regions, we provide evidence that for the sites that are located in structured regions, neither the surface accessibility alone nor the averaged measures calculated from the residue microenvironments utilized by FEATURE were sufficient to achieve high accuracy. The key benefit of the graphlet representation is its ability to capture neighborhood similarities in protein structures via enumerating the patterns of local connectivity in the corresponding labeled graphs.
DOI: 10.1021/ja803143g
发表时间: 2008-09-17
影响因子: 15
作者:
Espinoza-Fonseca, L. Michel;Kast, David;Thomas, David D.
通讯作者: Thomas, David D.
DOI: 10.1006/jmbi.1994.1657
发表时间: 1994-10-21
影响因子: 5.6
作者:
ARTYMIUK, PJ;POIRRETTE, AR;WILLETT, P
通讯作者: WILLETT, P
DOI: 10.1038/nbt0302-301
发表时间: 2002-03-01
影响因子: 46.9
作者:
Ficarro, SB;McCleland, ML;White, FM
通讯作者: White, FM
DOI: 10.1074/mcp.m700564-mcp200
发表时间: 2008-07-01
影响因子: 7
作者:
Collins, Mark O.;Yu, Lu;Choudhary, Jyoti S.
通讯作者: Choudhary, Jyoti S.
DOI: 10.1006/jmbi.2001.5009
发表时间: 2001-09-28
影响因子: 5.6
作者:
Elcock, AH
通讯作者: Elcock, AH