Kernel methods for predicting protein-protein interactions

Kernel methods for predicting protein-protein interactions
复制标题

DOI:
10.1093/bioinformatics/bti1016
复制
发表时间:
2005-06-01
期刊:
影响因子:
5.8
通讯作者:
Noble, WS
Noble, WS
中科院分区:
生物学3区
文献类型:
--
作者:
Ben-Hur, A;Noble, WS

文献摘要

被引文献

相似文献

动机:尽管高通量方法在发现蛋白质-蛋白质相互作用方面取得了进展,但即使是研究得很充分的模型生物的相互作用网络也充其量是粗略的,这突显了继续需要计算方法来帮助实验者寻找新的相互作用。结果:我们提出了一种结合使用数据源的核心方法来预测蛋白质-蛋白质相互作用,包括蛋白质序列、基因本体注释、网络的局部属性和其他物种中的同源相互作用。虽然文献中提出的蛋白质核提供了单个蛋白质之间的相似性,但预测相互作用需要一对蛋白质之间的核。我们提出了一种将单个蛋白质之间的核转换为蛋白质对之间的核的成对核,并结合支持向量机分类器说明了该核的有效性。此外,我们通过组合几个基于k-mer频率、基序和结构域内容的基于序列的核,并使用基于其他数据源的特征进一步增强成对序列核,从而获得了更好的性能。我们使用BIND数据库中的数据来预测酵母中的物理相互作用。在1%的假阳性率下,分类器检索到一组可信交互的近80%。因此,我们展示了我们的方法做出准确预测的能力,尽管在交互数据库中已知存在相当大比例的假阳性。
Motivation: Despite advances in high-throughput methods for discovering protein-protein interactions, the interaction networks of even well-studied model organisms are sketchy at best, highlighting the continued need for computational methods to help direct experimentalists in the search for novel interactions.Results: We present a kernel method for predicting protein-protein interactions using a combination of data sources, including protein sequences, Gene Ontology annotations, local properties of the network, and homologous interactions in other species. Whereas protein kernels proposed in the literature provide a similarity between single proteins, prediction of interactions requires a kernel between pairs of proteins. We propose a pairwise kernel that converts a kernel between single proteins into a kernel between pairs of proteins, and we illustrate the kernel's effectiveness in conjunction with a support vector machine classifier. Furthermore, we obtain improved performance by combining several sequence-based kernels based on k-mer frequency, motif and domain content and by further augmenting the pairwise sequence kernel with features that are based on other sources of data.We apply our method to predict physical interactions in yeast using data from the BIND database. At a false positive rate of 1% the classifier retrieves close to 80% of a set of trusted interactions. We thus demonstrate the ability of our method to make accurate predictions despite the sizeable fraction of false positives that are known to exist in interaction databases.