Learning to predict protein-protein interactions from protein sequences

Learning to predict protein-protein interactions from protein sequences
复制标题

DOI:
10.1093/bioinformatics/btg352
复制
发表时间:
2003-10-12
期刊:
影响因子:
5.8
通讯作者:
Rzhetsky, A
Rzhetsky, A
中科院分区:
生物学3区
文献类型:
--
作者:
Gomez, SM;Noble, WS;Rzhetsky, A

文献摘要

被引文献

相似文献

为了了解细胞的分子机制,我们需要了解使细胞发挥功能的大量蛋白质-蛋白质相互作用。高吞吐量技术提供了一些关于这些相互作用的数据,但到目前为止,这些数据相当嘈杂。因此,预测蛋白质-蛋白质相互作用的计算技术可能具有重要价值。在计算机中预测相互作用的一种方法是从第一性原理中产生一个候选相互作用的详细模型。我们采用另一种方法,采用一个相对简单的模型,从大量数据中动态学习。在这项工作中,我们描述了一个吸引-排斥模型,其中一对蛋白质之间的相互作用被表示为与每个蛋白质长度上的小的,结构域或基序大小的特征相关的吸引力和排斥力的总和。该模型具有辨别性,可以同时从已知的相互作用和已知(或怀疑)不相互作用的蛋白质对中学习。该模型计算效率高,并且可以很好地扩展到非常大的数据集。在使用已知酵母相互作用的交叉验证比较中,吸引-排斥方法比几种竞争技术表现得更好。
In order to understand the molecular machinery of the cell, we need to know about the multitude of protein-protein interactions that allow the cell to function. High-throughput technologies provide some data about these interactions, but so far that data is fairly noisy. Therefore, computational techniques for predicting protein-protein interactions could be of significant value. One approach to predicting interactions in silico is to produce from first principles a detailed model of a candidate interaction. We take an alternative approach, employing a relatively simple model that learns dynamically from a large collection of data. In this work, we describe an attraction-repulsion model, in which the interaction between a pair of proteins is represented as the sum of attractive and repulsive forces associated with small, domain- or motif-sized features along the length of each protein. The model is discriminative, learning simultaneously from known interactions and from pairs of proteins that are known (or suspected) not to interact. The model is efficient to compute and scales well to very large collections of data. In a cross-validated comparison using known yeast interactions, the attraction-repulsion method performs better than several competing techniques.