An improved method for scoring protein-protein interactions using semantic similarity within the gene ontology.

An improved method for scoring protein-protein interactions using semantic similarity within the gene ontology.
复制标题

DOI:
10.1186/1471-2105-11-562
复制
发表时间:
2010-11-15
期刊:
影响因子:
3
通讯作者:
Bader GD
Bader GD
中科院分区:
生物学4区
文献类型:
--
作者:
Jain S;Bader GD

文献摘要

参考文献

被引文献

相似文献

语义相似性度量可用于评估蛋白质-蛋白质相互作用(PPI)的生理相关性。他们使用基因本体论(GO)等注释系统根据蛋白质的功能量化蛋白质之间的相似性。与不相互作用的蛋白质相比,在细胞中相互作用的蛋白质可能处于相似的位置或参与相似的生物过程。因此,在相互作用的蛋白质中,基因功能注释在语义上越相似,相互作用越可能是生理相关的。然而,用于PPI置信度评估的大多数语义相似性度量不考虑GO的细胞位置、分子功能和生物过程本体的不同类别中的术语层次结构的不相等深度,因此可能高估或低估相似性。我们描述了一种改进的算法,拓扑聚类语义相似性(TCSS),计算语义相似性之间的GO条款注释的蛋白质相互作用数据集。我们的算法,认为不同的深度的生物知识表示在GO图的不同分支。中心思想是将GO图划分为子图,并且如果参与蛋白质属于相同子图,则与它们属于不同子图相比,PPI评分更高。TCSS算法比我们评估的其他语义相似性测量技术在区分真假蛋白质相互作用以及与基因表达和蛋白质家族相关性方面的性能更好。我们在我们的酿酒酵母PPI数据集上显示出比Resnik(次佳方法)平均提高4.6倍的F1得分,在我们的智人PPI数据集上使用细胞组分,生物过程和分子功能GO注释提高2倍。
Semantic similarity measures are useful to assess the physiological relevance of protein-protein interactions (PPIs). They quantify similarity between proteins based on their function using annotation systems like the Gene Ontology (GO). Proteins that interact in the cell are likely to be in similar locations or involved in similar biological processes compared to proteins that do not interact. Thus the more semantically similar the gene function annotations are among the interacting proteins, more likely the interaction is physiologically relevant. However, most semantic similarity measures used for PPI confidence assessment do not consider the unequal depth of term hierarchies in different classes of cellular location, molecular function, and biological process ontologies of GO and thus may over-or under-estimate similarity. We describe an improved algorithm, Topological Clustering Semantic Similarity (TCSS), to compute semantic similarity between GO terms annotated to proteins in interaction datasets. Our algorithm, considers unequal depth of biological knowledge representation in different branches of the GO graph. The central idea is to divide the GO graph into sub-graphs and score PPIs higher if participating proteins belong to the same sub-graph as compared to if they belong to different sub-graphs. The TCSS algorithm performs better than other semantic similarity measurement techniques that we evaluated in terms of their performance on distinguishing true from false protein interactions, and correlation with gene expression and protein families. We show an average improvement of 4.6 times the F1 score over Resnik, the next best method, on our Saccharomyces cerevisiae PPI dataset and 2 times on our Homo sapiens PPI dataset using cellular component, biological process and molecular function GO annotations.
DOI: 10.1186/1471-2105-6-100
发表时间: 2005-04-18
期刊: BMC bioinformatics
影响因子: 3
作者:
Patil A;Nakamura H
通讯作者: Nakamura H
DOI: 10.1038/nature04532
发表时间: 2006-03-30
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Aloy, P;Superti-Furga, G
通讯作者: Superti-Furga, G
DOI: 10.1186/1471-2105-7-491
发表时间: 2006-11-07
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Lei, Zhengdeng;Dai, Yang
通讯作者: Dai, Yang
根据异质全基因组数据进行概率蛋白质功能预测。
DOI: 10.1371/journal.pone.0000337
发表时间: 2007-03-28
期刊: PLOS ONE
影响因子: 3.7
作者:
Nariai, Naoki;Kolaczyk, Eric D.;Kasif, Simon
通讯作者: Kasif, Simon
DOI: 10.1186/1471-2105-9-s5-s4
发表时间: 2008-04-29
期刊: BMC bioinformatics
影响因子: 3
作者:
Pesquita C;Faria D;Bastos H;Ferreira AE;Falcão AO;Couto FM
通讯作者: Couto FM