Fuzzy measures on the gene ontology for gene product similarity

Fuzzy measures on the gene ontology for gene product similarity
复制标题

DOI:
10.1109/tcbb.2006.37
复制
发表时间:
2006-07-01
影响因子:
4.5
通讯作者:
Mitchell, Joyce A.
Mitchell, Joyce A.
中科院分区:
工程技术3区
文献类型:
--
作者:
Popescu, Mihail;Keller, James M.;Mitchell, Joyce A.

文献摘要

被引文献

相似文献

生物信息学中最重要的对象之一是基因产物(蛋白质或RNA)。对于许多基因产物,功能信息总结在一组基因本体(GO)注释中。对于这些基因,包括基于GO或其他分类法中发现的术语的相似性度量是合理的。在本文中,我们介绍了几种新的措施,计算两个基因产品的相似性与GO条款。模糊度量相似性(FMS)的优点是,它考虑了上下文的两个完整的注释条款时,计算两个基因产品之间的相似性。当两个基因产物没有被注释的共同的分类术语,我们提出了一种方法,避免了零相似性的结果。为了解释注释可靠性的变化,我们提出了一个基于Choquet积分的相似性度量。这些相似性度量为生物学家寻找基因产物的功能信息提供了额外的工具。对代表三个蛋白质家族的一组194个序列的初始测试显示FMS和Choquet相似性与BLAST序列相似性的相关性高于传统的相似性度量,例如成对平均或成对最大。
One of the most important objects in bioinformatics is a gene product (protein or RNA). For many gene products, functional information is summarized in a set of Gene Ontology (GO) annotations. For these genes, it is reasonable to include similarity measures based on the terms found in the GO or other taxonomy. In this paper, we introduce several novel measures for computing the similarity of two gene products annotated with GO terms. The fuzzy measure similarity (FMS) has the advantage that it takes into consideration the context of both complete sets of annotation terms when computing the similarity between two gene products. When the two gene products are not annotated by common taxonomy terms, we propose a method that avoids a zero similarity result. To account for the variations in the annotation reliability, we propose a similarity measure based on the Choquet integral. These similarity measures provide extra tools for the biologist in search of functional information for gene products. The initial testing on a group of 194 sequences representing three proteins families shows a higher correlation of the FMS and Choquet similarities to the BLAST sequence similarities than the traditional similarity measures such as pairwise average or pairwise maximum.