A relation based measure of semantic similarity for Gene Ontology annotations.

A relation based measure of semantic similarity for Gene Ontology annotations.
复制标题

DOI:
10.1186/1471-2105-9-468
复制
发表时间:
2008-11-04
期刊:
影响因子:
3
通讯作者:
Dobson S
Dobson S
中科院分区:
生物学4区
文献类型:
--
作者:
Sheehan B;Quigley A;Gaudin B;Dobson S

文献摘要

参考文献

被引文献

相似文献

生物本体论中术语的语义相似性的各种度量,例如基因本体论(GO),已经被用于比较基因产物。这种相似性的测量已经被用于注释未表征的基因产物和将基因产物分组为功能组。有多种方法来测量语义相似性,或者使用本体的拓扑结构,与术语相关联的实例(基因产物),或者两者的混合。我们专注于语义相似性的实例级定义,同时使用包含在本体中的信息,无论是在本体的图形结构和术语之间的关系的语义,提供我们的实例级描述的约束。术语的语义相似性通过各种方法扩展到注释,无论是通过聚合操作,如最小值,最大值和平均值,或通过外推的方法。这些方法引入了关于术语的语义相似性如何与注释的语义相似性相关的假设,这些假设不一定反映术语如何彼此相关。我们利用的关系在GO的语义构造一个算法,称为SSA,提供了一个框架,自然扩展的基础上的实例的方法的术语的语义相似性,如Resnik的措施,描述注释,而不仅仅是术语的基础。我们的措施试图正确地解释如何通过他们的关系在本体论层次联合收割机。SSA使用这些关系来识别术语之间最具体的共同祖先。我们概述了一组情况下,条款可以联合收割机和关联的偏序约束与每种情况下,为了条款的特异性。这些情况构成了SSA算法的基础。这组相关的约束条件也提供了一组原则,对我们的方法进行任何改进都应该寻求满足这些原则。我们得出一个措施,利用所有可用的信息,而不引入假设的本体或数据的性质之间的语义相似性的注释。我们保留的原则,基于实例的语义相似性的术语在注释级别的方法。因此,我们的措施更好地描述了与基因产物相关的注释中包含的信息,因此更适合于通过它们的注释来表征和分类基因产物。
Various measures of semantic similarity of terms in bio-ontologies such as the Gene Ontology (GO) have been used to compare gene products. Such measures of similarity have been used to annotate uncharacterized gene products and group gene products into functional groups. There are various ways to measure semantic similarity, either using the topological structure of the ontology, the instances (gene products) associated with terms or a mixture of both. We focus on an instance level definition of semantic similarity while using the information contained in the ontology, both in the graphical structure of the ontology and the semantics of relations between terms, to provide constraints on our instance level description. Semantic similarity of terms is extended to annotations by various approaches, either though aggregation operations such as min, max and average or through an extrapolative method. These approaches introduce assumptions about how semantic similarity of terms relates to the semantic similarity of annotations that do not necessarily reflect how terms relate to each other. We exploit the semantics of relations in the GO to construct an algorithm called SSA that provides the basis of a framework that naturally extends instance based methods of semantic similarity of terms, such as Resnik's measure, to describing annotations and not just terms. Our measure attempts to correctly interpret how terms combine via their relationships in the ontological hierarchy. SSA uses these relationships to identify the most specific common ancestors between terms. We outline the set of cases in which terms can combine and associate partial order constraints with each case that order the specificity of terms. These cases form the basis for the SSA algorithm. The set of associated constraints also provide a set of principles that any improvement on our method should seek to satisfy. We derive a measure of semantic similarity between annotations that exploits all available information without introducing assumptions about the nature of the ontology or data. We preserve the principles underlying instance based methods of semantic similarity of terms at the annotation level. As a result our measure better describes the information contained in annotations associated with gene products and as a result is better suited to characterizing and classifying gene products through their annotations.
DOI: 10.1186/1471-2105-7-491
发表时间: 2006-11-07
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Lei, Zhengdeng;Dai, Yang
通讯作者: Dai, Yang
DOI: 10.1109/tcbb.2005.50
发表时间: 2005-10-01
影响因子: 4.5
作者:
Sevilla, JL;Segura, V;Rubio, A
通讯作者: Rubio, A
DOI: 10.2307/431307
发表时间: 1987-09-01
影响因子: 0.8
作者:
ARRELL, D
通讯作者: ARRELL, D
DOI: 10.1109/tcbb.2006.37
发表时间: 2006-07-01
影响因子: 4.5
作者:
Popescu, Mihail;Keller, James M.;Mitchell, Joyce A.
通讯作者: Mitchell, Joyce A.
DOI: 10.1093/nar/26.1.73
发表时间: 1998-01-01
影响因子: 14.9
作者:
Cherry, JM;Adler, C;Botstein, D
通讯作者: Botstein, D