Investigating semantic similarity measures across the Gene Ontology: the relationship between sequence and annotation

Investigating semantic similarity measures across the Gene Ontology: the relationship between sequence and annotation
复制标题

DOI:
10.1093/bioinformatics/btg153
复制
发表时间:
2003-07-01
期刊:
影响因子:
5.8
通讯作者:
Goble, CA
Goble, CA
中科院分区:
生物学3区
文献类型:
--
作者:
Lord, PW;Stevens, RD;Goble, CA

文献摘要

被引文献

相似文献

动机:许多生物信息学数据资源不仅以序列的形式保存数据,而且还作为注释。在大多数情况下,注释被编写为科学的自然语言:这适合人类,但对于机器处理并不是特别有用。本体论提供了一种机制,通过该机制可以用能够进行这种处理的形式来表示知识。在这篇文章中,我们研究了使用本体标注来度量数据资源中条目之间的知识内容相似性或“语义相似性”。这允许生物信息学家以类似于在序列上执行的方式对注释执行相似性测量。生物信息学资源知识成分的语义相似性度量应该为生物学家提供一个新的分析工具。结果:通过与序列相似性的比较,我们给出了实验结果,以考察语义相似性的有效性。我们展示了一个简单的扩展,它允许对序列数据库中保存的知识进行语义搜索。
Motivation: Many bioinformatics data resources not only hold data in the form of sequences, but also as annotation. In the majority of cases, annotation is written as scientific natural language: this is suitable for humans, but not particularly useful for machine processing. Ontologies offer a mechanism by which knowledge can be represented in a form capable of such processing. In this paper we investigate the use of ontological annotation to measure the similarities in knowledge content or 'semantic similarity' between entries in a data resource. These allow a bioinformatician to perform a similarity measure over annotation in an analogous manner to those performed over sequences. A measure of semantic similarity for the knowledge component of bioinformatics resources should afford a biologist a new tool in their repetoire of analyses.Results: We present the results from experiments that investigate the validity of using semantic similarity by comparison with sequence similarity. We show a simple extension that enables a semantic search of the knowledge held within sequence databases.