FEDRR: fast, exhaustive detection of redundant hierarchical relations for quality improvement of large biomedical ontologies.

FEDRR: fast, exhaustive detection of redundant hierarchical relations for quality improvement of large biomedical ontologies.
复制标题

DOI:
10.1186/s13040-016-0110-8
复制
发表时间:
2016
期刊:
影响因子:
4.5
通讯作者:
Cui L
Cui L
中科院分区:
生物学3区
文献类型:
--
作者:
Xing G;Zhang GQ;Cui L

文献摘要

参考文献

被引文献

相似文献

冗余层次关系是指从一个概念到另一个概念的两条路径的模式,一条路径的长度为一(直接),另一条路径的长度大于一(间接)。每个冗余关系都代表一个可能的非预期缺陷,需要在本体质量保证过程中进行纠正。检测和消除冗余关系将有助于改善所有依赖相关本体系统作为知识源的方法的结果,例如概念之间语义距离的计算以及本体匹配和对齐。本文介绍了一种新颖且可扩展的方法,称为 FEDRR(快速、详尽的冗余关系检测),用于本体演化过程中的质量保证工作。 FEDRR结合了动态规划和拓扑排序的算法思想,在O(c·|V|+|E|)时间内穷举挖掘本体层次结构中的所有冗余层次关系,其中|V|是概念的数量,|E|是关系的数量,c 在实践中是一个常数。使用 FEDRR,我们对生物医学中两个最大的本体系统:SNOMED CT 和基因本体 (GO) 中的所有冗余 is-a 关系进行了详尽的搜索。在2015-09-01版本的SNOMED CT和2015-05-01版本的GO中分别发现了372和1609个冗余is-a关系。我们还对 UMLS(生物医学本体的大型集成存储库)中的 190 多个源词汇表进行了 FEDRR,并确定了包含冗余 is-a 关系的 6 个源。随机生成的本体也被用来进一步验证 FEDRR 的效率。 FEDRR 提供了一种普遍适用的、有效的工具,用于系统地检测大型本体系统中的冗余关系,以提高质量。
Redundant hierarchical relations refer to such patterns as two paths from one concept to another, one with length one (direct) and the other with length greater than one (indirect). Each redundant relation represents a possibly unintended defect that needs to be corrected in the ontology quality assurance process. Detecting and eliminating redundant relations would help improve the results of all methods relying on the relevant ontological systems as knowledge source, such as the computation of semantic distance between concepts and for ontology matching and alignment. This paper introduces a novel and scalable approach, called FEDRR – Fast, Exhaustive Detection of Redundant Relations – for quality assurance work during ontological evolution. FEDRR combines the algorithm ideas of Dynamic Programming with Topological Sort, for exhaustive mining of all redundant hierarchical relations in ontological hierarchies, in O(c·|V|+|E|) time, where |V| is the number of concepts, |E| is the number of the relations, and c is a constant in practice. Using FEDRR, we performed exhaustive search of all redundant is-a relations in two of the largest ontological systems in biomedicine: SNOMED CT and Gene Ontology (GO). 372 and 1609 redundant is-a relations were found in the 2015-09-01 version of SNOMED CT and 2015-05-01 version of GO, respectively. We have also performed FEDRR on over 190 source vocabularies in the UMLS - a large integrated repository of biomedical ontologies, and identified six sources containing redundant is-a relations. Randomly generated ontologies have also been used to further validate the efficiency of FEDRR. FEDRR provides a generally applicable, effective tool for systematic detecting redundant relations in large ontological systems for quality improvement.
DOI: 10.1109/bigdata.2014.7004301
发表时间: 2014-10
期刊: Proceedings : ... IEEE International Conference on Big Data. IEEE International Conference on Big Data
影响因子: --
作者:
Zhang GQ;Zhu W;Sun M;Tao S;Bodenreider O;Cui L
通讯作者: Cui L
DOI: 10.1093/nar/gkh036
发表时间: 2004-01-01
影响因子: 14.9
作者:
Harris, MA;Clark, J;White, R
通讯作者: White, R
DOI: 10.3233/sw-2011-0036
发表时间: 2012-01-01
期刊: SEMANTIC WEB
影响因子: 3
作者:
Giunchiglia, Fausto;Autayeu, Aliaksandr;Pane, Juan
通讯作者: Pane, Juan
DOI: 10.1186/2041-1480-2-6
发表时间: 2011-09-13
影响因子: 1.9
作者:
Kirsten T;Gross A;Hartung M;Rahm E
通讯作者: Rahm E
DOI: 10.1093/nar/gkh061
发表时间: 2004-01-01
影响因子: 14.9
作者:
Bodenreider, O
通讯作者: Bodenreider, O