Learning domain ontologies from document warehouses and dedicated web sites

Learning domain ontologies from document warehouses and dedicated web sites
复制标题

DOI:
10.1162/089120104323093276
复制
发表时间:
2004-06-01
影响因子:
9.3
通讯作者:
Velardi, P
Velardi, P
中科院分区:
计算机科学3区
文献类型:
--
作者:
Navigli, R;Velardi, P

文献摘要

被引文献

相似文献

我们提出了一种方法和工具,OntoLearn,旨在从网站中提取领域本体,更普遍的是从虚拟组织成员之间共享的文档中提取领域本体。 OntoLearn 首先从可用文档中提取领域术语。然后,对复杂的领域术语进行语义解释并以分层方式排列。最后,用检测到的领域概念对通用本体 WordNet 进行修剪和丰富。这种方法的主要新颖之处是语义解释,即复杂概念与复杂术语的关联。这涉及到为术语串中的每个单词找到适当的 WordNet 概念以及概念组件之间的适当概念关系。语义解释基于一种新的词义消歧算法,称为结构语义互连。
We present a method and a tool, OntoLearn, aimed at the extraction of domain ontologies from Web sites, and more generally from documents shared among the members of virtual organizations. OntoLearn first extracts a domain terminology from available documents. Then, complex domain terms are semantically interpreted and arranged in a hierarchical fashion. Finally, a general-purpose ontology, WordNet, is trimmed and enriched with the detected domain concepts. The major novel aspect of this approach is semantic interpretation, that is, the association of a complex concept with a complex term. This involves finding the appropriate WordNet concept for each word of a terminological string and the appropriate conceptual relations that hold among the concept components. Semantic interpretation is based on a new word sense disambiguation algorithm, called structural semantic interconnections.