Issues in learning an ontology from text.

Issues in learning an ontology from text.
复制标题

DOI:
10.1186/1471-2105-10-s5-s1
复制
发表时间:
2009-05-06
期刊:
影响因子:
3
通讯作者:
Zhang Z
Zhang Z
中科院分区:
生物学4区
文献类型:
--
作者:
Brewster C;Jupp S;Luciano J;Shotton D;Stevens RD;Zhang Z

文献摘要

被引文献

相似文献

任何领域的本体构建都是一个劳动密集型和复杂的过程。任何能够降低成本和提高效率的方法都有可能对生命科学产生重大影响。本文描述了一个实验中的本体构建从文本的动物行为领域。我们的目标是看看有多少可以做一个简单的和相对快速的方式使用语料库的期刊论文。我们使用了一系列预先存在的文本处理步骤,这里描述了为清理输入、导出一组术语以及在许多层次结构中构建这些术语所做的不同选择。我们描述了一些挑战,特别是集中本体适当的起点的异构语料库。主要使用自动化技术,我们能够构建一个18055术语的本体结构,动物行为术语的召回率为73%,但精度仅为26%。我们能够使用测试术语包含在本体中的有效性的词典-句法模式从新生本体中清除不需要的术语。我们使用相同的技术来测试剩余项之间的包含关系,以将结构添加到我们生成的最初广泛和浅的结构中。所有输出都可在。我们提出了一个系统的方法,为科学领域的本体或结构化词汇建设的初始步骤,需要有限的人力,可以作出贡献的本体学习和维护。该方法是有用的探索一个科学领域,并作为垫脚石正式严谨的本体论。从一个异构的语料库中识别的术语的过滤,专注于那些是本体的主题被确定为在本体学习的研究的主要挑战之一。
Ontology construction for any domain is a labour intensive and complex process. Any methodology that can reduce the cost and increase efficiency has the potential to make a major impact in the life sciences. This paper describes an experiment in ontology construction from text for the animal behaviour domain. Our objective was to see how much could be done in a simple and relatively rapid manner using a corpus of journal papers. We used a sequence of pre-existing text processing steps, and here describe the different choices made to clean the input, to derive a set of terms and to structure those terms in a number of hierarchies. We describe some of the challenges, especially that of focusing the ontology appropriately given a starting point of a heterogeneous corpus. Using mainly automated techniques, we were able to construct an 18055 term ontology-like structure with 73% recall of animal behaviour terms, but a precision of only 26%. We were able to clean unwanted terms from the nascent ontology using lexico-syntactic patterns that tested the validity of term inclusion within the ontology. We used the same technique to test for subsumption relationships between the remaining terms to add structure to the initially broad and shallow structure we generated. All outputs are available at . We present a systematic method for the initial steps of ontology or structured vocabulary construction for scientific domains that requires limited human effort and can make a contribution both to ontology learning and maintenance. The method is useful both for the exploration of a scientific domain and as a stepping stone towards formally rigourous ontologies. The filtering of recognised terms from a heterogeneous corpus to focus upon those that are the topic of the ontology is identified to be one of the main challenges for research in ontology learning.