Wikipedia Link Structure and Text Mining for Semantic Relation Extraction

Wikipedia Link Structure and Text Mining for Semantic Relation Extraction
复制标题

DOI:
--
复制
发表时间:
2008
期刊:
--
影响因子:
--
通讯作者:
Kotaro Nakayama;T. Hara;S. Nishio
Kotaro Nakayama;T. Hara;S. Nishio
中科院分区:
其他
文献类型:
--
作者:
Kotaro Nakayama;T. Hara;S. Nishio

文献摘要

被引文献

相似文献

维基百科,一个基于维基百科的合作百科全书,在互联网用户中已经成为一个巨大的现象。它涵盖了艺术,地理,历史,科学,体育和游戏等各个领域的大量概念。由于维基百科正在成为一个存储所有人类知识的数据库,维基百科挖掘是一种很有前途的方法,它将语义网和社交网联系在一起。K. a. Web 2.0)。事实上,在对Wikipedia挖掘的大量研究中,已经充分证明了Wikipedia作为一个知识抽取的语料库,特别是在概念之间的相关性度量方面具有很强的能力。然而,语义相关性只是关系的数值强度,而没有明确的关系类型。为了提取具有外显关系类型的可推断语义关系,我们不仅需要分析维基百科中的链接结构,还需要分析维基百科中的文本。在本文中,我们提出了一个一致的方法从维基百科的语义关系提取。该方法包括三个子过程高度优化维基百科挖掘:1)快速预处理,2)POS(部分语音)标签树分析,和3)主干提取。此外,我们的详细评估证明,链接结构挖掘提高了语义关系提取的准确性和可扩展性。
Wikipedia, a collaborative Wiki-based encyclopedia, has be- come a huge phenomenon among Internet users. It covers huge number of concepts of various fields such as Arts, Geography, History, Science, Sports and Games. Since it is becoming a database storing all human knowledge, Wikipedia mining is a promising approach that bridges the Semantic Web and the Social Web (a. k. a. Web 2.0). In fact, in the previ- ous researches on Wikipedia mining, it is strongly proved that Wikipedia has a remarkable capability as a corpus for knowledge extraction, espe- cially for relatedness measurement among concepts. However, semantic relatedness is just a numerical strength of a relation but does not have an explicit relation type. To extract inferable semantic relations with ex- plicit relation types, we need to analyze not only the link structure but also texts in Wikipedia. In this paper, we propose a consistent approach of semantic relation extraction from Wikipedia. The method consists of three sub-processes highly optimized for Wikipedia mining; 1) fast pre- processing, 2) POS (Part Of Speech) tag tree analysis, and 3) mainstay extraction. Furthermore, our detailed evaluation proved that link struc- ture mining improves both the accuracy and the scalability of semantic relations extraction.