Motif-Based Hyponym Relation Extraction from Wikipedia Hyperlinks

Motif-Based Hyponym Relation Extraction from Wikipedia Hyperlinks
复制标题

DOI:
10.1109/tkde.2013.183
复制
发表时间:
2014-10
影响因子:
8.9
通讯作者:
Bifan Wei;Jun Liu;Jian Ma;Q. Zheng;Wei Zhang;B. Feng
Bifan Wei;Jun Liu;Jian Ma;Q. Zheng;Wei Zhang;B. Feng
中科院分区:
计算机科学2区
文献类型:
--
作者:
Bifan Wei;Jun Liu;Jian Ma;Q. Zheng;Wei Zhang;B. Feng

文献摘要

被引文献

相似文献

发现领域术语之间的上下义关系是分类学学习和知识获取的一项基本任务。然而,各种领域语料库的巨大差异和缺乏标记的训练集使得这个任务对于基于文本内容的传统方法非常具有挑战性。维基百科的文章页面的超链接结构中发现,在这项研究中,包含反复出现的网络图案,表明超链接的概率是一个下位词超链接。为此,提出了一种基于维基百科超链接网络模体的上下义关系抽取方法。该方法从域的超链接结构中自动构造基于模体的特征,每个超链接映射为基于13种三节点模体的13维特征向量。该方法从Wikipedia中提取结构信息,并以启发式方式创建标记的训练集。从训练集中确定分类模型,用于下义关系提取。基于从维基百科获得的七个特定领域的数据集进行了两个实验来验证我们的方法。第一个实验,使用手动标记的数据,验证了基于模体的功能的有效性。第二个实验使用了自动标注的不同领域的训练集,实验结果表明,该方法的性能优于基于词汇句法模式的方法,并取得了与基于文本特征的方法相当的结果。实验结果表明,该方法具有较好的实用性和领域可扩展性。
Discovering hyponym relations among domain-specific terms is a fundamental task in taxonomy learning and knowledge acquisition. However, the great diversity of various domain corpora and the lack of labeled training sets make this task very challenging for conventional methods that are based on text content. The hyperlink structure of Wikipedia article pages was found to contain recurring network motifs in this study, indicating the probability of a hyperlink being a hyponym hyperlink. Hence, a novel hyponym relation extraction approach based on the network motifs of Wikipedia hyperlinks was proposed. This approach automatically constructs motif-based features from the hyperlink structure of a domain; every hyperlink is mapped to a 13-dimensional feature vector based on the 13 types of three-node motifs. The approach extracts structural information from Wikipedia and heuristically creates a labeled training set. Classification models were determined from the training sets for hyponym relation extraction. Two experiments were conducted to validate our approach based on seven domain-specific datasets obtained from Wikipedia. The first experiment, which utilized manually labeled data, verified the effectiveness of the motif-based features. The second experiment, which utilized an automatically labeled training set of different domains, showed that the proposed approach performs better than the approach based on lexico-syntactic patterns and achieves comparable result to the approach based on textual features. Experimental results show the practicability and fairly good domain scalability of the proposed approach.