Hacking Wikipedia for Hyponymy Relation Acquisition

Hacking Wikipedia for Hyponymy Relation Acquisition
复制标题

DOI:
--
复制
发表时间:
2008
期刊:
--
影响因子:
--
通讯作者:
Asuka Sumida;Kentaro Torisawa
Asuka Sumida;Kentaro Torisawa
中科院分区:
其他
文献类型:
--
作者:
Asuka Sumida;Kentaro Torisawa

文献摘要

被引文献

相似文献

本文描述了一种从维基百科中提取大量上下位关系的方法。维基百科的结构比一般的 HTML 文档更加一致,我们可以用简单的方法提取大量的上下位关系。在这项工作中,我们成功地从日文版维基百科中以 75.3% 的精度提取了超过 1.4 × 106 个上下位关系。据我们所知,这是最大的机器可读日语同义词库。本文的主要贡献是一种从维基百科的分层布局中获取下位词的方法。通过使用机器学习技术和模式匹配,我们能够从日语维基百科的层次布局中提取超过 6.3 × 105 的关系,其精度为 76.4%。剩余的上下位关系是通过现有的从定义语句和类别页面中提取关系的方法获得的。这意味着从分层布局中提取几乎使提取的关系数量增加了一倍。
This paper describes a method for extracting a large set of hyponymy relations from Wikipedia. The Wikipedia is much more consistently structured than generic HTML documents, and we can extract a large number of hyponymy relations with simple methods. In this work, we managed to extract more than 1.4 × 106 hyponymy relations with 75.3% precision from the Japanese version of the Wikipedia. To the best of our knowledge, this is the largest machine-readable thesaurus for Japanese. The main contribution of this paper is a method for hyponymy acquisition from hierarchical layouts in Wikipedia. By using a machine learning technique and pattern matching, we were able to extract more than 6.3 × 105 relations from hierarchical layouts in the Japanese Wikipedia, and their precision was 76.4%. The remaining hyponymy relations were acquired by existing methods for extracting relations from definition sentences and category pages. This means that extraction from the hierarchical layouts almost doubled the number of relations extracted.