Hacking Wikipedia for Hyponymy Relation Acquisition
Hacking Wikipedia for Hyponymy Relation Acquisition
复制标题
DOI:
--
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
Asuka Sumida;Kentaro Torisawa
中科院分区:
文献类型:
--
作者:
Asuka Sumida;Kentaro Torisawa
This paper describes a method for extracting a large set of hyponymy relations from Wikipedia. The Wikipedia is much more consistently structured than generic HTML documents, and we can extract a large number of hyponymy relations with simple methods. In this work, we managed to extract more than 1.4 × 106 hyponymy relations with 75.3% precision from the Japanese version of the Wikipedia. To the best of our knowledge, this is the largest machine-readable thesaurus for Japanese. The main contribution of this paper is a method for hyponymy acquisition from hierarchical layouts in Wikipedia. By using a machine learning technique and pattern matching, we were able to extract more than 6.3 × 105 relations from hierarchical layouts in the Japanese Wikipedia, and their precision was 76.4%. The remaining hyponymy relations were acquired by existing methods for extracting relations from definition sentences and category pages. This means that extraction from the hierarchical layouts almost doubled the number of relations extracted.