A simple WWW-based method for semantic word class acquisition
A simple WWW-based method for semantic word class acquisition
复制标题
DOI:
10.1075/cilt.292.27shi
复制
发表时间:
2007-12
期刊:
影响因子:
--
通讯作者:
Keiji Shinzato;Kentaro Torisawa
中科院分区:
文献类型:
--
作者:
Keiji Shinzato;Kentaro Torisawa
This chapter describes a simple method to obtain semantic word classes from html documents. We previously showed that itemizations in html documents can contain semantically coherent word classes. However, not all the itemizations are semantically coherent. Our goal is to provide a simple method to extract only semantically coherent itemizations from html documents. Our new method can perform this task by obtaining hit counts from an existing search engine 2n times for an itemization consisting of n items.