A simple WWW-based method for semantic word class acquisition

A simple WWW-based method for semantic word class acquisition
复制标题

DOI:
10.1075/cilt.292.27shi
复制
发表时间:
2007-12
期刊:
--
影响因子:
--
通讯作者:
Keiji Shinzato;Kentaro Torisawa
Keiji Shinzato;Kentaro Torisawa
中科院分区:
其他
文献类型:
--
作者:
Keiji Shinzato;Kentaro Torisawa

文献摘要

被引文献

相似文献

本章描述了一种从html文档中获取语义词类的简单方法。我们以前表明,在html文档中的itemizations可以包含语义一致的词类。然而,并不是所有的项目都是语义一致的。我们的目标是提供一个简单的方法,从html文档中只提取语义一致的项。我们的新方法可以执行此任务,从现有的搜索引擎获得命中计数2n次的项目组成的n个项目。
This chapter describes a simple method to obtain semantic word classes from html documents. We previously showed that itemizations in html documents can contain semantically coherent word classes. However, not all the itemizations are semantically coherent. Our goal is to provide a simple method to extract only semantically coherent itemizations from html documents. Our new method can perform this task by obtaining hit counts from an existing search engine 2n times for an itemization consisting of n items.