Automatic Construction of Polarity-Tagged Corpus from HTML Documents
Automatic Construction of Polarity-Tagged Corpus from HTML Documents
复制标题
从 HTML 文档自动构建极性标记语料库
DOI:
--
复制
发表时间:
2006
期刊:
影响因子:
--
通讯作者:
M. Kitsuregawa
中科院分区:
文献类型:
--
作者:
Nobuhiro Kaji;M. Kitsuregawa
This paper proposes a novel method of building polarity-tagged corpus from HTML documents. The characteristics of this method is that it is fully automatic and can be applied to arbitrary HTML documents. The idea behind our method is to utilize certain layout structures and linguistic pattern. By using them, we can automatically extract such sentences that express opinion. In our experiment, the method could construct a corpus consisting of 126,610 sentences.