Automatic Construction of Polarity-Tagged Corpus from HTML Documents

Automatic Construction of Polarity-Tagged Corpus from HTML Documents
复制标题

从 HTML 文档自动构建极性标记语料库

DOI:
--
复制
发表时间:
2006
期刊:
Annual Meeting of the Association for Computational Linguistics
影响因子:
--
通讯作者:
M. Kitsuregawa
M. Kitsuregawa
中科院分区:
--
文献类型:
--
作者:
Nobuhiro Kaji;M. Kitsuregawa

文献摘要

被引文献

相似文献

提出了一种基于HTML文档的极性标注语料库构建方法。该方法的特点是全自动,可应用于任意HTML文档。我们的方法背后的想法是利用某些布局结构和语言模式。通过使用它们,我们可以自动提取表达观点的句子。在我们的实验中,该方法可以构建由126,610个句子组成的语料库。
This paper proposes a novel method of building polarity-tagged corpus from HTML documents. The characteristics of this method is that it is fully automatic and can be applied to arbitrary HTML documents. The idea behind our method is to utilize certain layout structures and linguistic pattern. By using them, we can automatically extract such sentences that express opinion. In our experiment, the method could construct a corpus consisting of 126,610 sentences.