Weakly Supervised Joint Sentiment-Topic Detection from Text

Weakly Supervised Joint Sentiment-Topic Detection from Text
复制标题

DOI:
10.1109/tkde.2011.48
复制
发表时间:
2012-06-01
影响因子:
8.9
通讯作者:
Rueger, Stefan
Rueger, Stefan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Lin, Chenghua;He, Yulan;Rueger, Stefan

文献摘要

被引文献

相似文献

情感分析或观点挖掘旨在使用自动化工具检测文本中表达的观点、态度和感受等主观信息。本文提出了一种基于潜在狄利克雷分配(LDA)的新型概率建模框架,称为联合情感主题(JST)模型,可同时从文本中检测情感和主题。还研究了称为 Reverse-JST 的 JST 模型的重新参数化版本,它是通过反转建模过程中情感和主题生成的顺序而获得的。尽管 JST 相当于没有分层先验的 Reverse-JST,但大量实验表明,当添加情感先验时,JST 的表现始终优于 Reverse-JST。此外,与情感分类的监督方法在转移到其他领域时往往无法产生令人满意的性能不同,JST 的弱监督性质使其高度可移植到其他领域。这通过来自五个不同领域的数据集的实验结果得到了验证,其中 JST 模型甚至在某些数据集中优于现有的半监督方法,尽管没有使用标记文档。而且,JST 检测到的主题和主题情绪确实是连贯且信息丰富的。我们假设 JST 模型能够以开放的方式轻松满足网络大规模情感分析的需求。
Sentiment analysis or opinion mining aims to use automated tools to detect subjective information such as opinions, attitudes, and feelings expressed in text. This paper proposes a novel probabilistic modeling framework called joint sentiment-topic (JST) model based on latent Dirichlet allocation (LDA), which detects sentiment and topic simultaneously from text. A reparameterized version of the JST model called Reverse-JST, obtained by reversing the sequence of sentiment and topic generation in the modeling process, is also studied. Although JST is equivalent to Reverse-JST without a hierarchical prior, extensive experiments show that when sentiment priors are added, JST performs consistently better than Reverse-JST. Besides, unlike supervised approaches to sentiment classification which often fail to produce satisfactory performance when shifting to other domains, the weakly supervised nature of JST makes it highly portable to other domains. This is verified by the experimental results on data sets from five different domains where the JST model even outperforms existing semi-supervised approaches in some of the data sets despite using no labeled documents. Moreover, the topics and topic sentiment detected by JST are indeed coherent and informative. We hypothesize that the JST model can readily meet the demand of large-scale sentiment analysis from the web in an open-ended fashion.