Incorporating Sentiment Prior Knowledge for Weakly Supervised Sentiment Analysis

Incorporating Sentiment Prior Knowledge for Weakly Supervised Sentiment Analysis
复制标题

DOI:
10.1145/2184436.2184437
复制
发表时间:
2012-06
期刊:
ACM Trans. Asian Lang. Inf. Process.
影响因子:
--
通讯作者:
Yulan He
Yulan He
中科院分区:
其他
文献类型:
--
作者:
Yulan He

文献摘要

被引文献

相似文献

本文提出了两种新的方法,将情感先验知识的主题模型的弱监督情感分析的情感标签被认为是主题。一种是通过修改主题词分布的Dirichlet先验(LDA-DP),另一种是通过使用广义期望准则(LDA-GE)增加表达对词典词的情感标签期望的偏好的项来扩充模型目标函数。我们对英文电影评论数据和多领域情感数据集以及关于移动的手机、数码相机、MP3播放器和显示器的中文产品评论进行了广泛的实验。结果表明,虽然LDA-DP和LDA-GE都执行现有的弱监督情感分类算法,但它们更简单,计算效率更高,更适合于Web上的在线和实时情感分类。我们观察到LDA-GE比LDA-DP更有效,这表明在考虑采用主题模型进行情感分析时应该首选LDA-GE。此外,这两个模型都能够从文本中提取高度领域突出的极性词。
This article presents two novel approaches for incorporating sentiment prior knowledge into the topic model for weakly supervised sentiment analysis where sentiment labels are considered as topics. One is by modifying the Dirichlet prior for topic-word distribution (LDA-DP), the other is by augmenting the model objective function through adding terms that express preferences on expectations of sentiment labels of the lexicon words using generalized expectation criteria (LDA-GE). We conducted extensive experiments on English movie review data and multi-domain sentiment dataset as well as Chinese product reviews about mobile phones, digital cameras, MP3 players, and monitors. The results show that while both LDA-DP and LDA-GE perform comparably to existing weakly supervised sentiment classification algorithms, they are much simpler and computationally efficient, rendering them more suitable for online and real-time sentiment classification on the Web. We observed that LDA-GE is more effective than LDA-DP, suggesting that it should be preferred when considering employing the topic model for sentiment analysis. Moreover, both models are able to extract highly domain-salient polarity words from text.