A unified latent variable model for contrastive opinion mining

A unified latent variable model for contrastive opinion mining
复制标题

DOI:
10.1007/s11704-018-7073-5
复制
发表时间:
2019-08
影响因子:
4.2
通讯作者:
Ebuka Ibeke;Chenghua Lin;A. Wyner;M. Barawi
Ebuka Ibeke;Chenghua Lin;A. Wyner;M. Barawi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ebuka Ibeke;Chenghua Lin;A. Wyner;M. Barawi

文献摘要

相似文献

有大量且不断增长的语料库,人们在语料库中对同一主题表达对比意见。这导致了越来越多关于对比意见挖掘的研究。然而,现有的研究中有几个值得注意的问题。它们主要集中于从多个数据集合中挖掘对比意见,这些数据集合需要预先划分到各自的集合中。此外,现有的模型在提取的主题与语料库中表达主题的句子之间的关系方面是不透明的;这种不透明性无助于我们理解语料库中表达的观点。最后,对比意见主要是定性的,而不是定量的。针对这些问题,本文提出了一种新的统一潜在变量模型,该模型从单个和多个数据集合中挖掘对比意见,提取反映对比意见的句子,并测量对提取的主题的意见对比强度。实验结果表明,该模型在挖掘对比观点方面是有效的,在提取连贯的、信息丰富的情感主题方面优于我们的基线。我们进一步证明了我们的模型在文本数据的主题和情感分类方面的准确性,并将我们的结果与五条强基线进行了比较。
There are large and growing textual corpora in which people express contrastive opinions about the same topic. This has led to an increasing number of studies about contrastive opinion mining. However, there are several notable issues with the existing studies. They mostly focus on mining contrastive opinions from multiple data collections, which need to be separated into their respective collections beforehand. In addition, existing models are opaque in terms of the relationship between topics that are extracted and the sentences in the corpus which express the topics; this opacity does not help us understand the opinions expressed in the corpus. Finally, contrastive opinion is mostly analysed qualitatively rather than quantitatively. This paper addresses these matters and proposes a novel unified latent variable model (contraLDA), which: mines contrastive opinions from both single and multiple data collections, extracts the sentences that project the contrastive opinion, and measures the strength of opinion contrastiveness towards the extracted topics. Experimental results show the effectiveness of our model in mining contrasted opinions, which outperformed our baselines in extracting coherent and informative sentiment-bearing topics. We further show the accuracy of our model in classifying topics and sentiments of textual data, and we compared our results to five strong baselines.