Comparison of feature-level learning methods for mining online consumer reviews

Comparison of feature-level learning methods for mining online consumer reviews
复制标题

DOI:
10.1016/j.eswa.2012.02.158
复制
发表时间:
2012-08
期刊:
Expert Syst. Appl.
影响因子:
--
通讯作者:
Li Chen;Luole Qi;Feng Wang
Li Chen;Luole Qi;Feng Wang
中科院分区:
其他
文献类型:
--
作者:
Li Chen;Luole Qi;Feng Wang

文献摘要

被引文献

相似文献

特征级意见挖掘的任务通常包括从消费者评论中提取产品实体,识别与这些实体相关联的意见词,以及确定这些意见的极性(例如,正面、负面或中性)。近年来,人们提出了两种主要的方法来确定特征层面的意见:一种是基于模型的方法,如基于词汇化隐马尔可夫模型的方法(L-HMM);另一种是统计方法,如基于关联规则挖掘技术。然而,很少有工作比较这些算法在识别各种类型的评论元素方面的实际能力,如特征、意见、强化词、实体短语和不常见的实体。另一方面,很少有人注意使用更具区别性的学习模型来完成这些意见挖掘任务。在本文中,我们不仅在真实评论数据集上对这些方法进行了实验比较,而且特别采用了条件随机场(CRFS)模型,并与相关算法进行了比较。此外,对于基于CRFS的挖掘算法,我们测试了自标注过程在两种自动训练条件下的作用,并进一步确定了理想的学习函数组合,以优化其学习性能。对比实验最终表明,与其他方法相比,基于CRFS的方法在挖掘多个评论元素方面具有更好的准确性。
The tasks of feature-level opinion mining usually include the extraction of product entities from consumer reviews, the identification of opinion words that are associated with the entities, and the determining of these opinions’ polarities (e.g., positive, negative, or neutral). In recent years, two major approaches have been proposed to determine opinions at the feature level: model based methods such as the one based on lexicalized Hidden Markov Model (L-HMMs), and statistical methods like the association rule mining based technique. However, little work has compared these algorithms regarding their practical abilities in identifying various types of review elements, such as features, opinions, intensifiers, entity phrases and infrequent entities. On the other hand, little attentions has been paid to applying more discriminative learning models to accomplish these opinion mining tasks. In this paper, we not only experimentally compared these methods based on a real-world review dataset, but also in particular adopted the Conditional Random Fields (CRFs) model and evaluated its performance in comparison with related algorithms. Moreover, for CRFs-based mining algorithm, we tested the role of a self-tagging process in two automatic training conditions, and further identified the ideal combination of learning functions to optimize its learning performance. The comparative experiment eventually revealed the CRFs-based method’s outperforming accuracy in terms of mining multiple review elements, relative to other methods.