Spotting opinion spammers using behavioral footprints

Spotting opinion spammers using behavioral footprints
复制标题

DOI:
10.1145/2487575.2487580
复制
发表时间:
2013-08
期刊:
Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining
影响因子:
--
通讯作者:
Arjun Mukherjee;Abhinav Kumar;B. Liu;Junhui Wang;M. Hsu;M. Castellanos;Riddhiman Ghosh
Arjun Mukherjee;Abhinav Kumar;B. Liu;Junhui Wang;M. Hsu;M. Castellanos;Riddhiman Ghosh
中科院分区:
其他
文献类型:
--
作者:
Arjun Mukherjee;Abhinav Kumar;B. Liu;Junhui Wang;M. Hsu;M. Castellanos;Riddhiman Ghosh

文献摘要

被引文献

相似文献

诸如产品评论之类的有意见的社交媒体现在被个人和组织广泛用于他们的决策。然而,由于利益或名声的原因,人们试图通过意见垃圾邮件(例如,撰写虚假评论)来推广或降级某些目标产品。近年来,虚假评论检测引起了商业和研究界的极大关注。然而,由于监督学习和评估所需的人工标记的困难,该问题仍然具有高度挑战性。这项工作提出了一个新的角度来解决这个问题,通过建模的潜在的垃圾邮件。提出了一种无监督的作者垃圾邮件模型(ASM)。它在贝叶斯环境中工作,这有助于将作者的垃圾信息建模为潜在的,并允许我们利用评论者的各种观察到的行为足迹。直觉是,意见垃圾邮件发送者与非垃圾邮件发送者具有不同的行为分布。这在两个集群的潜在人口分布之间创建了分布分歧:垃圾邮件发送者和非垃圾邮件发送者。模型推断导致学习两个集群的种群分布。ASM的几个扩展也被认为是利用不同的先验知识。在现实生活中的亚马逊评论数据集上的实验证明了所提出的模型的有效性,其性能明显优于最先进的竞争对手。
Opinionated social media such as product reviews are now widely used by individuals and organizations for their decision making. However, due to the reason of profit or fame, people try to game the system by opinion spamming (e.g., writing fake reviews) to promote or to demote some target products. In recent years, fake review detection has attracted significant attention from both the business and research communities. However, due to the difficulty of human labeling needed for supervised learning and evaluation, the problem remains to be highly challenging. This work proposes a novel angle to the problem by modeling spamicity as latent. An unsupervised model, called Author Spamicity Model (ASM), is proposed. It works in the Bayesian setting, which facilitates modeling spamicity of authors as latent and allows us to exploit various observed behavioral footprints of reviewers. The intuition is that opinion spammers have different behavioral distributions than non-spammers. This creates a distributional divergence between the latent population distributions of two clusters: spammers and non-spammers. Model inference results in learning the population distributions of the two clusters. Several extensions of ASM are also considered leveraging from different priors. Experiments on a real-life Amazon review dataset demonstrate the effectiveness of the proposed models which significantly outperform the state-of-the-art competitors.