Identifying ground truth in opinion spam: an empirical survey based on review psychology

Identifying ground truth in opinion spam: an empirical survey based on review psychology
复制标题

识别垃圾评论中的基本事实:基于评论心理学的实证调查

DOI:
10.1007/s10489-020-01764-7
复制
发表时间:
2020
影响因子:
5.3
通讯作者:
Dingyu Yang
Dingyu Yang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ji;un Li;Xiaogang Wang;Liu Yang;Pengpeng Zhang;Dingyu Yang

文献摘要

被引文献

相似文献

由于垃圾评论的危害性很大,尤其是涉及不真实评论的评论,在过去的十年里引起了人们的极大关注。然而,缺乏注释,即基本真理问题,仍然是关键的挑战。这很困难,因为垃圾邮件发送者总是故意伪造他们的评论,即使是领域专家也无法区分这些评论。考虑到垃圾邮件发送者的明显意图,即提升或降低一项商品的声誉,存在通过考虑人群心理来标记它们的机会。到目前为止,一些研究已经应用、验证和提出了有用的证据,包括先前的、经验的、启发式的和模拟的伪真理。在本文中,在调查了真实和欺骗性评论者的不同动机后,我们通过考虑众包和专家垃圾邮件发送者这两个经典角色来调查最新的真相。对于每个角色,突出显示了与垃圾邮件攻击相关的几个主题,无论是否进行伪装,以及可能的离群值。对比分析得出了一些有趣的结论:1)专业垃圾邮件发送者的数据比众包垃圾邮件发送者的数据更具挑战性,可靠性也更低;2)大多数语言证据的可靠性不如行为足迹;3)异常活动与垃圾邮件目标一样值得信任,但它们不需要任何额外的支持,如用户配置文件;4)需要可接受努力的最可靠事实是偏差、突发性、分组垃圾邮件、超出阈值、审查分布、意见比例和垃圾邮件成本。此外,我们还介绍了未来研究的几个有前途的方向。总体而言,这项调查可能会揭示出新的角度,可以用来理解评论垃圾邮件,并提高任何反垃圾邮件平台的性能。
Because it is very harmful, opinion spam, especially that involving untruthful reviews, has attracted much attention in the last decade. However, the lack of annotations, i.e., the ground truth problem, still serves as the key challenge. It is difficult because spammers always deliberately forge their reviews, which cannot be distinguished even by field experts. Considering the obvious intention of spammers, i.e., to promote or demote an items reputation, the opportunity exists to label them by considering crowd psychology. To date, several studies have applied, verified, and presented helpful evidence, including prior, empirical, heuristic, and simulative pseudo truths. In this paper, after investigating both authentic and deceptive reviewers’ diverse motives, we survey state-of-the-art truth by considering two classical roles, e.g., crowdsourcing and expert spammers. For each role, several topics related to spam attacks either with or without disguising and possible outliers are highlighted. Comparison analyses led to some interesting conclusions: 1) data on professional spammers are more challenging to collect and less reliable than data on crowdsourcing spammers; 2) most linguistic evidences are less reliable than behavioral footprints; 3) abnormal activities are as trustworthy as spamming objectives, while they hardly need any extra support, such as the user profile; and 4) the top reliable facts requiring acceptable effort aredeviation,burstiness,grouped spamming,deviation over the threshold,review distribution,opinion proportionandspam cost. Moreover, we introduce several promising directions for future research. In general, this survey may shed light on new angles that can be used to understand review spam and to improve the performance of any anti-spam platforms.