Deceptive review detection using labeled and unlabeled data

Deceptive review detection using labeled and unlabeled data
复制标题

DOI:
10.1007/s11042-016-3819-y
复制
发表时间:
2016-08
影响因子:
3.6
通讯作者:
Jitendra Kumar Rout;Smriti Singh;S. K. Jena;Sambit Bakshi
Jitendra Kumar Rout;Smriti Singh;S. K. Jena;Sambit Bakshi
中科院分区:
计算机科学4区
文献类型:
--
作者:
Jitendra Kumar Rout;Smriti Singh;S. K. Jena;Sambit Bakshi

文献摘要

被引文献

相似文献

电子商务网站上有数以百万计的产品和服务,由于存在许多替代品,因此很难根据需求搜索最合适的产品。为了摆脱这种情况,最流行和最有用的方法是关注固执己见的社交媒体上其他人的评论,他们已经尝试过这些方法。几乎所有电子商务网站都为用户提供了对他们所体验的产品和服务的看法和体验的设施。个人、制造商和零售商越来越多地使用客户评论来做出购买和业务决策。由于没有对收到的评论进行审查,任何人都可以一致撰写任何最终导致评论垃圾邮件的内容。此外,在利润和/或宣传的欲望的驱动下,垃圾邮件发送者会生成综合评论来宣传某些产品/品牌并降低竞争对手的产品/品牌。随着时间的推移,欺骗性评论垃圾邮件大幅增长。在这项工作中,我们应用了监督和非监督技术来识别垃圾评论。最有效的特征集已被组装起来用于模型构建。情感分析也被纳入检测过程中。为了获得最佳性能,一些著名的分类器被应用于标记数据集。此外,对于未标记的数据,在计算出垃圾邮件检测所需的属性后使用聚类。此外,垃圾邮件评论者也很有可能对多媒体社交网络中的内容污染负责,因为现在许多用户都使用其社交网络登录来进行评论。最后,这项工作可以扩展到查找负责将虚假多媒体内容发布到各自社交网络中的可疑帐户。
Availability of millions of products and services on e-commerce sites makes it difficult to search the best suitable product according to the requirements because of existence of many alternatives. To get rid of this the most popular and useful approach is to follow reviews of others in opinionated social medias, who have already tried them. Almost all e-commerce sites provide facility to the users for giving views and experience of the product and services they experienced. The customers reviews are increasingly used by individuals, manufacturers and retailers for purchase and business decisions. As there is no scrutiny over the reviews received, anybody can write anything unanimously which conclusively leads to review spam. Moreover, driven by the desire of profit and/or publicity, spammers produce synthesized reviews to promote some products/brand and demote competitors products/brand. Deceptive review spam has seen a considerable growth overtime. In this work, we have applied supervised as well as unsupervised techniques to identify review spam. Most effective feature sets have been assembled for model building. Sentiment analysis has also been incorporated in the detection process. In order to get best performance some well-known classifiers were applied on labeled dataset. Further, for the unlabeled data, clustering is used after desired attributes were computed for spam detection. Additionally, there is a high chance that spam reviewers may also be held responsible for content pollution in multimedia social networks, because nowadays many users are giving the reviews using their social network logins. Finally, the work can be extended to find suspicious accounts responsible for posting fake multimedia contents into respective social networks.