Robust Spammer Detection in Microblogs: Leveraging User Carefulness

Robust Spammer Detection in Microblogs: Leveraging User Carefulness
复制标题

DOI:
10.1145/3086637
复制
发表时间:
2017-09-01
影响因子:
5
通讯作者:
Chen, Enhong
Chen, Enhong
中科院分区:
计算机科学3区
文献类型:
--
作者:
Fu, Hao;Xie, Xing;Chen, Enhong

文献摘要

被引文献

相似文献

微博网站,如Twitter和新浪微博,近年来已成为社交和分享信息的热门平台。垃圾邮件发送者也发现了这个新的机会,可以不公平地用未经请求的内容(即社交垃圾邮件)压倒正常用户。虽然每个人都很容易跟随合法用户,但最近的研究表明,合法用户和垃圾邮件发送者都因为不同的原因而跟随垃圾邮件发送者。也观察到用户有意寻找垃圾邮件发送者的证据。我们认为这种行为是检测垃圾邮件发送者的有用信息。在本文中,我们通过利用用户的“谨慎”来解决垃圾邮件发送者检测问题,这表明用户在跟踪潜在的垃圾邮件发送者时有多谨慎。我们提出了一个框架来衡量谨慎,并开发了一个监督学习算法来估计它的基础上已知的垃圾邮件发送者和合法用户。我们说明了如何可以提高检测算法的鲁棒性与援助所提出的措施。在新浪微博和Twitter两个拥有数百万用户的真实的数据集上进行了评测,并在新浪微博上进行了在线测试。结果表明,我们的方法确实抓住了谨慎,它是有效的检测垃圾邮件发送者。此外,我们发现我们的措施也是有益的其他应用,如链接预测。
Microblogging Web sites, such as Twitter and Sina Weibo, have become popular platforms for socializing and sharing information in recent years. Spammers have also discovered this new opportunity to unfairly overpower normal users with unsolicited content, namely social spams. Although it is intuitive for everyone to follow legitimate users, recent studies show that both legitimate users and spammers follow spammers for different reasons. Evidence of users seeking spammers on purpose is also observed. We regard this behavior as useful information for spammer detection. In this article, we approach the problem of spammer detection by leveraging the "carefulness" of users, which indicates how careful a user is when she is about to follow a potential spammer. We propose a framework to measure the carefulness and develop a supervised learning algorithm to estimate it based on known spammers and legitimate users. We illustrate how the robustness of the detection algorithms can be improved with aid of the proposed measure. Evaluation on two real datasets from SinaWeibo and Twitter withmillions of users are performed, as well as an online test on SinaWeibo. The results show that our approach indeed captures the carefulness, and it is effective for detecting spammers. In addition, we find that our measure is also beneficial for other applications, such as link prediction.