A Semi-Supervised Framework for Social Spammer Detection

A Semi-Supervised Framework for Social Spammer Detection
复制标题

DOI:
10.1007/978-3-319-18032-8_14
复制
发表时间:
2015-05
期刊:
--
影响因子:
--
通讯作者:
Zhaoxing Li;Xianchao Zhang;Hua Shen;Wenxin Liang;Zengyou He
Zhaoxing Li;Xianchao Zhang;Hua Shen;Wenxin Liang;Zengyou He
中科院分区:
其他
文献类型:
--
作者:
Zhaoxing Li;Xianchao Zhang;Hua Shen;Wenxin Liang;Zengyou He

文献摘要

被引文献

相似文献

垃圾邮件发送者创建了大量的妥协或虚假帐户,以传播有害信息的社交网络,如推特。识别社交垃圾邮件发送者已成为一个具有挑战性的问题。现有的社交垃圾邮件检测算法大多是基于监督学习的,需要大量的标记数据进行训练。然而,标记足够的训练集花费太多的资源,这使得监督学习不切实际的社会垃圾邮件检测。在本文中,我们提出了一个半监督的社会垃圾邮件检测框架(SSSD),它结合了监督分类模型与社会图上的排名计划。首先,我们用少量的标记数据训练原始分类器。其次,我们提出了一个排名模型来传播社会图上的信任和不信任。第三,我们选择的信任用户的分类器和排名评分判断为新的训练数据,并重新训练分类器。我们重复上面的所有步骤,直到分类器不能再细化。实验结果表明,在缺乏足够标记数据的情况下,该框架能够有效地检测出社交垃圾邮件发送者。
Spammers create large number of compromised or fake accounts to disseminate harmful information in social networks like Twitter. Identifying social spammers has become a challenging problem. Most of existing algorithms for social spammer detection are based on supervised learning, which needs a large amount of labeled data for training. However, labeling sufficient training set costs too much resources, which makes supervised learning impractical for social spammer detection. In this paper, we propose a semi-supervised framework for social spammer detection(SSSD), which combines the supervised classification model with a ranking scheme on the social graph. First, we train an original classifier with a small number of labeled data. Second, we propose a ranking model to propagate trust and distrust on the social graph. Third, we select confident users that are judged by the classifier and ranking scores as new training data and retrain the classifier. We repeat the all steps above until the classifier cannot be refined any more. Experimental results show that our framework can effectively detect social spammers in the condition of lacking sufficient labeled data.