Discover millions of fake followers in Weibo

Discover millions of fake followers in Weibo
复制标题

DOI:
10.1007/s13278-016-0324-2
复制
发表时间:
2016-12-01
影响因子:
2.8
通讯作者:
Lu, Jianguo
Lu, Jianguo
中科院分区:
其他
文献类型:
--
作者:
Zhang, Yi;Lu, Jianguo

文献摘要

被引文献

相似文献

微博是Twitter的中国版,Twitter吸引了数亿用户。就像其他在线社交网络(以下简称OSN)一样,微博也有大量的虚假账户。他们的创建是为了向客户出售他们的以下链接,他们希望增加他们的追随者数量。这些虚假账户很难单独识别,特别是当它们是由复杂的程序创建或由人类直接控制时。本文提出了一种新的虚假帐户检测方法,该方法基于这些帐户存在的目的:它们被创建为跟踪其目标,从而导致其客户的追随者列表之间的高度重叠。本文调查了关注者列表重复或几乎重复的顶级微博账户(以下称为近似重复)。发现近似重复是一项具有挑战性的任务。网络很大;数据整体不可用;成对比较非常昂贵。我们开发了一种基于抽样的方法来发现所有接近重复的顶级账户,这些账户至少有50,000名粉丝。在实验中,我们发现了395个近似重复,这导致我们有1190万个虚假账户(占总用户的4.56%),他们发送了7.411亿个链接(占整个边缘的9.50%)。此外,我们描述了四个典型的垃圾邮件发送者的结构,这些垃圾邮件发送者分为34组,并分析了每个组的属性。
Weibo is the Chinese counterpart of Twitter, which has attracted hundreds of millions of users. Just like other Online Social Networks (hereafter OSNs), Weibo has a large number of fake accounts. They are created to sell their following links to customers, who want to boost their follower counts. These bogus accounts are difficult to identify individually, especially when they are created by sophisticated programs or controlled by human beings directly. This paper proposes a novel fake account detection method that is based on the very purpose of the existence of these accounts: they are created to follow their targets en masse, resulting in high-overlapping between the follower lists of their customers. This paper investigates the top Weibo accounts whose follower lists duplicate or nearly duplicate each other (hereafter called near-duplicates). Discovering near-duplicates is a challenging task. The network is large; the data in its entirety are not available; the pair-wise comparison is very expensive. We developed a sampling-based approach to discover all the near-duplicates of the top accounts, who have at least 50,000 followers. In the experiment, we found 395 near-duplicates, which leads us to 11.90 million fake accounts (4.56 % of total users) who send 741.10 million links (9.50 % of the entire edges). Furthermore, we characterize four typical structures of the spammers, cluster these spammers into 34 groups, and analyze the properties of each group.