Detecting Blog Spams using the Vocabulary Size of All Substrings in Their Copies

Detecting Blog Spams using the Vocabulary Size of All Substrings in Their Copies
复制标题

DOI:
--
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
K. Narisawa;Yasuhiro Yamada;Daisuke Ikeda;Masayuki Takeda
K. Narisawa;Yasuhiro Yamada;Daisuke Ikeda;Masayuki Takeda
中科院分区:
其他
文献类型:
--
作者:
K. Narisawa;Yasuhiro Yamada;Daisuke Ikeda;Masayuki Takeda

文献摘要

被引文献

相似文献

本文解决了在博客条目中检测博客垃圾邮件的问题,这些垃圾邮件是博客站点上未经请求的消息。与垃圾邮件不同,典型的博客垃圾邮件是为了提高垃圾邮件发送者网站的PageRank而产生的,因此需要大量的博客垃圾邮件副本,并且所有副本都包含站点的url。因此,副本的数量,我们称之为频率,似乎是发现这类博客垃圾邮件的一个很好的关键。然而,如果频率大于某个阈值,则频率不足以检测算法将条目检测为博客垃圾邮件,原因如下:很难收集包含博客条目所有副本的网页;因此,输入数据仅包含条目的几个副本,其数量可能小于预定义的阈值;因此,基于频率的垃圾邮件检测算法无法检测到。我们提出了一种基于词汇表大小的垃圾邮件检测方法,而不是基于频率的方法,词汇表大小是频率相同的子字符串的数量。所提出的方法利用了正常博客条目中子字符串的词汇表大小遵循Zipf分布而博客垃圾邮件中的词汇表大小不遵循Zipf分布的事实。我们使用人工数据和从实际博客条目中收集的Web数据,通过实验证明了它的有效性。使用Web数据的实验表明,该方法可以检测到博客垃圾邮件,即使它的频率不是很大,并且该方法可以在给定的博客条目中同时发现所有带有一些副本的博客垃圾邮件。在一个英文博客网站上发现了一个用中文写的似乎是中国电影广告的博客垃圾邮件。结果表明,该方法与语言无关。我们也显示了可扩展性*。目前的隶属关系是用户科学研究所,九州大学,Hakozaki 6-10-1。所提出的方法在输入大小方面使用了大量的文本数据。
This paper addresses the problem of detecting blog spams , which are unsolicited messages on blog sites, among blog entries. Unlike a spam mail, a typical blog spam is produced to increase the PageRank for the spammer’s Web sites, and so many copies of the blog spam are necessary and all of them contain URLs of the sites. Therefore the number of the copies, we call it the frequency , seems to be a good key to find this type of blog spams. The frequency is not, however, sufficient for detection algorithms which detect an entry as a blog spam if the frequency is greater than some threshold value, because of the following reasons: it is very difficult to collect Web pages including all copies of a blog entry; therefore an input data contains only a few copies of the entry whose number may be smaller than the predefined threshold; and thus a frequency based spam detection algorithm fails to detect. Instead of frequency based approaches, we propose a spam detection method based on the vocabulary size , which is the number of substrings whose frequencies are the same. The proposed method utilizes the fact that the vocabulary size of substrings in normal blog entries follows the Zipf’s distribution but the vocabulary size in blog spams does not. We show its effectiveness by experiments, using both artificial data and Web data collected from actual blog entries. Experiments using Web data show that the proposed method can detect a blog spam even if the frequency of it is not so large, and that the method finds all blog spams with some copies simultaneously in given blog entries. A blog spam written in Chinese, which seems to be advertisements for Chinese movies, is found from an English blog site. This result shows that the proposed method is independent from the language. We also show the scalability ∗ The current affiliation is the User Science Institute, Kyushu University, Hakozaki 6-10-1 . of the proposed method with respect to input size using a huge size of text data.