Detecting splogs using similarities of splog HTML structures

Detecting splogs using similarities of splog HTML structures
复制标题

DOI:
10.1145/2108616.2108661
复制
发表时间:
2010-01
期刊:
--
影响因子:
--
通讯作者:
Taichi Katayama;Takayuki Yoshinaka;T. Utsuro;Yasuhide Kawada;T. Fukuhara
Taichi Katayama;Takayuki Yoshinaka;T. Utsuro;Yasuhide Kawada;T. Fukuhara
中科院分区:
其他
文献类型:
--
作者:
Taichi Katayama;Takayuki Yoshinaka;T. Utsuro;Yasuhide Kawada;T. Fukuhara

文献摘要

相似文献

垃圾博客或splogs是托管垃圾帖子的博客,使用机器生成或劫持的内容创建,其唯一目的是托管广告或增加目标网站的inlinks数量。在这些splogs中,本文的重点是检测一组splogs,估计是由同一个垃圾邮件发送者创建的。我们特别表明,由同一个垃圾邮件发送者创建的splog之间的html结构的相似性有助于提高splog检测的性能。在度量html结构的相似性时,我们从html文件的DOM树中提取一个块列表(最小内容单元)。我们发现,估计由相同的垃圾邮件发送者创建的splog的html文件往往有类似的DOM树,这种趋势是非常有效的splog检测。
Spam blogs or splogs are blogs hosting spam posts, created using machine generated or hijacked content for the sole purpose of hosting advertisements or increasing the number of inlinks of target sites. Among those splogs, this paper focuses on detecting a group of splogs which are estimated to be created by an identical spammer. We especially show that similarities of html structures among those splogs created by an identical spammer contribute to improving the performance of splog detection. In measuring similarities of html structures, we extract a list of blocks (minimum unit of content) from the DOM tree of a html file. We show that the html files of splogs estimated to be created by an identical spammer tend to have similar DOM trees and this tendency is quite effective in splog detection.