Know your neighbors: web spam detection using the web topology

Know your neighbors: web spam detection using the web topology
复制标题

DOI:
10.1145/1277741.1277814
复制
发表时间:
2007-07
期刊:
--
影响因子:
--
通讯作者:
C. Castillo;D. Donato;A. Gionis;Vanessa Murdock;F. Silvestri
C. Castillo;D. Donato;A. Gionis;Vanessa Murdock;F. Silvestri
中科院分区:
其他
文献类型:
--
作者:
C. Castillo;D. Donato;A. Gionis;Vanessa Murdock;F. Silvestri

文献摘要

被引文献

相似文献

Web垃圾邮件会显著降低搜索引擎结果的质量。因此,商业搜索引擎有很大的动机来有效和准确地检测垃圾页面。在本文中,我们提出了一个垃圾邮件检测系统,结合了基于链接和基于内容的功能,并利用Web图的拓扑结构,利用网页之间的链接依赖关系。我们发现,链接的主机往往属于同一类:要么都是垃圾邮件或都是非垃圾邮件。我们展示了三种方法,将Web图拓扑结构到我们的基础分类器获得的预测:(i)聚类的主机图,并通过多数表决分配集群中所有主机的标签,(ii)传播预测的标签到相邻的主机,以及(iii)使用相邻主机的预测标签作为新的功能和重新训练分类器。其结果是一个准确的系统,用于检测Web垃圾邮件,测试了一个大型的公共数据集,使用的算法,可以在实践中应用到大规模的Web数据。
Web spam can significantly deteriorate the quality of search engine results. Thus there is a large incentive for commercial search engines to detect spam pages efficiently and accurately. In this paper we present a spam detection system that combines link-based and content-based features, and uses the topology of the Web graph by exploiting the link dependencies among the Web pages. We find that linked hosts tend to belong to the same class: either both are spam or both are non-spam. We demonstrate three methods of incorporating the Web graph topology into the predictions obtained by our base classifier: (i) clustering the host graph, and assigning the label of all hosts in the cluster by majority vote, (ii) propagating the predicted labels to neighboring hosts, and (iii) using the predicted labels of neighboring hosts as new features and retraining the classifier. The result is an accurate system for detecting Web spam, tested on a large and public dataset, using algorithms that can be applied in practice to large-scale Web data.