Web Spam Detection with Anti-Trust Rank

Web Spam Detection with Anti-Trust Rank
复制标题

DOI:
--
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
Vijay Krishnan;R. Raj
Vijay Krishnan;R. Raj
中科院分区:
其他
文献类型:
--
作者:
Vijay Krishnan;R. Raj

文献摘要

被引文献

相似文献

网络上的垃圾页面使用各种技术来人为地在搜索引擎结果中获得高排名。人类专家可以很好地识别垃圾页面和信息质量可疑的页面,但是对于大量页面,使用人类努力实际上是不可行的。类似于信任排名算法[1],我们提出了一种选择由人类评估的页面种子集的方法。然后,我们使用的链接结构的网页和人工标记的种子集,以检测其他垃圾网页。我们在WebGraph数据集上的实验[3]表明,我们的方法在从小种子集中检测垃圾页面方面非常有效,并且除了检测具有较高pagerank的页面外,平均而言,比Trust Rank算法实现了更高的垃圾页面检测精度[10,11]。
Spam pages on the web use various techniques to artificially achieve high rankings in search engine results. Human experts can do a good job of identifying spam pages and pages whose information is of dubious quality, but it is practically infeasible to use human effort for a large number of pages. Similar to the Trust Rank algorithm [1], we propose a method of selecting a seed set of pages to be evaluated by a human. We then use the link structure of the web and the manually labeled seed set, to detect other spam pages. Our experiments on the WebGraph dataset [3] show that our approach is very effective at detecting spam pages from a small seed set and achieves higher precision of spam page detection than the Trust Rank algorithm, apart from detecting pages with higher pageranks [10, 11], on an average.