On-line spam filter fusion

On-line spam filter fusion
复制标题

DOI:
10.1145/1148170.1148195
复制
发表时间:
2006-08
期刊:
Proceedings of the 29th annual international ACM SIGIR conference on Research and development in information retrieval
影响因子:
--
通讯作者:
T. Lynam;G. Cormack;D. Cheriton
T. Lynam;G. Cormack;D. Cheriton
中科院分区:
其他
文献类型:
--
作者:
T. Lynam;G. Cormack;D. Cheriton

文献摘要

被引文献

相似文献

我们表明,一组独立开发的垃圾邮件过滤器可以以简单的方式组合起来,以提供比任何单个过滤器更好的过滤。 TREC 2005 垃圾邮件赛道上评估的 53 个垃圾邮件过滤器的结果经过事后组合,以模拟过滤器的并行在线操作。使用 TREC 方法对综合结果进行评估,与最佳过滤器相比,结果提高了两倍以上。最简单的方法——对各个过滤器返回的二元分类进行平均——会产生非常好的结果。一种新方法——根据各个过滤器返回的分数来平均对数赔率估计——产生了更好的结果,并为基于 SVM 和逻辑回归的堆叠方法提供了输入。堆叠方法似乎提供了进一步的改进,但仅限于非常大的语料库。在堆叠方法中,逻辑回归产生更好的结果。最后,我们证明可以选择滤波器的先验小子集,这些子集在组合时仍然远远优于最佳的单个滤波器。
We show that a set of independently developed spam filters may be combined in simple ways to provide substantially better filtering than any of the individual filters. The results of fifty-three spam filters evaluated at the TREC 2005 Spam Track were combined post-hoc so as to simulate the parallel on-line operation of the filters. The combined results were evaluated using the TREC methodology, yielding more than a factor of two improvement over the best filter. The simplest method -- averaging the binary classifications returned by the individual filters -- yields a remarkably good result. A new method -- averaging log-odds estimates based on the scores returned by the individual filters -- yields a somewhat better result, and provides input to SVM- and logistic-regression-based stacking methods. The stacking methods appear to provide further improvement, but only for very large corpora. Of the stacking methods, logistic regression yields the better result. Finally, we show that it is possible to select a priori small subsets of the filters that, when combined, still outperform the best individual filter by a substantial margin.