Spam Filtering Using Statistical Data Compression Models

Spam Filtering Using Statistical Data Compression Models
复制标题

DOI:
--
复制
发表时间:
2006-12
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Andrej Bratko;G. Cormack;B. Filipič;T. Lynam;B. Zupan
Andrej Bratko;G. Cormack;B. Filipič;T. Lynam;B. Zupan
中科院分区:
其他
文献类型:
--
作者:
Andrej Bratko;G. Cormack;B. Filipič;T. Lynam;B. Zupan

文献摘要

被引文献

相似文献

垃圾邮件过滤是文本分类中的一个特殊问题,其特点是过滤器面临着一个主动的对手,这个对手不断地试图逃避过滤。由于垃圾邮件不断发展,大多数实际应用是基于在线用户反馈,该任务要求快速,增量和鲁棒的学习算法。本文提出了一种基于自适应统计数据压缩模型的垃圾邮件过滤方法。这些模型的性质允许它们被用作基于字符级或二进制序列的概率文本分类器。通过将消息建模为序列,标记化和其他容易出错的预处理步骤被完全省略,从而产生一种非常健壮的方法。这些模型也可以快速构建,并且可以增量更新。我们评估两种不同的压缩算法的过滤性能;动态马尔可夫压缩和预测部分匹配。我们的实证评估结果表明,压缩模型优于目前建立的垃圾邮件过滤器,以及在以前的研究中提出的一些方法。
Spam filtering poses a special problem in text categorization, of which the defining characteristic is that filters face an active adversary, which constantly attempts to evade filtering. Since spam evolves continuously and most practical applications are based on online user feedback, the task calls for fast, incremental and robust learning algorithms. In this paper, we investigate a novel approach to spam filtering based on adaptive statistical data compression models. The nature of these models allows them to be employed as probabilistic text classifiers based on character-level or binary sequences. By modeling messages as sequences, tokenization and other error-prone preprocessing steps are omitted altogether, resulting in a method that is very robust. The models are also fast to construct and incrementally updateable. We evaluate the filtering performance of two different compression algorithms; dynamic Markov compression and prediction by partial matching. The results of our empirical evaluation indicate that compression models outperform currently established spam filters, as well as a number of methods proposed in previous studies.