Online Discriminative Spam Filter Training

Online Discriminative Spam Filter Training
复制标题

DOI:
--
复制
发表时间:
2006-07
期刊:
--
影响因子:
--
通讯作者:
Joshua Goodman;Wen-tau Yih
Joshua Goodman;Wen-tau Yih
中科院分区:
其他
文献类型:
--
作者:
Joshua Goodman;Wen-tau Yih

文献摘要

被引文献

相似文献

我们描述了一种非常简单的判别训练垃圾邮件过滤器的技术。我们在TREC安然垃圾邮件语料库上的结果对火腿来说是最好的。1%衡量标准,其次是1-ROCA衡量标准。对于Mr. X语料库,我们的1-ROCA测量是第二好,第三好。1%的措施。我们使用一个非常简单的特征提取器(主题和标题中的所有单词)。我们的学习算法也很简单:一个逻辑回归模型的梯度下降。
We describe a very simple technique for discriminatively training a spam filter. Our results on the TREC Enron spam corpus would have been the best for the Ham at .1% measure, and second best by the 1-ROCA measure. For the Mr. X corpus, our 1-ROCA measure was a close second best, and third best by the Ham at .1% measure. We use a very simple feature extractor (all words in the subject and headers). Our learning algorithm is also very simple: gradient descent of a logistic regression model.