TREC 2007 Spam Track Overview

TREC 2007 Spam Track Overview
复制标题

DOI:
--
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
G. Cormack
G. Cormack
中科院分区:
其他
文献类型:
--
作者:
G. Cormack

文献摘要

被引文献

相似文献

TREC的垃圾邮件跟踪使用一个标准的测试框架,该框架提供了一组按时间顺序排列的电子邮件消息,并通过垃圾邮件过滤器进行分类。在过滤任务中,邮件一次一个地呈现给过滤器,过滤器产生一个二元判断(垃圾邮件或火腿[即非垃圾邮件]),并将其与人类判断的黄金标准进行比较。过滤器还产生垃圾邮件分数,旨在反映分类邮件是垃圾邮件的可能性,这是事后ROC(接收者操作特征)分析的主题。对四种不同形式的用户反馈进行建模:即时反馈,每条消息的黄金标准在分类后立即传达给过滤器;延迟反馈,黄金标准稍后传达给过滤器(或可能永远不会),以便对不时地阅读电子邮件并且可能不勤奋地报告过滤器的错误的用户进行建模;利用部分反馈,仅将用于电子邮件接收者的子集的黄金标准传输到过滤器,以便模拟一些用户从不报告过滤器错误的情况;利用主动在线学习(由D.来自塔夫茨大学的Sculley [5])过滤器允许对一定配额的消息请求即时反馈,该配额远远小于总数。两个测试语料库-电子邮件消息加上黄金标准的判断-被用来评估主题过滤器。一个公共语料库(trec 07 p)被分发给参与者,他们使用跟踪提供的工具包实现框架和四种反馈在语料库上运行过滤器。一个私有语料库(MrX 3)没有分发给参与者;相反,参与者提交了使用工具包在私有数据上运行的过滤器实现。12个小组参加了这一跟踪,每个小组在四种反馈模式(立即;延迟;部分;主动)中的每一种模式下提交最多四个过滤器进行评估。
TREC’s Spam Track uses a standard testing framework that presents a set of chronologically ordered email messages a spam filter for classification. In the filtering task, the messages are presented one at at time to the filter, which yields a binary judgement (spam or ham [i.e. non-spam]) which is compared to a humanadjudicated gold standard. The filter also yields a spamminess score, intended to reflect the likelihood that the classified message is spam, which is the subject of post-hoc ROC (Receiver Operating Characteristic) analysis. Four different forms of user feedback are modeled: with immediate feedback the gold standard for each message is communicated to the filter immediately following classification; with delayed feedback the gold standard is communicated to the filter sometime later (or potentially never), so as to model a user reading email from time to time and perhaps not diligently reporting the filter’s errors; with partial feedback the gold standard for only a subset of email recipients is transmitted to the filter, so as to model the case of some users never reporting filter errors; with active on-line learning (suggested by D. Sculley from Tufts University [5]) the filter is allowed to request immediate feedback for a certain quota of messages which is considerably smaller than the total number. Two test corpora – email messages plus gold standard judgements – were used to evaluate subject filters. One public corpus (trec07p) was distributed to participants, who ran their filters on the corpora using a track-supplied toolkit implementing the framework and the four kinds of feedback. One private corporus (MrX 3) was not distributed to participants; rather, participants submitted filter implementations that were run, using the toolkit, on the private data. Twelve groups participated in the track, each submitting up to four filters for evaluation in each of the four feedback modes (immediate; delayed; partial; active).