The Enron Corpus: A New Dataset for Email Classification Research

The Enron Corpus: A New Dataset for Email Classification Research
复制标题

DOI:
10.1007/978-3-540-30115-8_22
复制
发表时间:
2004-09
期刊:
--
影响因子:
--
通讯作者:
Bryan Klimt;Yiming Yang
Bryan Klimt;Yiming Yang
中科院分区:
其他
文献类型:
--
作者:
Bryan Klimt;Yiming Yang

文献摘要

被引文献

相似文献

将电子邮件信息自动分类到用户指定的文件夹中以及从按时间顺序排列的电子邮件流中提取信息已经成为文本学习研究中有趣的领域。然而,缺乏大型基准集合一直是研究问题和评估解决方案的障碍。在本文中,我们引入安然语料库作为一个新的测试平台。我们分析了它在电子邮件文件夹预测方面的适用性,并在各种条件下提供了最先进的分类器(支持向量机)的基线结果,包括单独使用单个部分(From, to, Subject和body)作为分类器输入的情况,以及将所有部分与回归权重结合使用。
Automated classification of email messages into user-specific folders and information extraction from chronologically ordered email streams have become interesting areas in text learning research. However, the lack of large benchmark collections has been an obstacle for studying the problems and evaluating the solutions. In this paper, we introduce the Enron corpus as a new test bed. We analyze its suitability with respect to email folder prediction, and provide the baseline results of a state-of-the-art classifier (Support Vector Machines) under various conditions, including the cases of using individual sections (From, To, Subject and body) alone as the input to the classifier, and using all the sections in combination with regression weights.