Support vector machines for spam categorization

Support vector machines for spam categorization
复制标题

DOI:
10.1109/72.788645
复制
发表时间:
1999-09-01
影响因子:
--
通讯作者:
Vapnik, VN
Vapnik, VN
中科院分区:
其他
文献类型:
--
作者:
Drucker, H;Wu, DH;Vapnik, VN

文献摘要

被引文献

相似文献

我们研究了支持向量机(SVM)的使用,通过将其与其他三种分类算法(Ripper, Rocchio和boosting决策树)进行比较,将电子邮件分类为垃圾邮件或非垃圾邮件。这四种算法在两个不同的数据集上进行了测试:一个数据集的特征数量被限制为1000个最佳特征,另一个数据集的维数超过7000,SVM在使用二进制特征时表现最佳。对于这两个数据集,增强树和支持向量机在准确性和速度方面都有可接受的测试性能。然而,SVM的训练时间明显更短。
We study the use of support vector machines (SVM's) In classifying e-mail as spam or nonspam by comparing it to three other classification algorithms: Ripper, Rocchio, and boosting decision trees, These four algorithms were tested on two different data sets: one data set where the number of features were constrained to the 1000 best features and another data set where the dimensionality was over 7000, SVM's performed best when using binary features. For both data sets, boosting trees and SVM's had acceptable test performance in terms of accuracy and speed. However, SVM's had significantly less training time.