Semantic Feature Selection for Text with Application to Phishing Email Detection

Semantic Feature Selection for Text with Application to Phishing Email Detection
复制标题

DOI:
10.1007/978-3-319-12160-4_27
复制
发表时间:
2013-11
期刊:
--
影响因子:
--
通讯作者:
Rakesh M. Verma;Nabil Hossain
Rakesh M. Verma;Nabil Hossain
中科院分区:
其他
文献类型:
--
作者:
Rakesh M. Verma;Nabil Hossain

文献摘要

被引文献

相似文献

在网络钓鱼攻击中,毫无戒心的受害者通常通过电子邮件被引诱到旨在窃取敏感信息的网站,如银行/信用卡账号、帐户登录信息等。每年,互联网用户因此而损失数十亿美元。本文提出了一种基于统计t-检验和WordNet的文本语义特征选择方法,并通过设计语义和统计相结合的分类器对电子邮件中的文本进行分析,验证了该方法在钓鱼邮件检测中的有效性。我们的特征选择方法是通用的,也适用于其他涉及基于文本的分析的应用程序。我们的纯邮件正文分类器对钓鱼邮件的检测准确率达到95%以上,误检率为2.24%。由于它使用了语义,我们的特征选择方法对自适应攻击具有健壮性,并且避免了机器学习分类器需要频繁重新训练的问题。
In a phishing attack, an unsuspecting victim is lured, typically via an email, to a web site designed to steal sensitive information such as bank/credit card account numbers, login information for accounts, etc. Each year Internet users lose billions of dollars to this scourge. In this paper, we present a general semantic feature selection method for text problems based on the statistical t-test and WordNet, and we show its effectiveness on phishing email detection by designing classifiers that combine semantics and statistics in analyzing the text in the email. Our feature selection method is general and useful for other applications involving text-based analysis as well. Our emailbody-text-onlyclassifier achieves more than 95 % accuracy on detecting phishing emails with a false positive rate of 2.24 %. Due to its use of semantics, our feature selection method is robust against adaptive attacks and avoids the problem of frequent retraining needed by machine learning classifiers.