Semantic Feature Selection for Text with Application to Phishing Email Detection
Semantic Feature Selection for Text with Application to Phishing Email Detection
复制标题
DOI:
10.1007/978-3-319-12160-4_27
复制
发表时间:
2013-11
期刊:
影响因子:
--
通讯作者:
Rakesh M. Verma;Nabil Hossain
中科院分区:
文献类型:
--
作者:
Rakesh M. Verma;Nabil Hossain
In a phishing attack, an unsuspecting victim is lured, typically via an email, to a web site designed to steal sensitive information such as bank/credit card account numbers, login information for accounts, etc. Each year Internet users lose billions of dollars to this scourge. In this paper, we present a general semantic feature selection method for text problems based on the statistical t-test and WordNet, and we show its effectiveness on phishing email detection by designing classifiers that combine semantics and statistics in analyzing the text in the email. Our feature selection method is general and useful for other applications involving text-based analysis as well. Our emailbody-text-onlyclassifier achieves more than 95 % accuracy on detecting phishing emails with a false positive rate of 2.24 %. Due to its use of semantics, our feature selection method is robust against adaptive attacks and avoids the problem of frequent retraining needed by machine learning classifiers.