URL-based Phishing Detection using the Entropy of Non-Alphanumeric Characters

URL-based Phishing Detection using the Entropy of Non-Alphanumeric Characters
复制标题

DOI:
10.1145/3366030.3366064
复制
发表时间:
2019-12
期刊:
Proceedings of the 21st International Conference on Information Integration and Web-based Applications & Services
影响因子:
--
通讯作者:
Eint Sandi Aung;H. Yamana
Eint Sandi Aung;H. Yamana
中科院分区:
其他
文献类型:
--
作者:
Eint Sandi Aung;H. Yamana

文献摘要

相似文献

网络钓鱼是一种个人信息盗窃行为,网络钓鱼者引诱用户窃取敏感信息。已经开发了使用各种技术的网络钓鱼检测机制。我们的假设是,网络钓鱼者创建虚假网站时,网页中的信息尽可能少,这使得通过分析网页内容来进行基于内容和视觉相似性的检测变得困难。为了克服这个问题,我们重点关注使用统一资源定位器 (URL) 来检测网络钓鱼。由于之前的工作提取了特定的特殊字符特征,因此我们假设非字母数字 (NAN) 字符分布会极大地影响基于 URL 的检测的性能。因此,我们提出了一种称为 NAN 字符熵的新功能,用于基于 URL 的网络钓鱼检测。平衡和不平衡数据集的实验评估显示,平衡数据集上的 ROC AUC 为 96%,不平衡数据集上的 ROC AUC 为 89%,这使得 ROC AUC 比不采用我们提出的特征增加了 5% 到 6%。
Phishing is a type of personal information theft in which phishers lure users to steal sensitive information. Phishing detection mechanisms using various techniques have been developed. Our hypothesis is that phishers create fake websites with as little information as possible in a webpage, which makes it difficult for content- and visual similarity-based detections by analyzing the webpage content. To overcome this, we focus on the use of Uniform Resource Locators (URLs) to detect phishing. Since previous work extracts specific special-character features, we assume that non-alphanumeric (NAN) character distributions highly impact the performance of URL-based detection. We hence propose a new feature called the entropy of NAN characters for URL-based phishing detection. Experimental evaluation with balanced and imbalanced datasets shows 96% ROC AUC on the balanced dataset and 89% ROC AUC on the imbalanced dataset, which increases the ROC AUC as 5 to 6% from without adopting our proposed feature.