Improving Auto-Detection of Phishing Websites using Fresh-Phish Framework

Improving Auto-Detection of Phishing Websites using Fresh-Phish Framework
复制标题

DOI:
10.4018/ijmdem.2018010104
复制
发表时间:
2018
期刊:
Int. J. Multim. Data Eng. Manag.
影响因子:
--
通讯作者:
H. Shirazi;Kyle Haefner;I. Ray
H. Shirazi;Kyle Haefner;I. Ray
中科院分区:
其他
文献类型:
--
作者:
H. Shirazi;Kyle Haefner;I. Ray

文献摘要

相似文献

互联网居民正遭受着越来越频繁和复杂的网络钓鱼攻击。伴随着看起来可信的网站的电子邮件正在诱使用户,他们在不知不觉中交出了自己的凭据,损害了他们的隐私和安全。将这些钓鱼网站列入黑名单等方法变得站不住脚,无法跟上虚假网站爆炸的步伐。对恶意网站的检测必须自动化,并能够适应这种不断演变的社会工程形式。有一个改进的框架,以前被实现称为“Fresh-Phish”,用于为钓鱼网站创建当前的机器学习数据。改进的框架使用了总共28个不同的网站功能,这些功能使用Python进行查询,然后构建一个大型的标记数据集,并通过几个机器学习分类器对该数据集进行分析,以确定哪个是最准确的。这个修改后的框架通过在可能的情况下使用整数值而不是二进制值来提高对这些特征建模的准确性。本文不仅分析了该技术的准确性,还分析了训练模型所需的时间。
Denizens of the Internet are under a barrage of phishing attacks of increasing frequency and sophistication. Emails accompanied by authentic looking websites are ensnaring users who, unwittingly, hand over their credentials compromising both their privacy and security. Methods such as the blacklisting of these phishing websites become untenable and cannot keep pace with the explosion of fake sites. Detection of nefarious websites must become automated and be able to adapt to this ever-evolving form of social engineering. There is an improved framework that was previously implemented called “Fresh-Phish”, for creating current machine-learning data for phishing websites. The improved framework uses a total of 28 different website features that query using python, then a large labeled dataset is built and analyze over several machine learning classifiers against this dataset to determine which is the most accurate. This modified framework improves the accuracy of modeling those features by using integer rather than binary values where possible. This article analyzes not just the accuracy of the technique, but also how long it takes to train the model.