Diverse Datasets and a Customizable Benchmarking Framework for Phishing

Diverse Datasets and a Customizable Benchmarking Framework for Phishing
复制标题

DOI:
10.1145/3375708.3380313
复制
发表时间:
2020-03
期刊:
Proceedings of the Sixth International Workshop on Security and Privacy Analytics
影响因子:
--
通讯作者:
Victor Zeng;Shahryar Baki;Ayman El Aassal;Rakesh M. Verma;Luis F. T. Moraes;Avisha Das
Victor Zeng;Shahryar Baki;Ayman El Aassal;Rakesh M. Verma;Luis F. T. Moraes;Avisha Das
中科院分区:
其他
文献类型:
--
作者:
Victor Zeng;Shahryar Baki;Ayman El Aassal;Rakesh M. Verma;Luis F. T. Moraes;Avisha Das

文献摘要

相似文献

网络钓鱼是一个具有挑战性的问题,许多研究人员在几篇论文中使用了许多不同的数据集和技术来解决这个问题。当提出新的特征或方法时,研究人员通常用有限的指标、数据集和参数来测试他们提出的方法。因此,需要一个基准框架和数据集来尽可能全面地评估这些系统。在本文中,我们讨论:(i)我们在创建和传播网络钓鱼电子邮件、网站和URL检测的各种代表性数据集方面所做的努力,以及(ii) PhishBench,我们的网络钓鱼检测系统基准测试框架。PhishBench允许研究人员在提供的数据上轻松有效地评估和比较特征和分类方法。
Phishing is a challenging problem that has been addressed by many researchers in several papers using many different datatsets and techniques~\citedas2019sok. Researchers usually test their proposed methods with limited metrics, datasets, and parameters when presenting new features or approach(es). Hence, the need arises for a benchmarking framework and dataset to evaluate such systems as comprehensively as possible. In this paper, we discuss: (i) our efforts on the creation and dissemination of diverse and representative datasets for phishing email, website and URL detection, and (ii) PhishBench, our framework for benchmarking phishing detection systems. PhishBench allows researchers to evaluate and compare features and classification approaches easily and efficiently on the provided data.