Adversarial Autoencoder Data Synthesis for Enhancing Machine Learning-Based Phishing Detection Algorithms

Adversarial Autoencoder Data Synthesis for Enhancing Machine Learning-Based Phishing Detection Algorithms
复制标题

DOI:
10.1109/tsc.2023.3234806
复制
发表时间:
2023-07
影响因子:
8.1
通讯作者:
H. Shirazi;Shashika Ranga Muramudalige;I. Ray;A. Jayasumana;Haonan Wang
H. Shirazi;Shashika Ranga Muramudalige;I. Ray;A. Jayasumana;Haonan Wang
中科院分区:
计算机科学2区
文献类型:
--
作者:
H. Shirazi;Shashika Ranga Muramudalige;I. Ray;A. Jayasumana;Haonan Wang

文献摘要

相似文献

有监督的机器学习经常被用来检测钓鱼网站。然而,用于训练目的的网络钓鱼数据的稀缺限制了分类器的性能。此外,机器学习算法容易受到敌意攻击:对攻击数据的微小扰动可能会绕过分类器。这些问题降低了机器学习对网络钓鱼检测的效率。我们提出了两种基于生成性对抗网络(GAN)的方法,它们综合钓鱼和合法样本来模拟真实世界的网站。关于真实世界数据集的信息是从十个公开可用的网络钓鱼数据集中获得的,AAE(对抗性自动编码器)和WGAN(Wasserstein Gan)使用这些数据来生成合成数据。使用真实数据和合成数据,我们演示了如何实现更高性能和更强抵抗对手攻击的分类器。我们提出了一组假设,并通过实验验证了它们:(I)合成样本和实际样本的不可区分性,(Ii)分类器对对抗性攻击的敏感性,(Iii)通过在包含正确标记的合成样本的较大数据集上进行训练来缓解对抗性攻击,以及(Iv)在大数据集上训练的分类器具有更好的性能。我们的AAE和WGAN已经在广泛的数据集上进行了培训,使我们对其广泛的适用性持乐观态度。
Supervised machine learning is often used to detect phishing websites. However, the scarcity of phishing data for training purposes limits the classifier's performance. Further, machine learning algorithms are prone to adversarial attacks: small perturbations on attack data can bypass the classifier. These problems make machine learning less effective for phishing detection. We propose two Generative Adversarial Network (GAN) based approaches that synthesize phishing and legitimate samples to mimic real-world websites. Information about real-world datasets is obtained from ten publicly available phishing datasets which are used by the AAE (Adversarial Autoencoder) and WGAN (Wasserstein GAN) for generating synthetic data. Using both real and synthesized data, we demonstrate how to implement classifiers with higher performance and more resistance to adversarial attacks. We propose a set of hypotheses and validate them through experiments to demonstrate: (i) indistinguishability of synthesized samples from actual ones, (ii) susceptibility of classifiers to adversarial attacks, (iii) mitigating adversarial attacks by training on larger datasets that include correctly labeled synthesized samples, and (iv) better performance of classifiers trained on large datasets. Our AAE and WGAN have been trained on a wide range of datasets, making us optimistic about its widespread applicability.