Visualizing and Interpreting RNN Models in URL-based Phishing Detection

Visualizing and Interpreting RNN Models in URL-based Phishing Detection
复制标题

基于 URL 的网络钓鱼检测中 RNN 模型的可视化和解释

DOI:
10.1145/3381991.3395602
复制
发表时间:
2020
期刊:
ACM Symposium on Access Control Models and Technologies
影响因子:
--
通讯作者:
Yue, Chuan
Yue, Chuan
中科院分区:
--
文献类型:
--
作者:
Feng, Tao;Yue, Chuan

文献摘要

参考文献

被引文献

相似文献

现有研究表明,使用传统的机器学习技术,简单地基于URL的特征进行网络钓鱼检测可以非常有效。在本文中,我们探索了深度学习的方法,并建立了四个仅利用URL的词汇特征来检测钓鱼攻击的递归神经网络模型。我们收集了150万个URL作为数据集,结果表明,我们的RNN模型可以达到99%以上的检测准确率,而不需要任何专业知识来手动识别特征。然而,众所周知,RNN和其他深度学习技术在很大程度上仍处于黑箱中。了解深度学习模型的内部结构对于改进和正确应用模型是非常重要和非常必要的。因此,在这项工作中,我们进一步开发了几种独特的可视化技术来深入解释RNN模型如何在内部实现出色的网络钓鱼检测性能。特别是,我们识别并回答了六个重要的研究问题,表明我们的四个RNN模型(1)是相辅相成的,可以组合成一个精度更高的集成模型,(2)可以很好地捕获传统机器学习方法中人工提取并用于网络钓鱼检测的相关特征,(3)可以帮助识别有用的新特征,以提高传统机器学习方法的准确性。我们在这项工作中的技术和经验可以帮助研究人员有效地将深度学习技术应用于解决其他现实世界的安全或隐私问题。
Existing studies have demonstrated that using traditional machine learning techniques, phishing detection simply based on the features of URLs can be very effective. In this paper, we explore the deep learning approach and build four RNN (Recurrent Neural Network) models that only use lexical features of URLs for detecting phishing attacks. We collect 1.5 million URLs as the dataset and show that our RNN models can achieve a higher than 99% detection accuracy without the need of any expert knowledge to manually identify the features. However, it is well known that RNNs and other deep learning techniques are still largely in black boxes. Understanding the internals of deep learning models is important and highly desirable to the improvement and proper application of the models. Therefore, in this work, we further develop several unique visualization techniques to intensively interpret how RNN models work internally in achieving the outstanding phishing detection performance. Especially, we identify and answer six important research questions, showing that our four RNN models (1) are complementary to each other and can be combined into an ensemble model with even better accuracy, (2) can well capture the relevant features that were manually extracted and used in the traditional machine learning approach for phishing detection, and (3) can help identify useful new features to enhance the accuracy of the traditional machine learning approach. Our techniques and experience in this work could be helpful for researchers to effectively apply deep learning techniques in addressing other real-world security or privacy problems.
DOI: 10.1145/1242572.1242659
发表时间: 2007-05
期刊: --
影响因子: --
作者:
Yue Zhang;Jason I. Hong;L. Cranor
通讯作者: Yue Zhang;Jason I. Hong;L. Cranor
网络钓鱼是魔鬼:重新思考 Web 单点登录系统安全
DOI: --
发表时间: 2013
期刊: USENIX Workshop on Large-Scale Exploits and Emergent Threats
影响因子: --
作者:
Chuan Yue
通讯作者: Chuan Yue
DOI: --
发表时间: 2016-06
期刊: ArXiv
影响因子: --
作者:
Hendrik Strobelt;Sebastian Gehrmann;Bernd Huber;H. Pfister;Alexander M. Rush
通讯作者: Hendrik Strobelt;Sebastian Gehrmann;Bernd Huber;H. Pfister;Alexander M. Rush
在现实场景中评估网络钓鱼 URL 识别的半监督方法
DOI: --
发表时间: 2011
期刊: International Conference on Email and Anti-Spam
影响因子: --
作者:
Binod Gyawali;T. Solorio;M. Montes;Brad Wardman;Gary Warner
通讯作者: Gary Warner