Everything Is in the Name - A URL Based Approach for Phishing Detection

Everything Is in the Name - A URL Based Approach for Phishing Detection
复制标题

一切皆在名称中 - 基于 URL 的网络钓鱼检测方法

DOI:
--
复制
发表时间:
2019
期刊:
International Conference on Cyber Security Cryptography and Machine Learning
影响因子:
--
通讯作者:
S. Lodha
S. Lodha
中科院分区:
--
文献类型:
--
作者:
Harshal Tupsamudre;A. Singh;S. Lodha

文献摘要

被引文献

相似文献

网络钓鱼攻击是对网络安全最常见的威胁之一,在这种攻击中,用户被欺骗,在一个被欺骗的网站上泄露敏感信息。大多数现代web浏览器使用已确认的网络钓鱼url黑名单来对抗网络钓鱼攻击。然而,黑名单方法的一个主要缺点是它对新生成的网络钓鱼无效。基于机器学习的技术依赖于从URL(例如,URL长度和词袋)或网页(例如,TF-IDF和表单字段)中提取的特征,被认为在识别新的网络钓鱼攻击方面更有效。与基于页面的功能相比,使用基于URL的功能的主要好处是,机器学习模型甚至可以在网页浏览器加载页面之前对新的URL进行动态分类,从而避免其他潜在的危险,例如飞车下载攻击和加密劫持攻击。
Phishing attack, in which a user is tricked into revealing sensitive information on a spoofed website, is one of the most common threat to cybersecurity. Most modern web browsers counter phishing attacks using a blacklist of confirmed phishing URLs. However, one major disadvantage of the blacklist method is that it is ineffective against newly generated phishes. Machine learning based techniques that rely on features extracted from URL (e.g., URL length and bag-of-words) or web page (e.g., TF-IDF and form fields) are considered to be more effective in identifying new phishing attacks. The main benefit of using URL based features over page based features is that the machine learning model can classify new URLs on-the-fly even before the page is loaded by the web browser, thus avoiding other potential dangers such as drive-by download attacks and cryptojacking attacks.