Evaluating a semisupervised approach to phishing url identification in a realistic scenario

Evaluating a semisupervised approach to phishing url identification in a realistic scenario
复制标题

在现实场景中评估网络钓鱼 URL 识别的半监督方法

DOI:
--
复制
发表时间:
2011
期刊:
International Conference on Email and Anti-Spam
影响因子:
--
通讯作者:
Gary Warner
Gary Warner
中科院分区:
--
文献类型:
--
作者:
Binod Gyawali;T. Solorio;M. Montes;Brad Wardman;Gary Warner

文献摘要

被引文献

相似文献

钓鱼网站已成为窃取敏感信息的常见方法,例如互联网用户的用户名,密码和信用卡详细信息。我们提出了一种半监督机器学习方法来检测一组网络钓鱼和垃圾邮件的网址钓鱼网址。垃圾邮件是这些URL的来源。实际上,通过这些垃圾邮件收到的钓鱼URL的数量比其他URL要少。我们的研究是针对一个现实的情况下,一个高度不平衡的数据集包含钓鱼和垃圾邮件的网址1:654的比例检测钓鱼网址。为了训练学习算法,需要标记URL,其中手动标记是一种常见的方法。鉴于手动标记来自大型数据集的所有URL是不可行的,我们建议通过手动标记仅10%的URL并使用半监督学习算法来减少手动干预。我们比较所提出的方法与监督学习方法。评估结果表明,我们的建议是有竞争力的,如果它与适当的特征选择和欠采样技术相结合。
Phishing sites have become a common approach to steal sensitive information, such as usernames, passwords and credit card details of the internet users. We propose a semisupervised machine learning approach to detect phishing URLs from a set of phishing and spam URLs. Spam emails are the source of these URLs. In reality, the number of phishing URLs received through these spam emails is fewer compared to other URLs. Our study is targeted to detect phishing URLs in a realistic scenario of a highly imbalanced data set containing phishing and spam URLs with 1:654 ratio. To train a learning algorithm labeled URLs are needed, where manual labeling is a common approach. Given that it is not feasible to manually label all the URLs from large data sets, we propose reducing manual intervention by labeling only 10% of the URLs manually and using a semisupervised learning algorithm. We compare the proposed approach with a supervised learning approach. Evaluation results show that our proposal is competitive if it is applied in combination with appropriate feature selection and undersampling techniques.