New filtering approaches for phishing email

New filtering approaches for phishing email
复制标题

DOI:
10.3233/jcs-2010-0371
复制
发表时间:
2010-01-01
影响因子:
1.2
通讯作者:
Strobel, Siehyun
Strobel, Siehyun
中科院分区:
其他
文献类型:
--
作者:
Bergholz, Andre;De Beer, Jan;Strobel, Siehyun

文献摘要

被引文献

相似文献

网络钓鱼电子邮件通常包含来自可靠来源的消息,要求用户单击指向网站的链接,并要求用户输入密码或其他机密信息。大多数网络钓鱼电子邮件的目的是从金融机构取款或获取私人信息。网络钓鱼在过去几年中急剧增加,对全球安全和经济构成严重威胁。针对网络钓鱼有许多可能的对策。这些范围从communicationoriented的方法,如认证协议的黑名单,以内容为基础的过滤approaches.We认为,前两种方法目前没有广泛实施或表现出赤字。因此,基于内容的网络钓鱼过滤器是必要的,并广泛使用,以提高通信安全。提取捕获电子邮件的内容和结构属性的许多特征。随后,统计分类器使用这些特征在标记为火腿(合法),垃圾邮件或网络钓鱼的电子邮件的训练集上进行训练。这个分类器,然后可以应用到电子邮件流估计类的新传入的电子邮件。在本文中,我们描述了一些新的功能,特别适合于识别网络钓鱼电子邮件。这些包括电子邮件主题的低维描述的统计模型,电子邮件文本和外部链接的顺序分析,嵌入式徽标的检测以及隐藏盐的指标。隐藏的盐是故意添加或扭曲的内容不被读者察觉。对于实证评估,我们已经获得了一个大型的现实语料库的电子邮件预标记为垃圾邮件,网络钓鱼,火腿(合法)。在实验中,我们的方法优于其他公布的方法分类钓鱼电子邮件。我们讨论了这些结果的实际应用,这种方法在工作流程中的电子邮件提供商的影响。最后,我们描述了一种策略,过滤器可以更新和适应新类型的网络钓鱼。
Phishing emails usually contain a message from a credible looking source requesting a user to click a link to a website where she/he is asked to enter a password or other confidential information. Most phishing emails aim at withdrawing money from financial institutions or getting access to private information. Phishing has increased enormously over the last years and is a serious threat to global security and economy. There are a number of possible countermeasures to phishing. These range from communicationoriented approaches like authentication protocols over blacklisting to content-based filtering approaches.We argue that the first two approaches are currently not broadly implemented or exhibit deficits. Therefore content-based phishing filters are necessary and widely used to increase communication security. A number of features are extracted capturing the content and structural properties of the email. Subsequently a statistical classifier is trained using these features on a training set of emails labeled as ham (legitimate), spam or phishing. This classifier may then be applied to an email stream to estimate the classes of new incoming emails.In this paper we describe a number of novel features that are particularly well-suited to identify phishing emails. These include statistical models for the low-dimensional descriptions of email topics, sequential analysis of email text and external links, the detection of embedded logos as well as indicators for hidden salting. Hidden salting is the intentional addition or distortion of content not perceivable by the reader. For empirical evaluation we have obtained a large realistic corpus of emails prelabeled as spam, phishing, and ham (legitimate). In experiments our methods outperform other published approaches for classifying phishing emails. We discuss the implications of these results for the practical application of this approach in the workflow of an email provider. Finally we describe a strategy how the filters may be updated and adapted to new types of phishing.