On Network-level Clusters for Spam Detection

On Network-level Clusters for Spam Detection
复制标题

DOI:
--
复制
发表时间:
2010-02
期刊:
--
影响因子:
--
通讯作者:
Zhiyun Qian;Z. Morley Mao;Yinglian Xie;Fang Yu
Zhiyun Qian;Z. Morley Mao;Yinglian Xie;Fang Yu
中科院分区:
其他
文献类型:
--
作者:
Zhiyun Qian;Z. Morley Mao;Yinglian Xie;Fang Yu

文献摘要

被引文献

相似文献

基于IP的黑名单是过滤垃圾邮件的有效方法。然而,在黑名单中建立和维护单个IP地址是困难的,因为新的恶意主机不断出现,并且它们的IP地址也可能随着时间的推移而改变。为了缓解这个问题,研究人员提出用IP集群替换黑名单中的单个IP地址,例如,BGP集群。在本文中,我们仔细研究了基于IP集群的方法的准确性,以了解其有效性和基本局限性。基于这样的理解,我们提出并实现了一种新的聚类方法,同时考虑网络起源和DNS信息,并将其与SpamAssassin,一个流行的垃圾邮件过滤系统广泛使用的今天。将我们的方法应用于在大型大学部门收集的7个月的电子邮件跟踪,与直接应用各种基于公共IP的黑名单相比,我们可以将误报率降低50%,而不会增加误报率。此外,使用蜜罐电子邮件帐户和真实的用户帐户,我们表明,我们的方法可以捕获30%-50%的垃圾邮件,今天通过SpamAssassin滑。
IP-based blacklist is an effective way to filter spam emails. However, building and maintaining individual IP addresses in the blacklist is difficult, as new malicious hosts continuously appear and their IP addresses may also change over time. To mitigate this problem, researchers have proposed to replace individual IP addresses in the blacklist with IP clusters, e.g., BGP clusters. In this paper, we closely examine the accuracy of IP-cluster-based approaches to understand their effectiveness and fundamental limitations. Based on such understanding, we propose and implement a new clustering approach that considers both network origin and DNS information, and incorporate it with SpamAssassin, a popular spam filtering system widely used today. Applying our approach to a 7-month email trace collected at a large university department, we can reduce the false negative rate by 50% compared with directly applying various public IP-based blacklists without increasing the false positive rate. Furthermore, using honeypot email accounts and real user accounts, we show that our approach can capture 30% 50% of the spam emails that slip through SpamAssassin today.