A Decade of Mal-Activity Reporting: A Retrospective Analysis of Internet Malicious Activity Blacklists

A Decade of Mal-Activity Reporting: A Retrospective Analysis of Internet Malicious Activity Blacklists
复制标题

恶意活动报告的十年:互联网恶意活动黑名单的回顾分析

DOI:
10.1145/3321705.3329834
复制
发表时间:
2019
期刊:
Proceedings of the 2019 ACM Asia Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Kanchana Thilakarathna
Kanchana Thilakarathna
中科院分区:
--
文献类型:
--
作者:
Benjamin Zi Hao Zhao;Muhammad Ikram;H. Asghar;M. Kâafar;Abdelberi Chaabane;Kanchana Thilakarathna

文献摘要

被引文献

相似文献

本文重点关注公共黑名单对互联网恶意活动(简称恶意活动)的报告,目的是对多年来报告的内容进行系统的描述,更重要的是,报告活动的演变。使用22个黑名单的初始种子,涵盖2007年1月至2017年6月期间,我们收集了超过5100万份涉及全球662K唯一IP地址的恶意活动报告。利用Wayback Machine、防病毒(AV)工具报告和几个额外的公共数据集(例如,BGP路由视图和互联网注册表),我们丰富了历史元信息,包括地理位置(国家),自治系统(AS)的数量和类型的不良活动的数据。此外,我们使用最初标记的约157万个不良活动数据集(从公共黑名单中获得)来训练机器学习分类器,以分类通过其他来源获得的剩余未标记的约4400万个不良活动数据集。我们将我们收集的独特数据集(和使用的脚本)公开以供进一步研究。该论文的主要贡献是一种新的报告收集方法,采用机器学习方法对报告的活动进行分类,对数据集进行表征,最重要的是对不良活动报告行为进行时间分析。受P2P行为建模的启发,我们的分析表明,一些类别的不良活动(例如,网络钓鱼)和少量恶意活动源是持久的,这表明基于黑名单的预防系统是无效的,或者具有不合理的长更新周期。我们的分析还表明,资源可以更好地利用,重点放在严重的不良活动的贡献者,这构成了大部分的不良活动。
This paper focuses on reporting of Internet malicious activity (or mal-activity in short) by public blacklists with the objective of providing a systematic characterization of what has been reported over the years, and more importantly, the evolution of reported activities. Using an initial seed of 22 blacklists, covering the period from January 2007 to June 2017, we collect more than 51 million mal-activity reports involving 662K unique IP addresses worldwide. Leveraging the Wayback Machine, antivirus (AV) tool reports and several additional public datasets (e.g., BGP Route Views and Internet registries) we enrich the data with historical meta-information including geo-locations (countries), autonomous system (AS) numbers and types of mal-activity. Furthermore, we use the initially labelled dataset of ~1.57 million mal-activities (obtained from public blacklists) to train a machine learning classifier to classify the remaining unlabeled dataset of ~44 million mal-activities obtained through additional sources. We make our unique collected dataset (and scripts used) publicly available for further research. The main contributions of the paper are a novel means of report collection, with a machine learning approach to classify reported activities, characterization of the dataset and, most importantly, temporal analysis of mal-activity reporting behavior. Inspired by P2P behavior modeling, our analysis shows that some classes of mal-activities (e.g., phishing) and a small number of mal-activity sources are persistent, suggesting that either blacklist-based prevention systems are ineffective or have unreasonably long update periods. Our analysis also indicates that resources can be better utilized by focusing on heavy mal-activity contributors, which constitute the bulk of mal-activities.