DeepRan: Attention-based BiLSTM and CRF for Ransomware Early Detection and Classification

DeepRan: Attention-based BiLSTM and CRF for Ransomware Early Detection and Classification
复制标题

DOI:
10.1007/s10796-020-10017-4
复制
发表时间:
2020-06
影响因子:
5.9
通讯作者:
K. Roy;Qian Chen
K. Roy;Qian Chen
中科院分区:
计算机科学3区
文献类型:
--
作者:
K. Roy;Qian Chen

文献摘要

被引文献

相似文献

勒索软件是一种自我传播的恶意软件,对受感染计算机的文件系统进行加密,以勒索受害者获得经济利益。数以百计的学校、医院和地方政府受到勒索软件的影响,平均造成12.1天的系统停机时间(Siegel 2019)。该研究旨在开发一种基于深度学习的检测器DeepRanfor勒索软件早期检测和分类,以防止网络范围内的数据加密。DeepRan应用基于注意力的双向长短期记忆(BiLSTM)和完全连接(FC)层来模拟运营企业系统中主机的正常状态,并从从裸机服务器收集的大量环境主机日志数据中检测异常活动。DeepRan还通过使用条件随机场(CRF)模型扩展基于注意力的BiLSTM,将异常活动分类为候选勒索软件攻击之一。采用词频-逆文档频率(TF-IDF)方法从高维主机测井数据中提取语义信息。增量学习技术用于扩展模型的现有知识,以防止DeepRan质量随时间推移而下降。我们开发了一个裸金属服务器的测试床,并收集了两个用户63天的正常主机日志(IRB批准)。在受害主机上执行17次勒索软件攻击,受感染主机的日志数据用于验证DeepRan。实验结果表明,DeepRan对于勒索软件早期检测的检测准确率为99.87%(F1得分为99.02%)。该检测器还实现了96.5%的准确率,将异常事件分类为17个候选勒索软件家族之一。增量学习的应用被验证为随着时间的推移提高模型质量的有效技术。
Ransomware is a self-propagating malware encrypting file systems of the compromised computers to extort victims for financial gains. Hundreds of schools, hospitals, and local government municipalities have been disrupted by ransomware that already caused 12.1 days of system downtime on average (Siegel 2019). This study aims at developing a deep learning-based detectorDeepRanfor ransomware early detection and classification to prevent network-wide data encryption. DeepRan applies an attention-based bi-directional Long Short Term Memory (BiLSTM) with a fully connected (FC) layer to model normalcy of hosts in an operational enterprise system and detects abnormal activity from a large volume of ambient host logging data collected from bare metal servers. DeepRan also classifies abnormal activity as one of the candidate ransomware attacks by extending attention-based BiLSTM with a Conditional Random Fields (CRF) model. The Term Frequency-Inverse Document Frequency (TF-IDF) method is applied to extract semantic information from high dimensional host logging data. An incremental learning technique is used to extend the model’s existing knowledge to prevent DeepRan quality degradation over time. We develop a testbed of bare metal servers and collect normal host logs of two users for 63 days (IRB-approved). 17 ransomware attacks are executed on the victim hosts, and the infected host logging data is used for validating DeepRan. Experimental results present that DeepRan produces 99.87% detection accuracy (F1-score of 99.02%) for ransomware early detection. The detector also achieves 96.5%accuracy to classify abnormal events as one of 17 candidate ransomware families. The application of incremental learning is validated as an efficient technique to enhance model quality over time.