Sorting Ransomware from Malware Utilizing Machine Learning Methods with Dynamic Analysis

Sorting Ransomware from Malware Utilizing Machine Learning Methods with Dynamic Analysis
复制标题

利用机器学习方法和动态分析对勒索软件和恶意软件进行分类

DOI:
10.1145/3565287.3617632
复制
发表时间:
2023
期刊:
ACM
影响因子:
--
通讯作者:
Li, Feng
Li, Feng
中科院分区:
--
文献类型:
--
作者:
Schoenbachler, Joshua;Krishnan, Vinay;Agarwal, Garvit;Li, Feng

文献摘要

相似文献

勒索软件攻击在过去十几年里大幅增长,扰乱了涉及个人数据的企业。在本文中,我们讨论了使用机器学习技术来识别勒索软件、恶意软件和良性软件。我们从互联网上的储存库收集了数据样本,并参考了先前研究中的数据集,这为我们的方法提供了基础。我们使用杜鹃沙盒™手动收集勒索软件、恶意软件和良性软件样本。我们筛选了某些功能组,以测试和确定感染过程中的某些活动/进程是否可以用于正确区分勒索软件与恶意软件和良性软件。这些功能组代表正在运行的应用程序中的相关进程:网络活动、注册表/事件进程和文件交互。使用几种机器学习(ML)模型对数据集进行分析,这些模型包括随机森林、支持向量机(SVM)、梯度提升和使用二进制分类的决策树。识别勒索软件和良性软件的最佳分类器是随机森林和支持向量机,F1得分为86%,F1得分为82%,随机森林的总体准确率为85%。除了勒索软件与良性软件,我们还比较了恶意软件和勒索软件数据。梯度增强分类器和决策树在区分勒索软件和恶意软件方面具有100%的准确率。这一高结果可能部分是由较小的恶意软件和勒索软件数据集造成的。总体而言,我们能够成功区分勒索软件与恶意软件和良性软件。
Ransomware attacks have grown significantly in the past dozen years and have disrupted businesses that engage with personal data. In this paper, we discuss the identification of ransomware, malware, and benign software from one another using machine learning techniques. We collected data samples from repositories on the internet as well as referencing a dataset from a previous study that provided a basis for our approach. We collected ransomware, malware, and benign software samples manually using Cuckoo Sandbox™. We filtered on certain feature groups to test and determine if certain activity/processes in the infection process could be used to correctly distinguish ransomware from malware and benign software. These feature groups represent correlated processes within a running application: network activity, registry/events processes, and file interactions. The datasets were analyzed using several machine learning (ML) models which included Random Forest, Support Vector Machines (SVM), Gradient Boosting, and Decision Trees using binary classification. The best classifiers for distinctly identifying ransomware from benign software were Random Forest and SVM with an f1- score of 86% and an f1-score of 82% as well as an 85% in overall accuracy for Random Forest. In addition to ransomware versus benign software, we also compared malware software to ransomware data. Yielding a 100% accuracy in performance, Gradient Boosting Classifier and Decision Trees were the best at distinguishing ransomware from malware software. This high result may partially be caused by a smaller malware and ransomware dataset. Overall, we were able to successfully distinguish ransomware from malware and benign software.