Analyzing and Explaining Black-Box Models for Online Malware Detection

Analyzing and Explaining Black-Box Models for Online Malware Detection
复制标题

DOI:
10.1109/access.2023.3255176
复制
发表时间:
2023
期刊:
影响因子:
3.9
通讯作者:
Harikha Manthena;Jeffrey Kimmell;Mahmoud Abdelsalam;Maanak Gupta
Harikha Manthena;Jeffrey Kimmell;Mahmoud Abdelsalam;Maanak Gupta
中科院分区:
计算机科学3区
文献类型:
--
作者:
Harikha Manthena;Jeffrey Kimmell;Mahmoud Abdelsalam;Maanak Gupta

文献摘要

相似文献

近年来,大量的研究集中在分析机器学习(ML)模型用于恶意软件检测的有效性上。这些方法包括从决策树和聚类等方法到更复杂的方法,如支持向量机(SVM)和深度神经网络。特别是,神经网络已被证明在检测复杂和高级恶意软件方面非常有效。然而,这有一个警告。神经网络是出了名的复杂。因此,他们所做的决定通常只是被接受,而没有质疑为什么模型会做出那个特定的决定。神经网络的黑箱特性对研究人员提出了挑战,需要探索如何解释支持向量机和神经网络等黑箱模型及其决策过程。透明度和可解释性为专家和恶意软件分析师提供了ML模型决策的保证和可信度。此外,它还有助于生成可用于加强网络威胁情报共享的综合报告。因此,这种急需的分析促使我们在本文中探索ML模型在在线恶意软件检测领域的可解释性和可解释性。在本文中,我们使用Shapley加性解释(SHAP)可解释性技术来实现解释不同ML模型的结果的高效性能,例如SVM线性、SVM- rbf(径向基函数)、随机森林(RF)、前馈神经网络(FFNN)和卷积神经网络(CNN)模型在在线恶意软件数据集上训练。为了解释这些模型的输出,将KernalSHAP、TreeSHAP和DeepSHAP等可解释性技术应用于获得的结果。
In recent years, a significant amount of research has focused on analyzing the effectiveness of machine learning (ML) models for malware detection. These approaches have ranged from methods such as decision trees and clustering to more complex approaches like support vector machine (SVM) and deep neural networks. In particular, neural networks have proven to be very effective in detecting complex and advanced malware. This, however, comes with a caveat. Neural networks are notoriously complex. Therefore, the decisions that they make are often just accepted without questioning why the model made that specific decision. The black box characteristic of neural networks has challenged researchers to explore methods to explain black-box models such as SVM and neural networks and their decision-making process. Transparency and explainability give the experts and malware analysts assurance and trustworthiness about the ML models’ decisions. In addition, it helps in generating comprehensive reports that can be used to enhance cyber threat intelligence sharing. As such, this much-needed analysis drives our work in this paper to explore the explainability and interpretability of ML models in the field of online malware detection. In this paper, we used the Shapley Additive exPlanations (SHAP) explainability technique to achieve efficient performance in interpreting the outcome of different ML models such as SVM Linear, SVM-RBF (Radial Basis Function), Random Forest (RF), Feed-Forward Neural Net (FFNN), and Convolutional Neural Network (CNN) models trained on an online malware dataset. To explain the output of these models, explainability techniques such as KernalSHAP, TreeSHAP, and DeepSHAP are applied to the obtained results.