End-to-end Multimodel Deep Learning for Malware Classification

End-to-end Multimodel Deep Learning for Malware Classification
复制标题

DOI:
10.1109/ijcnn48605.2020.9207120
复制
发表时间:
2020-07
期刊:
2020 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
E. Snow;M. Alam;Alexander M. Glandon;K. Iftekharuddin
E. Snow;M. Alam;Alexander M. Glandon;K. Iftekharuddin
中科院分区:
其他
文献类型:
--
作者:
E. Snow;M. Alam;Alexander M. Glandon;K. Iftekharuddin

文献摘要

相似文献

恶意软件(Malware)旨在对计算机造成不必要的或破坏性的影响。由于现代社会依赖计算机运行,恶意软件有可能造成无法估量的损害。因此,开发有效打击恶意软件的技术至关重要。随着多态恶意软件的流行,传统的反恶意软件技术无法跟上新恶意软件的出现速度。这对开发高效和强大的恶意软件检测技术提出了重大挑战。克服这一挑战的一种方法是在已知恶意软件家族中对新恶意软件进行分类。已经提出了几种机器学习方法来解决恶意软件分类问题。然而,这些技术依赖于从恶意软件数据中提取的手工设计的特征,这对于分类新的恶意软件可能不是有效的。深度学习模型在解决图像和文本分类等各种分类任务方面取得了巨大成功。最近的深度学习技术能够直接从输入数据中提取特征。因此,本文提出了一个端到端的多模型深度学习框架(以下简称多模型学习),以解决具有挑战性的恶意软件分类问题。该模型利用三种不同的深度神经网络架构,从恶意软件数据的不同属性中联合学习有意义的特征。端到端学习可同时优化所有处理步骤,从而提高模型的准确性和可推广性。该模型的性能使用广泛使用和公开可用的Microsoft恶意软件挑战数据集进行测试,并与最先进的基于深度学习的恶意软件分类管道进行比较。我们的研究结果表明,所提出的模型实现了与最先进的方法相当的性能,同时使用端到端多模型学习提供更快的训练。
Malicious software (malware) is designed to cause unwanted or destructive effects on computers. Since modern society is dependent on computers to function, malware has the potential to do untold damage. Therefore, developing techniques to effectively combat malware is critical. With the rise in popularity of polymorphic malware, conventional anti-malware techniques fail to keep up with the rate of emergence of new malware. This poses a major challenge towards developing an efficient and robust malware detection technique. One approach to overcoming this challenge is to classify new malware among families of known malware. Several machine learning methods have been proposed for solving the malware classification problem. However, these techniques rely on hand-engineered features extracted from malware data which may not be effective for classifying new malware. Deep learning models have shown paramount success for solving various classification tasks such as image and text classification. Recent deep learning techniques are capable of extracting features directly from the input data. Consequently, this paper proposes an end-to-end deep learning framework for multimodels (henceforth, multimodel learning) to solve the challenging malware classification problem. The proposed model utilizes three different deep neural network architectures to jointly learn meaningful features from different attributes of the malware data. End-to-end learning optimizes all processing steps simultaneously, which improves model accuracy and generalizability. The performance of the model is tested with the widely used and publicly available Microsoft Malware Challenge Dataset and is compared with the state-of-the-art deep learning-based malware classification pipeline. Our results suggest that the proposed model achieves comparable performance to the state-of-the-art methods while offering faster training using end-to-end multimodel learning.