Robust Neural Malware Detection Models for Emulation Sequence Learning

Robust Neural Malware Detection Models for Emulation Sequence Learning
复制标题

用于仿真序列学习的鲁棒神经恶意软件检测模型

DOI:
--
复制
发表时间:
2018
期刊:
IEEE Military Communications Conference
影响因子:
--
通讯作者:
K. Selvaraj
K. Selvaraj
中科院分区:
--
文献类型:
--
作者:
Rakshit Agrawal;J. W. Stokes;M. Marinescu;K. Selvaraj

文献摘要

被引文献

相似文献

恶意软件或恶意软件在计算机安全方面提出了不断发展的挑战。这些以恶意文件形式或隐藏在合法文件中的嵌入式代码片段会对能够运行恶意命令序列的系统造成重大风险。恶意软件作者甚至使用多态性来重新排序这些命令,并创建几个恶意变体。然而,如果在安全环境中执行,则可以对仿真命令序列执行早期恶意软件检测。本文中提出的模型利用通过仿真获得的顺序数据来执行神经恶意软件检测。这些模型通过从这些序列中学习恶意事件操作的存在和共同出现的模式来瞄准恶意操作的核心。我们的模型可以捕获整个事件序列,并直接使用已知的目标标签进行训练。这些端到端学习模型由两种常用的结构提供支持-长短期记忆(LSTM)网络和卷积神经网络(CNN),以前提出的顺序恶意软件分类模型处理不超过200个事件。攻击者可以通过将任何恶意活动延迟到文件开头之后来逃避检测。我们提出了专门的模型,可以处理非常长的序列,同时成功地执行恶意软件检测,以有效的方式。我们提出了一个实现的卷积分割的长序列的方法,以解决这个漏洞,并操作长序列。我们在一个由634,249个文件序列组成的大型数据集上展示了我们的结果,这些文件序列非常长。
Malicious software, or malware, presents a continuously evolving challenge in computer security. These embedded snippets of code in the form of malicious files or hidden within legitimate files cause a major risk to systems with their ability to run malicious command sequences. Malware authors even use polymorphism to reorder these commands and create several malicious variations. However, if executed in a secure environment, one can perform early malware detection on emulated command sequences. The models presented in this paper leverage this sequential data derived via emulation in order to perform Neural Malware Detection. These models target the core of the malicious operation by learning the presence and pattern of co-occurrence of malicious event actions from within these sequences. Our models can capture entire event sequences and be trained directly using the known target labels. These end-to-end learning models are powered by two commonly used structures - Long Short-Term Memory (LSTM) Networks and Convolutional Neural Networks (CNNs), Previously proposed sequential malware classification models process no more than 200 events. Attackers can evade detection by delaying any malicious activity beyond the beginning of the file. We present specialized models that can handle extremely long sequences while successfully performing malware detection in an efficient way. We present an implementation of the Convoluted Partitioning of Long Sequences approach in order to tackle this vulnerability and operate on long sequences. We present our results on a large dataset consisting of 634,249 file sequences, with extremely long file sequences.