N-gram MalGAN: Evading machine learning detection via feature n-gram

N-gram MalGAN: Evading machine learning detection via feature n-gram
复制标题

DOI:
10.1016/j.dcan.2021.11.007
复制
发表时间:
2021-11
期刊:
Digit. Commun. Networks
影响因子:
--
通讯作者:
Enmin Zhu;Jianjie Zhang;Jijie Yan;Kongyang Chen;Chongzhi Gao
Enmin Zhu;Jianjie Zhang;Jijie Yan;Kongyang Chen;Chongzhi Gao
中科院分区:
其他
文献类型:
--
作者:
Enmin Zhu;Jianjie Zhang;Jijie Yan;Kongyang Chen;Chongzhi Gao

文献摘要

被引文献

相似文献

近年来,已经引入了许多具有不同特征策略的对抗性恶意软件示例,特别是GAN及其变体,以处理安全威胁,例如,逃避机器学习检测器的检测。然而,这些解决方案仍然存在部署复杂或运行时间长的问题。在本文中,我们提出了一个n-gram MalGAN方法来解决这些问题。我们从自然语言处理(NLP)领域借用了n-gram的概念,以扩展MalGAN中对抗性恶意软件示例的特征源。通常,n-gram MalGAN直接从可执行文件的十六进制字节码中获得特征向量。它可以用简单的程序语言(例如,C++),不需要任何可执行文件的先验知识或任何专业的特征提取工具。这些功能在功能上是独立的,因此可以添加到恶意程序的非功能区域,以保持其原始的可执行性。通过这种方式,n-gram可以使对抗性攻击更容易和更方便。实验结果表明,在适当的分组率下,MalGAN对不同的机器学习算法的规避率至少为88.58%,对随机森林算法的规避率甚至达到100%.
In recent years, many adversarial malware examples with different feature strategies, especially GAN and its variants, have been introduced to handle the security threats, e.g., evading the detection of machine learning detectors. However, these solutions still suffer from problems of complicated deployment or long running time. In this paper, we propose an n-gram MalGAN method to solve these problems. We borrow the idea of n-gram from the Natural Language Processing (NLP) area to expand feature sources for adversarial malware examples in MalGAN. Generally, the n-gram MalGAN obtains the feature vector directly from the hexadecimal bytecodes of the executable file. It can be implemented easily and conveniently with a simple program language (e.g., C++), with no need for any prior knowledge of the executable file or any professional feature extraction tools. These features are functionally independent and thus can be added to the non-functional area of the malicious program to maintain its original executability. In this way, the n-gram could make the adversarial attack easier and more convenient. Experimental results show that the evasion rate of the n-gram MalGAN is at least 88.58% to attack different machine learning algorithms under an appropriate group rate, growing to even 100% for the Random Forest algorithm.