Binary Black-Box Attacks Against Static Malware Detectors with Reinforcement Learning in Discrete Action Spaces

Binary Black-Box Attacks Against Static Malware Detectors with Reinforcement Learning in Discrete Action Spaces
复制标题

DOI:
10.1109/spw53761.2021.00021
复制
发表时间:
2021-05
期刊:
2021 IEEE Security and Privacy Workshops (SPW)
影响因子:
--
通讯作者:
Mohammadreza Ebrahimi;Jason L. Pacheco;Weifeng Li;J. Hu;Hsinchun Chen
Mohammadreza Ebrahimi;Jason L. Pacheco;Weifeng Li;J. Hu;Hsinchun Chen
中科院分区:
其他
文献类型:
--
作者:
Mohammadreza Ebrahimi;Jason L. Pacheco;Weifeng Li;J. Hu;Hsinchun Chen

文献摘要

相似文献

最近基于机器学习和深度学习的静态恶意软件检测器在识别看不见的恶意软件变体方面显示出突破性的性能。因此,它们越来越多地被用来降低动态恶意软件分析和手动签名识别的成本。尽管他们取得了成功,但研究表明,他们很容易受到敌意恶意软件的攻击,在恶意软件攻击中,敌手巧妙地修改已知的可执行恶意软件,以欺骗恶意软件检测器将其识别为良性文件。最近的研究表明,大规模自动创建这些敌意恶意软件变体有助于提高恶意软件检测器的健壮性。为简洁起见,我们将此过程称为对抗性恶意软件示例生成(AMG)。大多数AMG方法依赖于关于探测器的结构或参数的先验知识,这在实践中通常是不存在的。此外,这些方法中的大多数仅限于附加修改,即在不修改其原始内容的情况下将内容附加到恶意软件可执行文件。在这项研究中,我们提出了一种新的强化学习方法AMG-VAC,它将变分行动者-批评者(VAC)扩展到修改本质上是离散的非连续动作空间。我们在两个基于机器学习的恶意软件检测器上对所提出的AMG-VAC的规避性能进行了评估。虽然该方法在统计上显著优于现有的非RL和基于RL的AMG方法,但我们表明,所获得的规避动作序列对于揭示恶意软件检测器的漏洞是有用的。
Recent machine learning- and deep learning-based static malware detectors have shown breakthrough performance in identifying unseen malware variants. As a result, they are increasingly being adopted to lower the cost of dynamic malware analysis and manual signature identification. Despite their success, studies have shown that they can be vulnerable to adversarial malware attacks, in which an adversary modifies a known malware executable subtly to fool the malware detector into recognizing it as a benign file. Recent studies have shown that automatically crafting these adversarial malware variants at scale is beneficial to improve the robustness of malware detectors. For conciseness, we refer to this process as Adversarial Malware example Generation (AMG). Most AMG methods rely on prior knowledge about the architecture or parameters of the detector, which is not often available in practice. Moreover, the majority of these methods are restricted to additive modifications that append contents to the malware executable without modifying its original content. In this study, we offer a novel Reinforcement Learning (RL) method, AMG-VAC, which extends Variational Actor-Critic (VAC) to non-continuous action spaces where modifications are inherently discrete. We evaluate the evasion performance of the proposed AMG-VAC on two reputable machine learning-based malware detectors. While the proposed method outperforms extant non-RL and RL-based AMG methods by statistically significant margins, we show that the obtained evasive action sequences are useful in shedding light on malware detectors’ vulnerabilities.