Adversarial Perturbations Against Deep Neural Networks for Malware Classification

Adversarial Perturbations Against Deep Neural Networks for Malware Classification
复制标题

DOI:
--
复制
发表时间:
2016-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Kathrin Grosse;Nicolas Papernot;Praveen Manoharan;M. Backes;P. Mcdaniel
Kathrin Grosse;Nicolas Papernot;Praveen Manoharan;M. Backes;P. Mcdaniel
中科院分区:
其他
文献类型:
--
作者:
Kathrin Grosse;Nicolas Papernot;Praveen Manoharan;M. Backes;P. Mcdaniel

文献摘要

被引文献

相似文献

深度神经网络,像许多其他机器学习模型一样,最近被证明对恶意输入缺乏鲁棒性。这些输入是通过微小但仔细选择的扰动从常规输入中导出的,这些扰动将机器学习模型欺骗到所需的错误分类中。在这个新兴领域的现有工作主要是特定的图像分类领域,因为高熵的图像可以方便地操纵,而不改变图像的整体视觉外观。然而,目前尚不清楚此类攻击如何转化为更安全敏感的应用程序,如恶意软件检测-这可能会对样本生成带来重大挑战,并可能导致严重的失败后果。在本文中,我们展示了如何为用作恶意软件分类器的神经网络构建高效的对抗性样本制作攻击。与计算机视觉领域相比,恶意软件分类的应用领域在对抗性样本制作问题中引入了额外的约束:(i)连续的、可微分的输入域被离散的、通常是二进制的输入所取代;以及(ii)保持视觉外观不变的宽松条件被要求等效的功能行为所取代。我们使用DREBIN Android恶意软件数据集训练了许多不同的恶意软件分类器实例,证明了这些攻击的可行性。我们还评估了在何种程度上可以利用对抗性手工制作的潜在防御机制来设置恶意软件分类。虽然特征约简并没有被证明具有积极的影响,但对逆向制作的样本进行蒸馏和重新训练显示出有希望的结果。
Deep neural networks, like many other machine learning models, have recently been shown to lack robustness against adversarially crafted inputs. These inputs are derived from regular inputs by minor yet carefully selected perturbations that deceive machine learning models into desired misclassifications. Existing work in this emerging field was largely specific to the domain of image classification, since the high-entropy of images can be conveniently manipulated without changing the images' overall visual appearance. Yet, it remains unclear how such attacks translate to more security-sensitive applications such as malware detection - which may pose significant challenges in sample generation and arguably grave consequences for failure. In this paper, we show how to construct highly-effective adversarial sample crafting attacks for neural networks used as malware classifiers. The application domain of malware classification introduces additional constraints in the adversarial sample crafting problem when compared to the computer vision domain: (i) continuous, differentiable input domains are replaced by discrete, often binary inputs; and (ii) the loose condition of leaving visual appearance unchanged is replaced by requiring equivalent functional behavior. We demonstrate the feasibility of these attacks on many different instances of malware classifiers that we trained using the DREBIN Android malware data set. We furthermore evaluate to which extent potential defensive mechanisms against adversarial crafting can be leveraged to the setting of malware classification. While feature reduction did not prove to have a positive impact, distillation and re-training on adversarially crafted samples show promising results.