Can We Leverage Predictive Uncertainty to Detect Dataset Shift and Adversarial Examples in Android Malware Detection?

Can We Leverage Predictive Uncertainty to Detect Dataset Shift and Adversarial Examples in Android Malware Detection?
复制标题

DOI:
10.1145/3485832.3485916
复制
发表时间:
2021-09
期刊:
Proceedings of the 37th Annual Computer Security Applications Conference
影响因子:
--
通讯作者:
Deqiang Li;Tian Qiu;Shuo Chen;Qianmu Li;Shouhuai Xu
Deqiang Li;Tian Qiu;Shuo Chen;Qianmu Li;Shouhuai Xu
中科院分区:
其他
文献类型:
--
作者:
Deqiang Li;Tian Qiu;Shuo Chen;Qianmu Li;Shouhuai Xu

文献摘要

被引文献

相似文献

深度学习检测恶意软件(Malware)的方法是有前景的,但尚未解决数据集迁移的问题,即与测试集关联的示例及其标签的联合分布不同于训练集的分布。这个问题导致深度学习模型在没有用户注意的情况下退化。为了缓解这一问题,一种方法是让分类器不仅预测给定示例上的标签,而且将其不确定性(或置信度)呈现在预测的标签上,从而防御者可以决定是否使用预测的标签。虽然这种方法很直观,也很重要,但它的功能和局限性还没有得到很好的理解。在本文中,我们进行了一项实证研究,以评估恶意软件检测器的预测不确定性质量。具体地说,我们重新设计和构建了24个Android恶意软件检测器(通过使用6种校准方法改造4个现成的检测器),并使用9个指标量化它们的不确定性,其中包括3个处理数据失衡的指标。我们的主要发现是:(I)预测不确定性确实有助于在数据集移动的情况下实现可靠的恶意软件检测,但无法应对敌意逃避攻击;(Ii)近似贝叶斯方法有望校准和推广恶意软件检测器以应对数据集迁移,但无法应对对抗性逃避攻击;(Iii)对抗性逃避攻击可以使校准方法失效,并且量化与对抗性实例预测标签相关的不确定性是一个悬而未决的问题(即,使用预测不确定性来检测对抗性实例是无效的)。
The deep learning approach to detecting malicious software (malware) is promising but has yet to tackle the problem of dataset shift, namely that the joint distribution of examples and their labels associated with the test set is different from that of the training set. This problem causes the degradation of deep learning models without users’ notice. In order to alleviate the problem, one approach is to let a classifier not only predict the label on a given example but also present its uncertainty (or confidence) on the predicted label, whereby a defender can decide whether to use the predicted label or not. While intuitive and clearly important, the capabilities and limitations of this approach have not been well understood. In this paper, we conduct an empirical study to evaluate the quality of predictive uncertainties of malware detectors. Specifically, we re-design and build 24 Android malware detectors (by transforming four off-the-shelf detectors with six calibration methods) and quantify their uncertainties with nine metrics, including three metrics dealing with data imbalance. Our main findings are: (i) predictive uncertainty indeed helps achieve reliable malware detection in the presence of dataset shift, but cannot cope with adversarial evasion attacks; (ii) approximate Bayesian methods are promising to calibrate and generalize malware detectors to deal with dataset shift, but cannot cope with adversarial evasion attacks; (iii) adversarial evasion attacks can render calibration methods useless, and it is an open problem to quantify the uncertainty associated with the predicted labels of adversarial examples (i.e., it is not effective to use predictive uncertainty to detect adversarial examples).