Investigating Labelless Drift Adaptation for Malware Detection

Investigating Labelless Drift Adaptation for Malware Detection
复制标题

DOI:
10.1145/3474369.3486873
复制
发表时间:
2021-11
期刊:
Proceedings of the 14th ACM Workshop on Artificial Intelligence and Security
影响因子:
--
通讯作者:
Zeliang Kan;Feargus Pendlebury;Fabio Pierazzi;L. Cavallaro
Zeliang Kan;Feargus Pendlebury;Fabio Pierazzi;L. Cavallaro
中科院分区:
其他
文献类型:
--
作者:
Zeliang Kan;Feargus Pendlebury;Fabio Pierazzi;L. Cavallaro

文献摘要

被引文献

相似文献

恶意软件的演变长期以来一直困扰着基于机器学习的检测系统,因为恶意软件作者开发了创新的策略来逃避检测和追逐利润。这会导致概念漂移,因为测试分布偏离训练,导致性能下降,需要不断监测和适应。在这项工作中,我们分析了DroidEvolver使用的自适应策略,DroidEvolver是一种最先进的学习系统,它使用伪标签进行自我更新,以避免与获得新的地面实况相关的高开销。在消除了原始评估中存在的实验偏差来源后,我们发现了这些伪标签的生成和整合中的一些缺陷,导致模型本身中毒时性能迅速下降。我们提出了DroidEvolver++,一个更强大的变种DroidEvolver,以解决这些问题,并强调伪标签在解决概念漂移的作用。我们测试了适应策略对不同程度的伪标签噪声的容忍度,并提出了采用方法来确保仅使用高质量的伪标签进行更新。最后,我们得出结论,使用伪标签仍然是一个很有前途的解决方案,标签容量的限制,但在设计更新机制,以避免负反馈循环和自我中毒的性能具有灾难性的影响时,必须非常小心。
The evolution of malware has long plagued machine learning-based detection systems, as malware authors develop innovative strategies to evade detection and chase profits. This induces concept drift as the test distribution diverges from the training, causing performance decay that requires constant monitoring and adaptation. In this work, we analyze the adaptation strategy used by DroidEvolver, a state-of-the-art learning system that self-updates using pseudo-labels to avoid the high overhead associated with obtaining a new ground truth. After removing sources of experimental bias present in the original evaluation, we identify a number of flaws in the generation and integration of these pseudo-labels, leading to a rapid onset of performance degradation as the model poisons itself. We propose DroidEvolver++, a more robust variant of DroidEvolver, to address these issues and highlight the role of pseudo-labels in addressing concept drift. We test the tolerance of the adaptation strategy versus different degrees of pseudo-label noise and propose the adoption of methods to ensure only high-quality pseudo-labels are used for updates. Ultimately, we conclude that the use of pseudo-labeling remains a promising solution to limitations on labeling capacity, but great care must be taken when designing update mechanisms to avoid negative feedback loops and self-poisoning which have catastrophic effects on performance.