Continuous Learning for Android Malware Detection

Continuous Learning for Android Malware Detection
复制标题

DOI:
10.48550/arxiv.2302.04332
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Yizheng Chen;Zhoujie Ding;David A. Wagner
Yizheng Chen;Zhoujie Ding;David A. Wagner
中科院分区:
其他
文献类型:
--
作者:
Yizheng Chen;Zhoujie Ding;David A. Wagner

文献摘要

相似文献

机器学习方法可以非常准确地检测 Android 恶意软件。然而,这些分类器有一个致命弱点,即概念漂移:由于恶意软件应用程序和良性应用程序的发展,它们很快就会过时且无效。我们的研究发现,在使用一年的数据训练 Android 恶意软件分类器后,在新测试样本上部署 6 个月后,F1 分数迅速从 0.99 下降到 0.76。在本文中,我们提出了解决 Android 恶意软件分类器概念漂移问题的新方法。由于机器学习技术需要不断部署,因此我们使用主动学习:选择新的样本供分析师进行标记,然后将标记的样本添加到训练集中以重新训练分类器。我们的关键思想是,基于相似性的不确定性对于概念漂移更加稳健。因此,我们将对比学习与主动学习结合起来。我们提出了一种新的分层对比学习方案和一种新的样本选择技术来持续训练 Android 恶意软件分类器。我们的评估表明,与之前发布的主动学习方法相比,这带来了显着的改进。我们的方法将假阴性率从 14%(最佳基线)降低到 9%,同时也降低了假阳性率(从 0.86% 到 0.48%)。此外,与过去的方法相比,我们的方法在七年时间内保持了更一致的性能。
Machine learning methods can detect Android malware with very high accuracy. However, these classifiers have an Achilles heel, concept drift: they rapidly become out of date and ineffective, due to the evolution of malware apps and benign apps. Our research finds that, after training an Android malware classifier on one year's worth of data, the F1 score quickly dropped from 0.99 to 0.76 after 6 months of deployment on new test samples. In this paper, we propose new methods to combat the concept drift problem of Android malware classifiers. Since machine learning technique needs to be continuously deployed, we use active learning: we select new samples for analysts to label, and then add the labeled samples to the training set to retrain the classifier. Our key idea is, similarity-based uncertainty is more robust against concept drift. Therefore, we combine contrastive learning with active learning. We propose a new hierarchical contrastive learning scheme, and a new sample selection technique to continuously train the Android malware classifier. Our evaluation shows that this leads to significant improvements, compared to previously published methods for active learning. Our approach reduces the false negative rate from 14% (for the best baseline) to 9%, while also reducing the false positive rate (from 0.86% to 0.48%). Also, our approach maintains more consistent performance across a seven-year time period than past methods.