Transcend: Detecting Concept Drift in Malware Classification Models

Transcend: Detecting Concept Drift in Malware Classification Models
复制标题

DOI:
--
复制
发表时间:
2017
期刊:
--
影响因子:
--
通讯作者:
Roberto Jordaney;K. Sharad;Santanu Kumar Dash;Zhi Wang;D. Papini;I. Nouretdinov;L. Cavallaro
Roberto Jordaney;K. Sharad;Santanu Kumar Dash;Zhi Wang;D. Papini;I. Nouretdinov;L. Cavallaro
中科院分区:
其他
文献类型:
--
作者:
Roberto Jordaney;K. Sharad;Santanu Kumar Dash;Zhi Wang;D. Papini;I. Nouretdinov;L. Cavallaro

文献摘要

被引文献

相似文献

构建恶意软件行为的机器学习模型被广泛接受为有效恶意软件分类的灵丹妙药。然而,构建可持续学习模型的一个关键要求是在各种恶意软件样本上进行训练。不幸的是,恶意软件发展迅速,因此很难(如果不是不可能的话)概括学习模型以反映未来的、以前看不到的行为。因此,从长远来看,大多数恶意软件分类器变得不可持续,随着恶意软件的不断发展而迅速过时。在这项工作中,我们提出了Transcend,这是一个在部署过程中识别体内老化分类模型的框架,远远早于机器学习模型的性能开始下降。这与传统方法有很大的不同,传统方法在观察到性能不佳时会回顾性地重新训练老化模型。我们的方法使用部署期间看到的样本与用于训练模型的样本进行统计比较,从而构建预测质量的指标。我们展示了Transcend如何基于Android和Windows恶意软件的两个单独的案例研究来识别概念漂移,在模型由于过时的训练而开始做出一贯糟糕的决策之前发出红旗。
Building machine learning models of malware behavior is widely accepted as a panacea towards effective malware classification. A crucial requirement for building sustainable learning models, though, is to train on a wide variety of malware samples. Unfortunately, malware evolves rapidly and it thus becomes hard—if not impossible—to generalize learning models to reflect future, previously-unseen behaviors. Consequently, most malware classifiers become unsustainable in the long run, becoming rapidly antiquated as malware continues to evolve. In this work, we propose Transcend, a framework to identify aging classification models in vivo during deployment, much before the machine learning model’s performance starts to degrade. This is a significant departure from conventional approaches that retrain aging models retrospectively when poor performance is observed. Our approach uses a statistical comparison of samples seen during deployment with those used to train the model, thereby building metrics for prediction quality. We show how Transcend can be used to identify concept drift based on two separate case studies on Android andWindows malware, raising a red flag before the model starts making consistently poor decisions due to out-of-date training.