SATTVA: SpArsiTy inspired classificaTion of malware VAriants

SATTVA: SpArsiTy inspired classificaTion of malware VAriants
复制标题

SATTVA:受稀疏性启发的恶意软件变体分类

DOI:
10.1145/2756601.2756616
复制
发表时间:
2015
期刊:
Proceedings of the 3rd ACM Workshop on Information Hiding and Multimedia Security
影响因子:
--
通讯作者:
B. S. Manjunath
B. S. Manjunath
中科院分区:
--
文献类型:
--
作者:
L. Nataraj;S. Karthikeyan;B. S. Manjunath

文献摘要

被引文献

相似文献

如今生成的恶意软件数量正以惊人的速度增长。然而,多项研究表明,大多数新恶意软件只是现有恶意软件的变种。快速检测这些变体在阻止新攻击方面发挥着有效作用。在本文中,我们提出了一种使用稀疏表示框架检测恶意软件变体的新方法。利用大多数恶意软件变体在结构上存在微小差异的事实,我们将新的/未知的恶意软件样本建模为训练集中其他恶意软件的稀疏线性组合。残留误差最小的类别被分配给未知恶意软件。对两个标准恶意软件数据集 Malheur 数据集和 Malimg 数据集的实验表明,我们的方法优于当前最先进的方法,分类准确率分别为 98.55% 和 92.83%。此外,通过使用置信度度量来拒绝异常值,我们在两个数据集上获得了 100% 的准确度,但代价是丢弃了一小部分异常值。最后,我们在两个大型恶意软件数据集上评估我们的技术:攻击性计算数据集(2,124 个类别,42,480 个恶意软件)和 Anubis 数据集(209 个类别,36,784 个样本)。在这两个数据集上,我们的方法获得了 77% 的平均分类准确率,从而使其适用于现实世界的恶意软件分类。
There is an alarming increase in the amount of malware that is generated today. However, several studies have shown that most of these new malware are just variants of existing ones. Fast detection of these variants plays an effective role in thwarting new attacks. In this paper, we propose a novel approach to detect malware variants using a sparse representation framework. Exploiting the fact that most malware variants have small differences in their structure, we model a new/unknown malware sample as a sparse linear combination of other malware in the training set. The class with the least residual error is assigned to the unknown malware. Experiments on two standard malware datasets, Malheur dataset and Malimg dataset, show that our method outperforms current state of the art approaches and achieves a classification accuracy of 98.55\% and 92.83\% respectively. Further, by using a confidence measure to reject outliers, we obtain 100\% accuracy on both datasets, at the expense of throwing away a small percentage of outliers. Finally, we evaluate our technique on two large scale malware datasets: Offensive Computing dataset (2,124 classes, 42,480 malware) and Anubis dataset (209 classes, 36,784 samples). On both datasets our method obtained an average classification accuracy of 77\%, thus making it applicable to real world malware classification.