An Optimized Positive-Unlabeled Learning Method for Detecting a Large Scale of Malware Variants

An Optimized Positive-Unlabeled Learning Method for Detecting a Large Scale of Malware Variants
复制标题

DOI:
10.1109/dsc47296.2019.8937650
复制
发表时间:
2019-11
期刊:
2019 IEEE Conference on Dependable and Secure Computing (DSC)
影响因子:
--
通讯作者:
Jixin Zhang;Mohammad Faham Khan;Xiaodong Lin;Zheng Qin
Jixin Zhang;Mohammad Faham Khan;Xiaodong Lin;Zheng Qin
中科院分区:
其他
文献类型:
--
作者:
Jixin Zhang;Mohammad Faham Khan;Xiaodong Lin;Zheng Qin

文献摘要

被引文献

相似文献

恶意软件(Malware)能够快速演变成许多不同的变种并逃避现有的检测机制,使得传统的基于签名的恶意软件检测系统失效。许多研究人员通过使用机器学习提出了先进的恶意软件检测技术。虽然基于机器学习的技术在检测各种恶意软件变体方面表现良好,但在满足行业中的真实的场景时仍然存在一些问题。由于新的恶意软件变体数量增长迅速,标记数据成本高昂,需要大量劳动力,因此公司无法标记这些样本中的每一个。他们倾向于标记一小部分恶意软件样本,并将其余未标记的样本视为良性样本,其中原始恶意软件样本被视为错误标记。这导致决策边界的偏差,严重限制了准确性。为了解决这样的问题,本文提出了一种成本敏感的提升方法,用恶意未标记的可执行文件训练无偏检测模型,以提高准确性。沿着,为了有效地检测恶意软件变种,我们提出了一个字节共生矩阵作为可执行文件的字节流的表示,以直接检测恶意软件变种。实验结果表明,当未标记数据中包含不同比例的误标记阳性数据时,优化后的机器学习方法可以达到80%~ 90%的准确率,而原有的机器学习方法只能达到50%~ 85%的准确率。
Malicious softwares (Malware) are able to quickly evolve into many different variants and evade existing detection mechanisms, rendering the ineffectiveness of traditional signature-based malware detection systems. Many researchers have proposed advanced malware detection techniques by using Machine Learning. Although the machine learning based techniques perform well in detecting a wide range of malware variants, there still remain some problems when meeting the real scene in the industry. Since the volume of new malware variants grows fast and labelling data is expensive and takes a lot of labor, companies cannot label every one of those samples. They tend to label a small part of the malware samples and treat the rest of the unlabeled samples as benign samples in which the original malware samples are treated as mislabeled. This causes a bias of decision boundary which severely limits the accuracy. To address such a problem, in this paper, we propose a cost-sensitive boosting method to train an unbiased detection model with the malicious-unlabeled executables to improve the accuracy. Along with that, in order to detect malware variants efficiently, we propose a byte co-occurrence matrix as a representation of byte streams of executables to detect malware variants directly. Experimental results show that the machine learning methods optimized by our approach can achieve 80% to 90% accuracy while the original machine learning methods can only achieve 50% to 85% accuracy when the unlabeled data contain different rates of mislabeled positive data.