A machine learning based approach to detect malicious android apps using discriminant system calls

A machine learning based approach to detect malicious android apps using discriminant system calls
复制标题

DOI:
10.1016/j.future.2018.11.021
复制
发表时间:
2019-05-01
影响因子:
7.5
通讯作者:
Conti, Mauro
Conti, Mauro
中科院分区:
计算机科学2区
文献类型:
--
作者:
Vinod, P.;Zemmari, Akka;Conti, Mauro

文献摘要

被引文献

相似文献

Android框架的开放性和用户信任度的提高引起了恶意软件作者的注意。从众多应用程序商店下载应用程序(简称app)的势头刺激了移动的恶意软件的扩散。现在的威胁是由于恶意软件的复杂性被编写来绕过基于签名的检测器。在本文中,我们研究了系统调用来解决Android操作系统上的移动的恶意软件。为此,我们首先使用机器学习来提取系统调用。然后,我们进行了经验估计的系统调用来自不同的数据集,采用人类的互动和随机输入。在使用两种特征选择方法,即加权系统调用的绝对差异(ADWSC)和使用大群体测试的排名系统调用(RSLPT)对合成系统调用进行了深入的实验后,我们在五个数据集上验证了结果。在1.0的曲线下面积中生成的所有分类器的准确度超过99.9%,表明所提出的方法的适当性和有效性。最后,我们评估了分类器对抗对抗攻击的有效性,发现分类器容易受到数据中毒和标签翻转攻击。通过中毒恶意软件样本创建的对抗性示例导致分类器性能在扰动12-18个突出属性时显着下降。此外,我们实施了类标签中毒攻击,在改变50个恶意训练实例的标签时,分类准确率下降了50%。(C)2018爱思唯尔B. V.保留所有权利。
The openness of Android framework and the enhancement of users trust have gained the attention of malware writers. The momentum of downloaded applications (app for short) from numerous app stores has stimulated the proliferation of mobile malware. Now the threat is due to the sophistication in malware being written to bypass signature-based detectors. In this paper, we investigate system calls to tackle mobile malware on Android operating system. To do so, we first employed machine learning to extract system calls. We then performed the empirical estimation of system calls derived from diverse datasets employing human interaction and random inputs. After accomplishing intensive experiments on synthesized system calls with two feature selection approach, namely Absolute Difference of Weighted System Calls (ADWSC) and Ranked System Calls using Large Population Test (RSLPT), we validated the results on five datasets. All classifiers generated in Area Under Curve of 1.0 with an accuracy exceeding 99.9% suggest the appropriateness and efficacy of the proposed approach. Finally, we evaluated the effectiveness of classifier against adversarial attacks and found that the classifiers are vulnerable to data poisoning and label flipping attacks. Adversarial examples created by poisoning malware samples resulted in the significant drop of classifier performance on perturbing 12-18 prominent attributes. Moreover, we implemented class label poisoning attacks which brought down the classification accuracy by 50% on altering labels of 50 malicious training instances. (C) 2018 Elsevier B.V. All rights reserved.