Picking on the family: Disrupting android malware triage by forcing misclassification

Picking on the family: Disrupting android malware triage by forcing misclassification
复制标题

DOI:
10.1016/j.eswa.2017.11.032
复制
发表时间:
2018-04-01
影响因子:
8.5
通讯作者:
Clark, David
Clark, David
中科院分区:
计算机科学1区
文献类型:
--
作者:
Calleja, Alejandro;Martin, Alejandro;Clark, David

文献摘要

被引文献

相似文献

机器学习分类算法被广泛应用于不同的恶意软件分析问题,因为它们可以从示例中学习并在很少的人类输入中表现出色。用例包括在可疑恶意软件分类过程中根据家人的恶意样本标记。但是,自动化算法很容易受到攻击。攻击者可以仔细操纵样品以迫使算法产生特定的输出。在本文中,我们讨论了对Android恶意软件分类器的一次攻击。我们设计并实施了一种称为Lagodroid的原型工具,该工具将其作为输入恶意软件样本和目标家族,并修改样品以使其被归类为属于该家族的同时,同时保留其原始语义。我们的技术取决于搜索过程,该过程生成原始样本的变体而无需修改其语义。我们根据各种静态功能测试了针对ReceAldroid的Lagodoid,这是最近的开源,Android恶意软件分类器。雌雄同体成功地迫使Drebin数据集中的29个代表性恶意软件家族中的28个成功迫使错误分类。值得注意的是,它通过仅修改原始恶意软件的单个功能来做到这一点。平均而言,它在第一次搜索迭代中发现了第一个回避样本,并在4次迭代中收集到100%逃避人群。最后,我们介绍了ReveAldroid*,这是一种更强大的分类器,它实现了其他对抗性学习域中提出的几种技术。我们的实验表明,recealdroid*可以正确检测到高多洛氏菌产生的变体的99%。 (c)2017年作者。由Elsevier Ltd.出版。
Machine learning classification algorithms are widely applied to different malware analysis problems because of their proven abilities to learn from examples and perform relatively well with little human input. Use cases include the labelling of malicious samples according to families during triage of suspected malware. However, automated algorithms are vulnerable to attacks. An attacker could carefully manipulate the sample to force the algorithm to produce a particular output. In this paper we discuss one such attack on Android malware classifiers. We design and implement a prototype tool, called lagoDroid, that takes as input a malware sample and a target family, and modifies the sample to cause it to be classified as belonging to this family while preserving its original semantics. Our technique relies on a search process that generates variants of the original sample without modifying their semantics. We tested lagoDroid against RevealDroid, a recent, open source, Android malware classifier based on a variety of static features. IagoDroid successfully forces misclassification for 28 of the 29 representative malware families present in the DREBIN dataset. Remarkably, it does so by modifying just a single feature of the original malware. On average, it finds the first evasive sample in the first search iteration, and converges to a 100% evasive population within 4 iterations. Finally, we introduce RevealDroid*, a more robust classifier that implements several techniques proposed in other adversarial learning domains. Our experiments suggest that RevealDroid* can correctly detect up to 99% of the variants generated by lagoDroid. (C) 2017 The Authors. Published by Elsevier Ltd.