Android malware detection with weak ground truth data

Android malware detection with weak ground truth data
复制标题

DOI:
10.1109/bigdata.2016.7841008
复制
发表时间:
2016-12
期刊:
2016 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
J. DeLoach;Doina Caragea;Xinming Ou
J. DeLoach;Doina Caragea;Xinming Ou
中科院分区:
其他
文献类型:
--
作者:
J. DeLoach;Doina Caragea;Xinming Ou

文献摘要

被引文献

相似文献

对于Android恶意软件检测,精确的地面真相是一种罕见的商品。随着安全知识的发展,在某个时刻被认为是基本事实的东西可能会发生变化,曾经被认为是良性的应用程序可能会变成恶意的。数据标签中不可避免的噪声对创建有效的机器学习模型提出了挑战。我们的工作重点是学习分类器的Android恶意软件检测的方式,是方法论上的声音方面的不确定性和不断变化的地面真理的问题空间。我们利用了这样一个事实,即尽管数据标签具有不可忽视的噪声,但恶意软件标签比良性标签精确得多。虽然你可以确信一个应用程序是恶意的,但你永远不能确定一个良性的应用程序是真正良性的,或者只是一个未被检测到的恶意软件。基于这一见解,我们利用了一个修改后的逻辑回归分类器,该分类器允许我们仅从正面和未标记的数据中学习,而不对良性标签进行任何假设。我们发现标签正则化逻辑回归对于嘈杂的应用程序数据集以及具有有限数量正标签数据的数据集表现良好,这两种数据集都代表了现实世界的情况。
For Android malware detection, precise ground truth is a rare commodity. As security knowledge evolves, what may be considered ground truth at one moment in time may change, and apps once considered benign turn out to be malicious. The inevitable noise in data labels poses a challenge to creating effective machine learning models. Our work is focused on approaches for learning classifiers for Android malware detection in a manner that is methodologically sound with regard to the uncertain and ever-changing ground truth in the problem space. We leverage the fact that although data labels are unavoidably noisy, a malware label is much more precise than a benign label. While you can be confident that an app is malicious, you can never be certain that a benign app is really benign or just an undetected malware. Based on this insight, we leverage a modified Logistic Regression classifier that allows us to learn from only positive and unlabeled data, without making any assumptions about benign labels. We find Label Regularized Logistic Regression to perform well for noisy app datasets, as well as datasets where there is a limited amount of positive labeled data, both of which are representative of real-world situations.