Disentangled Representation Learning in Heterogeneous Information Network for Large-scale Android Malware Detection in the COVID-19 Era and Beyond

Disentangled Representation Learning in Heterogeneous Information Network for Large-scale Android Malware Detection in the COVID-19 Era and Beyond
复制标题

DOI:
10.1609/aaai.v35i9.16947
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Shifu Hou;Yujie Fan;Mingxuan Ju;Yanfang Ye;Wenqiang Wan;Kui Wang;Y. Mei;Qi Xiong;Fudong Shao
Shifu Hou;Yujie Fan;Mingxuan Ju;Yanfang Ye;Wenqiang Wan;Kui Wang;Y. Mei;Qi Xiong;Fudong Shao
中科院分区:
其他
文献类型:
--
作者:
Shifu Hou;Yujie Fan;Mingxuan Ju;Yanfang Ye;Wenqiang Wan;Kui Wang;Y. Mei;Qi Xiong;Fudong Shao

文献摘要

相似文献

在抗击新冠肺炎疫情的过程中,许多社交活动都转向了网络;社会对复杂的网络空间的压倒性依赖使其安全比以往任何时候都更加重要。在本文中,我们提出并开发了一个名为Dr.HIN的智能系统,以保护用户免受COVID-19时代及以后不断发展的Android恶意软件攻击。在Dr.HIN中,除了应用内容之外,我们建议考虑应用、开发者和移动设备之间更高层次的语义和社会关系,以全面描述Android应用;然后引入结构化异构信息网络(HIN)对复杂关系进行建模,并利用元路径引导策略从HIN中学习节点(即应用程序)表示。由于在复杂的开发生态系统中,恶意软件的表征可能与良性应用高度纠缠,因此学习隐藏在HIN嵌入中的潜在解释因素以检测不断演变的恶意软件提出了新的挑战。为了应对这一挑战,我们建议整合从不同角度(即应用内容、应用作者、应用安装)生成的域先验,设计一个对抗性解纠集器,以分离隐藏在HIN嵌入中的不同的、信息丰富的变化因素,用于大规模Android恶意软件检测。这是对HIN数据进行解纠缠表示学习的首次尝试。通过与基线和流行的移动安全产品进行比较,基于安全行业大规模和真实样本采集的实验结果表明,Dr.HIN在不断发展的Android恶意软件检测方面具有良好的性能。
In the fight against the COVID-19 pandemic, many social activities have moved online; society's overwhelming reliance on the complex cyberspace makes its security more important than ever. In this paper, we propose and develop an intelligent system named Dr.HIN to protect users against the evolving Android malware attacks in the COVID-19 era and beyond. In Dr.HIN, besides app content, we propose to consider higher-level semantics and social relations among apps, developers and mobile devices to comprehensively depict Android apps; and then we introduce a structured heterogeneous information network (HIN) to model the complex relations and exploit meta-path guided strategy to learn node (i.e., app) representations from HIN. As the representations of malware could be highly entangled with benign apps in the complex ecosystem of development, it poses a new challenge of learning the latent explanatory factors hidden in the HIN embeddings to detect the evolving malware. To address this challenge, we propose to integrate domain priors generated from different views (i.e., app content, app authorship, app installation) to devise an adversarial disentangler to separate the distinct, informative factors of variations hidden in the HIN embeddings for large-scale Android malware detection. This is the first attempt of disentangled representation learning in HIN data. Promising experimental results based on the large-scale and real sample collections from security industry demonstrate the performance of Dr.HIN in evolving Android malware detection, by comparison with baselines and popular mobile security products.