Recurrent Neural Networks Based Online Behavioural Malware Detection Techniques for Cloud Infrastructure

Recurrent Neural Networks Based Online Behavioural Malware Detection Techniques for Cloud Infrastructure
复制标题

DOI:
10.1109/access.2021.3077498
复制
发表时间:
2021-01-01
期刊:
影响因子:
3.9
通讯作者:
Sandhu, Ravi
Sandhu, Ravi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kimmel, Jeffrey C.;Mcdole, Andrew D.;Sandhu, Ravi

文献摘要

被引文献

相似文献

一些组织正在利用云技术和资源来运行一系列应用程序。这些服务可帮助企业节省底层基础设施的硬件管理、可扩展性和可维护性问题。亚马逊、微软和谷歌等主要云服务提供商(CSP)提供基础设施即服务(IaaS),以满足这些企业不断增长的需求。云平台利用率的提高使其成为攻击者的一个有吸引力的目标,从而使云服务的安全性成为CSP的首要任务。在这方面,恶意软件已被认为是对云基础设施(IaaS)最危险和最具破坏性的威胁之一。在本文中,我们研究了基于递归神经网络(RNN)的深度学习技术在云虚拟机(VM)中检测恶意软件的有效性。我们专注于两种主要的RNN架构:长短期记忆RNN(LSTM)和双向RNN(BIDI)。这些模型根据运行时细粒度进程系统特性(如CPU、内存和磁盘利用率),了解恶意软件随时间的行为。我们在40,680个恶意和良性样本的数据集上评估了我们的方法。进程级特征是使用在开放的在线云环境中运行的真实的恶意软件收集的,没有任何限制,这对于模拟实际的云提供商设置以及捕获隐身和复杂恶意软件的真实行为非常重要。我们的LSTM和BIDI模型在不同的评估指标下都达到了超过99%的高检测率。此外,进行分析研究,以了解输入数据表示的意义。我们的研究结果表明,在特定情况下,输入排序确实对训练的RNN模型的性能有一定的影响。
Several organizations are utilizing cloud technologies and resources to run a range of applications. These services help businesses save on hardware management, scalability and maintainability concerns of underlying infrastructure. Key cloud service providers (CSPs) like Amazon, Microsoft and Google offer Infrastructure as a Service (IaaS) to meet the growing demand of such enterprises. This increased utilization of cloud platforms has made it an attractive target to the attackers, thereby, making the security of cloud services a top priority for CSPs. In this respect, malware has been recognized as one of the most dangerous and destructive threats to cloud infrastructure (IaaS). In this paper, we study the effectiveness of Recurrent Neural Networks (RNNs) based deep learning techniques for detecting malware in cloud Virtual Machines (VMs). We focus on two major RNN architectures: Long Short Term Memory RNNs (LSTMs) and Bidirectional RNNs (BIDIs). These models learn the behavior of malware over time based on run-time fine-grained processes system features such as CPU, memory, and disk utilization. We evaluate our approach on a dataset of 40,680 malicious and benign samples. The process level features were collected using real malware running in an open online cloud environment with no restrictions, which is important to emulate practical cloud provider settings and also capture the true behaviour of stealth and sophisticated malware. Both our LSTM and BIDI models achieve high detection rates over 99% for different evaluation metrics. In addition, an analysis study is conducted to understand the significance of input data representations. Our results suggest that in particular cases, input ordering does have some affect on the performance of the trained RNN models.