A Security-Constrained Reinforcement Learning Framework for Software Defined Networks

A Security-Constrained Reinforcement Learning Framework for Software Defined Networks
复制标题

软件定义网络的安全约束强化学习框架

DOI:
10.1109/icc42927.2021.9500763
复制
发表时间:
2021
期刊:
ICC 2021 - IEEE International Conference on Communications
影响因子:
--
通讯作者:
D. Verma
D. Verma
中科院分区:
--
文献类型:
--
作者:
Anand Mudgerikar;E. Bertino;Jorge Lobo;D. Verma

文献摘要

被引文献

相似文献

强化学习(RL)是一种构建“智能”SDN控制器的有效技术,因为它具有无模型的性质,并且能够在线学习策略,而无需大量的训练数据。然而,由于RL代理旨在最大限度地提高功能并在没有约束的情况下探索环境,因此安全性可能会受到破坏。在本文中,我们提出了Jarvis-SDN,这是一个RL框架,通过考虑安全性来限制探索。在Jarvis-SDN中,RL代理学习“智能策略”,最大化功能,但不以安全为代价。从入侵检测系统(IDS)数据集获得的基于标准网络流的攻击特征不能用作策略,因为它们不符合RL框架的状态模型,因此具有较差的准确性和较高的误报率。为了解决这个问题,在Jarvis-SDN中用于约束探索的安全策略以半监督的方式从IDS数据集的数据包捕获中以“部分攻击签名”的形式学习,然后将其编码在基于RL的优化框架的目标函数中。这些签名是使用深度Q网络(DQN)学习的。我们的分析表明,对于常见的网络攻击,基于DQN的攻击特征比经典的机器学习技术(例如决策树、随机森林和深度神经网络(DNN))表现得更好。我们实例化我们的SDN控制器框架,目标是智能速率控制,以进一步分析攻击签名的有效性。
Reinforcement Learning (RL) is an effective technique for building ‘smart’ SDN controllers because of its model-free nature and ability to learn policies online without requiring extensive training data. However, as RL agents are geared to maximize functionality and explore the environment without constraints, security can be breached. In this paper, we propose Jarvis-SDN, a RL framework that constrains explorations by taking security into account. In Jarvis-SDN, the RL agent learns ‘intelligent policies’ which maximize functionality but not at the cost of security. Standard network flow based attack sig-natures obtained from intrusion detection system (IDS) datasets cannot be used as policies because they do not conform to the state model of the RL framework and thus have poor accuracy and high false positives. To address such issue, the security policies for constraining explorations in Jarvis-SDN are learnt in a semi-supervised manner in the form of ‘partial attack signatures’ from packet captures of IDS datasets that are then encoded in the objective function of the RL based optimization framework. These signatures are learnt using Deep Q-Networks (DQN). Our analysis shows that DQN based attack signatures perform better than classical machine learning techniques, like decision trees, random forests and deep neural networks (DNN), for common network attacks. We instantiate our framework for a SDN controller with the goal of intelligent rate control to further analyze the effectiveness of the attack signatures.