Robust Moving Target Defense Against Unknown Attacks: A Meta-reinforcement Learning Approach

Robust Moving Target Defense Against Unknown Attacks: A Meta-reinforcement Learning Approach
复制标题

DOI:
10.1007/978-3-031-26369-9_6
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Henger Li;Zizhan Zheng
Henger Li;Zizhan Zheng
中科院分区:
其他
文献类型:
--
作者:
Henger Li;Zizhan Zheng

文献摘要

相似文献

移动目标防御(MTD)提供了一个系统的框架,以实现在先进和隐形攻击的存在下的主动防御。为了在面对未知攻击策略时获得稳健的MTD,一种很有前途的方法是将连续的攻击者-防御者互动建模为两个人的马尔可夫博弈,并将防御者的问题表述为寻找Stackelberg均衡(或其变体),其中防御者和领导者以及攻击者作为追随者。然而,为了解决这个问题,现有的方法通常假设攻击者的类型(包括其物理、认知、计算能力和约束)是已知的,或者是从已知的分布中采样的。前者在实践中很少成立,因为对攻击者类型的初始猜测通常是不准确的,而后者即使在训练MTD策略和应用MTD策略之间没有分布变化的情况下也会导致次优解决方案。另一方面,在安全敏感的域中动态地收集足够的样本来覆盖各种攻击场景通常是不可行的。为了解决这一困境,我们在这项工作中提出了一个基于两阶段元强化学习的MTD框架。在训练阶段,使用从一组可能的攻击中采样的经验来学习元mtd策略。在测试阶段,使用少量样本快速调整元策略以应对实际攻击。研究表明,在面对不确定/未知攻击者类型和攻击行为时,我们的两阶段MTD防御获得了出色的性能。
Moving target defense (MTD) provides a systematic framework to achieving proactive defense in the presence of advanced and stealthy attacks. To obtain robust MTD in the face of unknown attack strategies, a promising approach is to model the sequential attacker-defender interactions as a two-player Markov game, and formulate the defender’s problem as finding the Stackelberg equilibrium (or a variant of it) with the defender and the leader and the attacker as the follower. To solve the game, however, existing approaches typically assume that the attacker type (including its physical, cognitive, and computational abilities and constraints) is known or is sampled from a known distribution. The former rarely holds in practice as the initial guess about the attacker type is often inaccurate, while the latter leads to suboptimal solutions even when there is no distribution shift between when the MTD policy is trained and when it is applied. On the other hand, it is often infeasible to collect enough samples covering various attack scenarios on the fly in security-sensitive domains. To address this dilemma, we propose a two-stage meta-reinforcement learning based MTD framework in this work. At the training stage, a meta-MTD policy is learned using experiences sampled from a set of possible attacks. At the test stage, the meta-policy is quickly adapted against a real attack using a small number of samples. We show that our two-stage MTD defense obtains superb performance in the face of uncertain/unknown attacker type and attack behavior.