Attacker-Centric View of a Detection Game against Advanced Persistent Threats

Attacker-Centric View of a Detection Game against Advanced Persistent Threats
复制标题

针对高级持续威胁的检测游戏以攻击者为中心的观点

DOI:
10.1109/tmc.2018.2814052
复制
发表时间:
2018-11-01
影响因子:
7.9
通讯作者:
Poor, H. Vincent
Poor, H. Vincent
中科院分区:
计算机科学2区
文献类型:
--
作者:
Xiao, Liang;Xu, Dongjin;Poor, H. Vincent

文献摘要

被引文献

相似文献

高级持续性威胁 (APT) 是网络安全的主要威胁,每年都会造成重大的财务和隐私损失。本文应用累积前景理论(CPT)来研究网络系统和 APT 攻击者分别做出主观决定来选择扫描间隔和攻击间隔时之间的相互作用。概率扭曲效应和框架效应都被用来模拟在纯策略博弈中的不确定攻击持续时间和混合策略博弈中的扫描间隔下,最终用户的主观决策与预期效用理论控制的客观决策的偏差。基于CPT的APT检测博弈结合了网络系统主观攻击者和安全代理的概率权重失真和框架效应,而不是像早期APT检测前景理论研究那样采用离散决策权重。推导了APT检测博弈的纳什均衡,表明如果评估效用的参照系较大,则主观攻击者会变得追求风险;如果评估效用的参照系较小,则主观攻击者会变得风险厌恶。提出了一种基于策略爬山(PHC)的检测方案,以增加策略不确定性来欺骗动态博弈中的攻击者,并开发了一种利用类似场景中的经验来初始化质量值的“热启动”技术,以加快基于 PHC 的检测的学习速度。提出了一个移动网络的实际例子来评估所提出的检测策略的性能。仿真结果表明,与标准 Q 学习策略相比,所提出的策略可以在存在攻击者的情况下提高检测性能,具有更高的数据保护级别和云实用性。
Advanced persistent threats (APTs) are a major threat to cyber-security, causing significant financial and privacy losses each year. In this paper, cumulative prospect theory (CPT) is applied to study the interactions between a cyber system and an APT attacker when each of them makes subjective decisions to choose their scan interval and attack interval, respectively. Both the probability distortion effect and the framing effect are applied to model the deviation of subjective decisions of end-users from the objective decisions governed by expected utility theory, under uncertain attack durations in a pure-strategy game and scan interval in a mixed-strategy game. The CPT-based APT detection game incorporates both the probability weighting distortion and the framing effect of the subjective attacker and security agent of the cyber system, rather than discrete decision weights, as in earlier prospect theoretic study of APT detection. The Nash equilibria of the APT detection game are derived, showing that a subjective attacker becomes risk-seeking if the frame of reference for evaluating the utility is large, and becomes risk-averse if the frame of reference for evaluating the utility is small. A policy hill-climbing (PHC) based detection scheme is proposed to increase the policy uncertainty to fool the attacker in the dynamic game, and a “hotbooting” technique that exploits experiences in similar scenarios to initialize the quality values is developed to accelerate the learning speed of PHC-based detection. A practical example of a mobile network is presented to evaluate the performance of the proposed detection strategy. Simulation results show that the proposed strategy can improve detection performance with a higher data protection level and utilities of the cloud in the presence of an attacker compared with a standard Q-learning strategy.