Automated Adversary-in-the-Loop Cyber-Physical Defense Planning

Automated Adversary-in-the-Loop Cyber-Physical Defense Planning
复制标题

DOI:
10.1145/3596222
复制
发表时间:
2023-05
影响因子:
2.3
通讯作者:
Sandeep Banik;Thiagarajan Ramachandran;A. Bhattacharya;S. Bopardikar
Sandeep Banik;Thiagarajan Ramachandran;A. Bhattacharya;S. Bopardikar
中科院分区:
--
文献类型:
--
作者:
Sandeep Banik;Thiagarajan Ramachandran;A. Bhattacharya;S. Bopardikar

文献摘要

相似文献

由于网络和物理组件的紧密集成和操作复杂性,网络物理系统(CPS)的安全性继续面临新的挑战。为了应对这些挑战,本文提出了一种基于领域感知的优化方法,通过模拟利用系统漏洞、CPS互连和物理组件动态的循环中的战略对手,以自动化的方式确定有效的CPS防御策略。我们的方法建立在基于马尔可夫决策过程(MDP)的对抗决策模型之上,该模型确定了CPS攻击图上的最佳网络(离散)和物理(连续)攻击行为。防御规划问题被建模为对手和防御者之间的非零和博弈。我们使用无模型强化学习方法来解决对手的问题,作为防御策略的函数。然后,我们使用贝叶斯优化(BO)为防御者找到一个近似的最佳响应,以针对产生的对手策略加强网络。这个过程要反复多次,以改进双方的策略。我们用智能建筑系统作为物理过程,在勒索软件启发的图形上展示了我们的方法的有效性。数值研究表明,对于不同防御者特定的网络加固成本,该方法收敛于纳什均衡。
Security of cyber-physical systems (CPS) continues to pose new challenges due to the tight integration and operational complexity of the cyber and physical components. To address these challenges, this article presents a domain-aware, optimization-based approach to determine an effective defense strategy for CPS in an automated fashion—by emulating a strategic adversary in the loop that exploits system vulnerabilities, interconnection of the CPS, and the dynamics of the physical components. Our approach builds on an adversarial decision-making model based on a Markov Decision Process (MDP) that determines the optimal cyber (discrete) and physical (continuous) attack actions over a CPS attack graph. The defense planning problem is modeled as a non-zero-sum game between the adversary and defender. We use a model-free reinforcement learning method to solve the adversary’s problem as a function of the defense strategy. We then employ Bayesian optimization (BO) to find an approximate best-response for the defender to harden the network against the resulting adversary policy. This process is iterated multiple times to improve the strategy for both players. We demonstrate the effectiveness of our approach on a ransomware-inspired graph with a smart building system as the physical process. Numerical studies show that our method converges to a Nash equilibrium for various defender-specific costs of network hardening.