Low latency cyberattack detection in smart grids with deep reinforcement learning

Low latency cyberattack detection in smart grids with deep reinforcement learning
复制标题

通过深度强化学习在智能电网中进行低延迟网络攻击检测

DOI:
10.1016/j.ijepes.2022.108265
复制
发表时间:
2022
影响因子:
5.2
通讯作者:
Wu, Jingxian
Wu, Jingxian
中科院分区:
工程技术2区
文献类型:
--
作者:
Li, Yaze;Wu, Jingxian

文献摘要

相似文献

研究了基于深度强化学习(DRL)的智能电网网络攻击低延迟检测方法。低延迟检测的目标是在保证高检测精度的同时最小化检测延迟。这与传统的检测方法主要关注检测精度,很少关注检测延迟不同。降低检测延迟可以减少恢复时间,从而最大限度地减少网络攻击造成的业务中断或经济损失。由于检测延迟是主要的设计度量,因此该算法采用带有扩展卡尔曼滤波(EKF)的非线性动态交流系统模型来实时捕获电网状态转换,而文献中的许多其他工作使用简化的线性直流模型。在马尔可夫决策过程(MDP)的框架上,利用连续状态空间深度q网络(DQN)开发了DRL检测算法。新的DQN设计有两个主要创新。首先,将MDP状态设计为交流动态估计残差的rao统计量的滑动窗口;所提出的状态公式可以实时准确地捕捉动态功率状态转换。其次,设计了一个新的奖励函数,允许在检测延迟和检测精度之间进行灵活的权衡。可以通过调整奖励函数中的单个参数来调整延迟-精度权衡。仿真结果表明,本文提出的基于dqn的DRL检测算法可以实现非常低的检测延迟和较高的检测精度。
This paper focuses on low latency detection of cyberattacks in smart grids with deep reinforcement learning (DRL). The objective of low latency detection is to minimize detection delay while ensuring high detection accuracy. This is different from conventional detection methods that focus mainly on detection accuracy and pay little attention to detection delay. A lower detection delay can reduce recovery time, thus minimizing service interruption or economic losses due to cyberattacks. Since detection delay is the main design metric, the algorithm is developed by using a non-linear dynamic AC system model with an extended Kalman filter (EKF) to capture power grid state transitions in real-time, while many other works in the literature use a simplified linear DC model. The DRL detection algorithm is developed by using a continuous state space deep Q-network (DQN) on the framework of a Markov decision process (MDP). The new DQN design has two main innovations. First, the MDP state is designed as a sliding window of Rao-statistics of the AC dynamic state estimation residues. The proposed state formulation can accurately capture dynamic power state transitions in real-time. Second, a new reward function is designed to allow a flexible trade-off between detection delays and detection accuracy. The delay-accuracy trade-off can be adjusted by tuning a single parameter in the reward function. Simulation results show that the proposed DQN-based DRL detection algorithm can achieve very low detection delays with high detection accuracy.