Reinforcement Learning Based Recloser Control for Distribution Cables With Degraded Insulation Level

Reinforcement Learning Based Recloser Control for Distribution Cables With Degraded Insulation Level
复制标题

DOI:
10.1109/tpwrd.2020.3002503
复制
发表时间:
2020-06
影响因子:
4.4
通讯作者:
Qiushi Cui;Syed Muhammad Yousaf Hashmy;Yang Weng;M. Dyer
Qiushi Cui;Syed Muhammad Yousaf Hashmy;Yang Weng;M. Dyer
中科院分区:
工程技术2区
文献类型:
--
作者:
Qiushi Cui;Syed Muhammad Yousaf Hashmy;Yang Weng;M. Dyer

文献摘要

被引文献

相似文献

公用事业不断观察老化电缆的电缆故障,这些电缆具有未知的退化基本绝缘水平(BIL)。其根本原因之一是与断路器重合闸相关的瞬态过电压(TOV)。为了解决这一问题,研究者们提出了一系列的受控切换方法,其中大部分属于确定性控制。然而,在电力系统中,特别是在配电网中,开关暂态过程受到随机性的冲击。针对暂态过电压建模困难的问题,提出了一种在不确定性和噪声存在下的无模型重合器随机控制方法。具体来说,为了捕获高维动力学模式,我们通过将时间序列奖励机制纳入深度Q网络(DQN)来制定重合器控制问题。同时,我们将我们对问题的物理理解嵌入到动作概率分配中,提出了一种不可行动作空间消除算法。通过PSCAD仿真,我们首先揭示了负载类型对电缆的TOV的影响。然后,为了减少所提出的强化学习(RL)控制方法在不同应用中的训练负担,我们建立了一个学习后的知识转移方法。在与我们的工业合作伙伴进行验证后,我们展示了几条学习曲线,以显示增强的性能。由于所提出的时序奖励机制和不可行动作消除方法,学习效率被证明是突出的。此外,知识转移的结果表明,方法的泛化能力。最后,与传统的方法进行了比较。实验结果表明,在三种方法中,该方法对抑制TOV现象最为有效.
Utilities continuously observe cable failures on aged cables that have an unknown degraded basic insulation level (BIL). One of the root causes is the transient overvoltage (TOV) associated with circuit breaker reclosing. To solve this problem, researchers propose a series of controlled switching methods, most of which belong to deterministic control. However, in power systems, especially in distribution networks, the switching transient is buffeted by stochasticity. Since it is hard to model transient overvoltage due to its complexity, we propose a model-free stochastic control method for reclosers under the existence of uncertainty and noise. Concretely, to capture high-dimensional dynamics patterns, we formulate the recloser control problem by incorporating the temporal sequence reward mechanism into a deep Q-network (DQN). Meanwhile, we embed our physical understanding of the problem into the action probability allocation and develop an infeasible-action-space-elimination algorithm. Through PSCAD simulation, we first reveal the impact of load types on cables’ TOVs. Then, to reduce the training burden for the proposed reinforcement learning (RL) control method in different applications, we establish a post-learning knowledge transfer method. After the validation with our industrial partner, we exhibit several learning curves to show the enhanced performance. The learning efficiency is proved to be outstanding due to the proposed time sequence reward mechanism and infeasible action elimination method. Moreover, the results on knowledge transfer demonstrate the capability of method generalization. Finally, a comparison with conventional methods is conducted. It illustrates the proposed method is most effective in mitigating the TOV phenomenon among three methods.