Reinforcement Learning Based Efficiency Optimization Scheme for the DAB DC-DC Converter With Triple-Phase-Shift Modulation

Reinforcement Learning Based Efficiency Optimization Scheme for the DAB DC-DC Converter With Triple-Phase-Shift Modulation
复制标题

DOI:
10.1109/tie.2020.3007113
复制
发表时间:
2021-08-01
影响因子:
7.7
通讯作者:
Blaabjerg, Frede
Blaabjerg, Frede
中科院分区:
计算机科学1区
文献类型:
--
作者:
Tang, Yuanhong;Hu, Weihao;Blaabjerg, Frede

文献摘要

被引文献

相似文献

为了提高双有源桥(DAB)DC-DC变换器的功率效率,提出了一种基于强化学习(RL)的三相移(TPS)调制的效率优化方案。更具体地,Q学习算法作为RL的典型算法,被应用于离线训练代理以获得优化的调制策略,然后训练的代理根据当前操作环境以实时方式为DAB DC-DC转换器在线提供控制决策。主要目标是获得最佳的相移角的DAB DC-DC转换器,它可以实现最大的功率效率,通过减少功率损耗。此外,在Q学习算法的离线训练过程中,考虑了TPS调制的所有可能的操作模式。这样,就可以成功地避开常规方案中选择最佳运行模式的繁琐过程。基于这些优点,提出的效率优化方案,使用RL可以实现优良的性能,为整个负载条件和电压转换比。最后,搭建了一台1.2KW的样机,仿真和实验结果表明,采用基于强化学习的优化方案可以提高电源效率。
Aim to improve the power efficiency of the dual-active-bridge (DAB) dc-dc converter, an efficiency optimization scheme with triple-phase-shift (TPS) modulation using reinforcement learning (RL) is proposed in this article. More specifically, the Q-learning algorithm, as a typical algorithm of the RL, is applied to train an agent offline to obtain an optimized modulation strategy, and then the trained agent provides control decisions online in a real-time manner for the DAB dc-dc converter according to the current operating environment. The main objective is to obtain the optimal phase-shift angles for the DAB dc-dc converter, which can achieve the maximum power efficiency by reducing the power losses. Moreover, all possible operation modes of the TPS modulation are considered during the offline training process of the Q-learning algorithm. Thus, the cumbersome process for selecting the optimal operation mode in the conventional schemes can be circumvented successfully. Based on these merits, the proposed efficiency optimization scheme using the RL can realize the excellent performances for the whole load conditions and voltage conversion ratios. Finally, a 1.2-KW prototyped is built, and the simulation and the experimental results demonstrate that the power efficiency can be improved by using the optimization scheme based on the RL.