Real-Time Energy Management for Plug-in Hybrid Electric Vehicles via Incorporating Double-Delay Q-Learning and Model Prediction Control

Real-Time Energy Management for Plug-in Hybrid Electric Vehicles via Incorporating Double-Delay Q-Learning and Model Prediction Control
复制标题

DOI:
10.1109/access.2022.3229468
复制
发表时间:
2022
期刊:
影响因子:
3.9
通讯作者:
Shiquan Shen;Shun-Hua Gao;Yonggang Liu;Yuanjian Zhang;Jiangwei Shen;Zheng Chen;Z. Lei
Shiquan Shen;Shun-Hua Gao;Yonggang Liu;Yuanjian Zhang;Jiangwei Shen;Zheng Chen;Z. Lei
中科院分区:
计算机科学3区
文献类型:
--
作者:
Shiquan Shen;Shun-Hua Gao;Yonggang Liu;Yuanjian Zhang;Jiangwei Shen;Zheng Chen;Z. Lei

文献摘要

相似文献

插电式混合动力汽车(phev)在提高燃油经济性、减少有害气体排放和缓解里程焦虑方面具有巨大优势,已被证明是一种较好的交通解决方案。而设计一种有效的能量管理策略,在电池和发动机之间进行能量分配,是提高插电式混合动力汽车动力系统性能的关键。为此,提出了一种结合双延迟q -学习和模型预测控制(MPC)的实时能量管理策略。首先,将插电式混合动力汽车的能量管理问题转化为非线性最优控制问题,提出了基于卷积神经网络的速度预测器来预测插电式混合动力汽车的速度;然后,在预测车速的基础上,采用双延迟Q-Learning算法解决MPC模块中的后退地平线最优问题。仿真验证了所提策略的性能,结果表明,在MPC中加入双延迟Q-Learning可以有效提高能量管理对动态环境的适应性,同时达到与基于离线随机动态规划策略相似的油耗。此外,该策略的单步计算时间小于23毫秒,突出了其在线实现的巨大潜力。
Plug-in hybrid electric vehicles (PHEVs) have been validated as a preferable solution to transportation due to its great advantages in fuel economy promotion, harmful emission reduction and mileage anxiety mitigation. While, designing an effective energy management strategy to allocate the power between battery and engine is critical to improve the performance of powertrain in PHEVs. To this end, a real-time energy management strategy is proposed via incorporating double-delay Q-Learning and model predictive control (MPC). First, the energy management for PHEV is transformed into a nonlinear optimal control problem, and the vehicle speed predictor based on convolutional neural network is proposed to forecast vehicle speed in MPC. Then, based on the predicted vehicle speed, the double-delay Q-Learning algorithm is implemented to solve the receding horizon optimal problem in the MPC module. The simulation is conducted to verify the performance of the proposed strategy, and the results showcase that incorporating the double-delay Q-Learning into MPC can effectively improve the adaptability of energy management to dynamic environment, and meanwhile achieve a similar fuel consumption of the offline stochastic dynamic programming-based strategy. In addition, the single-step computation time of the proposed strategy is less than 23 milliseconds, highlighting its significant potential in online implementation.