Reinforcement Learning of Optimal Input Excitation for Parameter Estimation With Application to Li-Ion Battery

Reinforcement Learning of Optimal Input Excitation for Parameter Estimation With Application to Li-Ion Battery
复制标题

最优输入激励强化学习在锂离子电池参数估计中的应用

DOI:
10.1109/tii.2023.3244342
复制
发表时间:
2023-11
影响因子:
12.3
通讯作者:
Rui Huang;J. Fogelquist;Xinfan Lin
Rui Huang;J. Fogelquist;Xinfan Lin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Rui Huang;J. Fogelquist;Xinfan Lin

文献摘要

被引文献

相似文献

系统状态和参数的估计和诊断在工业应用中是普遍存在的。估计通常使用输入和输出数据来执行,并且输入激励的质量对结果的准确性具有关键影响。因此,最优输入激励设计受到越来越多的研究关注。以前,输入设计被公式化为寻找激励序列的优化问题,该激励序列最大化与估计精度相关联的某个标准,例如,数据的信息内容。然而,这种方法也存在一些主要的缺点,包括易受不确定性(特别是目标参数)的影响和解的易处理性。在这篇文章中,强化学习(RL)框架被提出作为一种新的输入设计方法。我们设想的输入生成过程作为一个马尔可夫决策过程,并利用RL学习的最佳策略生成的输入激励。新方法通过策略的反馈机制提高了生成的输入序列的鲁棒性,通过强化学习的学习机制提高了输入序列的易处理性。该方法被应用于最佳激励设计,以估计关键的锂离子电池的电化学参数在模拟和实验。结果表明,新的RL为基础的框架显着优于传统的直接优化方法(一个数量级的信息水平更高)的目标参数估计的不确定性的存在下,并实现了小得多的估计误差相比,其他配置文件在实验中。所获得的RL策略可用于电池健康诊断和二次寿命电池的测试,以用于再利用应用。
Estimation and diagnosis of system states and parameters is ubiquitous in industrial applications. Estimation is often performed using input and output data, and the quality of input excitation has critical impact on the accuracy of the results. Therefore, optimal input excitation design has been receiving increasing research attention. Previously, input design is formulated as an optimization problem to find a sequence of excitation, which maximizes a certain criterion associated with estimation accuracy, e.g., the information content of the data. However, the practice suffers from several major drawbacks, including the susceptibility to uncertainty (especially that in target parameter) and tractability of solution. In this article, a reinforcement learning (RL) framework is proposed as a new approach for input design. We envision the input generation procedure as a Markov decision process, and leverage RL to learn an optimal policy for generating the input excitation. The new approach improves the robustness of the generated input sequence through the feedback mechanism of the policy, and tractability through the learning mechanism of RL. The methodology is applied to optimal excitation design for estimating critical lithium-ion battery electrochemical parameters in simulation and experiments. Results show that the new RL-based framework significantly outperforms the conventional direct optimization approach (with one order-of-magnitude higher information level) under the presence of uncertainty in the target parameter for estimation, and achieves substantially smaller estimation error compared with other profiles in experiment. The obtained RL policy could be used for battery health diagnostics and testing of second-life batteries for repurposing applications.