Reinforcement learning with Gaussian processes for condition-based maintenance

Reinforcement learning with Gaussian processes for condition-based maintenance
复制标题

DOI:
10.1016/j.cie.2021.107321
复制
发表时间:
2021-04
期刊:
Comput. Ind. Eng.
影响因子:
--
通讯作者:
Shenglin Peng;Q. Feng
Shenglin Peng;Q. Feng
中科院分区:
其他
文献类型:
--
作者:
Shenglin Peng;Q. Feng

文献摘要

被引文献

相似文献

基于状态的维护策略可有效提高复杂工程系统的可靠性和安全性,这些工程系统表现出不确定的退化现象。当底层过程具有马尔可夫属性时,此类顺序决策问题通常被建模为马尔可夫决策过程 (MDP)。最近,强化学习(RL)在解决具有大状态空间的 MDP 问题方面变得越来越有效。在本文中,我们将基于状态的维护问题建模为离散时间连续状态 MDP,而不离散系统的恶化条件。高斯过程回归用作函数近似来对强化学习中的状态转换和状态值函数进行建模。然后开发 RL 算法,通过分别迭代状态动作价值函数和状态价值函数来最小化长期平均成本(而不是常用的折扣奖励)。我们通过仿真实验验证了所提出算法的能力,并在电池维护决策问题的案例研究中证明了其优势。所提出的算法通过实现较低的长期平均成本而优于离散 MDP 方法。
Condition-based maintenance strategies are effective in enhancing reliability and safety for complex engineering systems that exhibit degradation phenomena with uncertainty. Such sequential decision-making problems are often modeled as Markov decision processes (MDPs) when the underlying process has a Markov property. Recently, reinforcement learning (RL) becomes increasingly efficient to address MDP problems with large state spaces. In this paper, we model the condition-based maintenance problem as a discrete-time continuous-state MDP without discretizing the deterioration condition of the system. The Gaussian process regression is used as function approximation to model the state transition and the value functions of states in reinforcement learning. A RL algorithm is then developed to minimize the long-run average cost (instead of the commonly-used discounted reward) with iterations on the state-action value function and the state value function, respectively. We verify the capability of the proposed algorithm by simulation experiments and demonstrate its advantages in a case study on a battery maintenance decision-making problem. The proposed algorithm outperforms the discrete MDP approach by achieving lower long-run average costs.