Deep reinforcement learning based preventive maintenance policy for serial production lines

Deep reinforcement learning based preventive maintenance policy for serial production lines
复制标题

DOI:
10.1016/j.eswa.2020.113701
复制
发表时间:
2020-12-01
影响因子:
8.5
通讯作者:
Arinez, Jorge
Arinez, Jorge
中科院分区:
计算机科学1区
文献类型:
--
作者:
Huang, Jing;Chang, Qing;Arinez, Jorge

文献摘要

被引文献

相似文献

在制造业中,预防性维护(PM)是一种常见的做法,通过更换/维修老化的机器或零件来减少随机机器故障。由于具有中间缓冲器的串行生产线的复杂性和随机性,关于何时何地需要进行预防性维护的决定是重要的。为了提高串行生产线的成本效率,提出了一种基于深度强化学习的PM策略获取方法。在学习过程中采用了一种新的串行生产线建模方法。提出了一种基于系统生产损失评估的奖励函数。采用基于双深度Q网络的算法学习PM策略。通过仿真研究,证明了该学习算法在提供PM策略方面的有效性,从而提高了吞吐量并降低了成本。有趣的是,学习的政策被发现经常进行“组维护”和“机会主义维护”,虽然他们的概念和规则在学习过程中没有提供。这一发现进一步证明了本文中的问题形式、算法和奖励函数设置是有效的。(c)2020爱思唯尔有限公司保留所有权利。
In the manufacturing industry, the preventive maintenance (PM) is a common practice to reduce random machine failures by replacing/repairing the aged machines or parts. The decision on when and where the preventive maintenance needs to be carried out is nontrivial due to the complex and stochastic nature of a serial production line with intermediate buffers. In order to improve the cost efficiency of the serial production lines, a deep reinforcement learning based approach is proposed to obtain PM policy. A novel modeling method for the serial production line is adopted during the learning process. A reward function is proposed based on the system production loss evaluation. The algorithm based on the Double Deep Q-Network is applied to learn the PM policy. Using the simulation study, the learning algorithm is proved effective in delivering PM policy that leads to an increased throughput and reduced cost. Interestingly, the learned policy is found to frequently conduct "group maintenance" and "opportunistic maintenance", although their concepts and rules are not provided during the learning process. This finding further demonstrates that the problem formulation, the proposed algorithm and the reward function setting in this paper are effective. (c) 2020 Elsevier Ltd. All rights reserved.