An MDP Model-Based Reinforcement Learning Approach for Production Station Ramp-Up Optimization: Q-Learning Analysis

An MDP Model-Based Reinforcement Learning Approach for Production Station Ramp-Up Optimization: Q-Learning Analysis
复制标题

DOI:
10.1109/tsmc.2013.2294155
复制
发表时间:
2014-01
期刊:
IEEE Transactions on Systems, Man, and Cybernetics: Systems
影响因子:
--
通讯作者:
Stefanos Doltsinis;P. Ferreira;N. Lohse
Stefanos Doltsinis;P. Ferreira;N. Lohse
中科院分区:
其他
文献类型:
--
作者:
Stefanos Doltsinis;P. Ferreira;N. Lohse

文献摘要

被引文献

相似文献

提升是引入新的或适应的制造系统的一个重要瓶颈。升级系统所需的努力和时间在很大程度上取决于人类决策过程的有效性,以选择最有希望的行动序列,以将系统改进到所需的性能水平。虽然现有的工作已经确定了影响逐步提高有效性的重要因素,但在支持这一过程中的决策方面所做的工作很少。本文将提升作为一种顺序的调整和调整过程,旨在使制造系统在尽可能短的时间内达到期望的性能。生产工位和机器是制造系统中的关键资源。它们通常在功能上是分离的,可以首先作为独立的提升问题来处理。因此,本文致力于建立一个马尔可夫决策过程(MDP)模型来形式化地描述生产站点的投产情况,并使其能够进行形式化分析。其目的是捕捉操作员对站点的适应或调整与站点的响应之间的因果关系,以提高过程的有效性。强化学习已经被认为是一种很有前途的方法,可以从经验中学习,发现更成功的决策策略。尤其是批处理学习可以在很少的数据下表现得很好。本文研究了Q批学习算法与MDP模型相结合的斜升过程的应用。该方法已应用于一个高度自动化的生产站,在那里进行了几个提升过程。分析了Q-学习算法的收敛性能随其参数的变化情况。最后,将学习到的策略进行了应用,并与以前的提升案例进行了比较。
Ramp-up is a significant bottleneck for the introduction of new or adapted manufacturing systems. The effort and time required to ramp-up a system is largely dependent on the effectiveness of the human decision making process to select the most promising sequence of actions to improve the system to the required level of performance. Although existing work has identified significant factors influencing the effectiveness of ramp-up, little has been done to support the decision making during the process. This paper approaches ramp-up as a sequential adjustment and tuning process that aims to get a manufacturing system to a desirable performance in the fastest possible time. Production stations and machines are the key resources in a manufacturing system. They are often functionally decoupled and can be treated in the first instance as independent ramp-up problems. Hence, this paper focuses on developing a Markov decision process (MDP) model to formalize ramp-up of production stations and enable their formal analysis. The aim is to capture the cause-and-effect relationships between an operator's adaptation or adjustment of a station and the station's response to improve the effectiveness of the process. Reinforcement learning has been identified as a promising approach to learn from ramp-up experience and discover more successful decision-making policies. Batch learning in particular can perform well with little data. This paper investigates the application of a Q-batch learning algorithm combined with an MDP model of the ramp-up process. The approach has been applied to a highly automated production station where several ramp-up processes are carried out. The convergence of the Q-learning algorithm has been analyzed along with the variation of its parameters. Finally, the learned policy has been applied and compared against previous ramp-up cases.