A least squares temporal difference actor–critic algorithm with applications to warehouse management

A least squares temporal difference actor–critic algorithm with applications to warehouse management
复制标题

最小二乘时间差演员批评算法及其在仓库管理中的应用

DOI:
10.1002/nav.21481
复制
发表时间:
2012
期刊:
Naval Research Logistics (NRL)
影响因子:
--
通讯作者:
I. Paschalidis
I. Paschalidis
中科院分区:
--
文献类型:
--
作者:
Reza Moazzez Estanjini;Keyong Li;I. Paschalidis

文献摘要

被引文献

相似文献

提出了一种求解马尔可夫决策问题的近似动态规划算法,并将其应用于仓库管理中的车辆调度问题。该算法是演员-评论家类型,并使用最小二乘时间差学习方法。它在系统的样本路径上运行,并在一个预先指定的类中优化策略,该类由一组简约的参数进行参数化。该方法适用于部分可观测的马尔可夫决策过程设置的状态变量的测量是潜在的损坏,和成本只能通过不完美的状态观测。我们表明,在合理的假设下,该算法收敛到一个局部最优参数集。我们还表明,不完美的成本观察不影响政策和算法最小化真正的期望成本。在仓库应用中,问题是如何调度配备传感器的叉车,以最大限度地降低涉及产品移动延迟和叉车维护的运营成本。我们考虑的情况下,标准的DP是计算棘手的。仿真结果证实了文章的理论主张,并表明,我们的算法比早期的演员-评论家算法收敛更顺利,同时大大优于在实践中使用的算法。© 2012 Wiley Periodicals,Inc.海军研究后勤,2012年
This article develops a new approximate dynamic programming (DP) algorithm for Markov decision problems and applies it to a vehicle dispatching problem arising in warehouse management. The algorithm is of the actor‐critic type and uses a least squares temporal difference learning method. It operates on a sample‐path of the system and optimizes the policy within a prespecified class parameterized by a parsimonious set of parameters. The method is applicable to a partially observable Markov decision process setting where the measurements of state variables are potentially corrupted, and the cost is only observed through the imperfect state observations. We show that under reasonable assumptions, the algorithm converges to a locally optimal parameter set. We also show that the imperfect cost observations do not affect the policy and the algorithm minimizes the true expected cost. In the warehouse application, the problem is to dispatch sensor‐equipped forklifts in order to minimize operating costs involving product movement delays and forklift maintenance. We consider instances where standard DP is computationally intractable. Simulation results confirm the theoretical claims of the article and show that our algorithm converges more smoothly than earlier actor–critic algorithms while substantially outperforming heuristics used in practice. © 2012 Wiley Periodicals, Inc. Naval Research Logistics, 2012