Scaling Up Approximate Value Iteration with Options: Better Policies with Fewer Iterations

Scaling Up Approximate Value Iteration with Options: Better Policies with Fewer Iterations
复制标题

通过选项扩大近似值迭代:用更少的迭代获得更好的策略

DOI:
--
复制
发表时间:
2014
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
Shie Mannor
Shie Mannor
中科院分区:
--
文献类型:
--
作者:
Timothy A. Mann;Shie Mannor

文献摘要

被引文献

相似文献

我们展示了选项(一类包含原始动作和时间扩展动作的控制结构)如何在具有连续状态空间的马尔可夫决策过程(MDPs)的规划中发挥重要作用。对带有选项的近似值迭代的收敛速度进行分析后发现,对于悲观的初始值函数估计,即使时间扩展动作是次优的且在状态空间中稀疏分布,与仅使用原始动作进行规划相比,选项也能够加快收敛速度。我们在最优替换任务和复杂库存管理任务中的实验结果证明了选项在实际中加快收敛的潜力。我们表明选项能促使更快地收敛到最优值函数,这意味着用更少的迭代次数得出更好的策略。
We show how options, a class of control structures encompassing primitive and temporally extended actions, can play a valuable role in planning in MDPs with continuous state-spaces. Analyzing the convergence rate of Approximate Value Iteration with options reveals that for pessimistic initial value function estimates, options can speed up convergence compared to planning with only primitive actions even when the temporally extended actions are suboptimal and sparsely scattered throughout the state-space. Our experimental results in an optimal replacement task and a complex inventory management task demonstrate the potential for options to speed up convergence in practice. We show that options induce faster convergence to the optimal value function, which implies deriving better policies with fewer iterations.