Reinforcement Learning for Mixed Open-loop and Closed-loop Control

Reinforcement Learning for Mixed Open-loop and Closed-loop Control
复制标题

混合开环和闭环控制的强化学习

DOI:
--
复制
发表时间:
1996
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
S. Zilberstein
S. Zilberstein
中科院分区:
--
文献类型:
--
作者:
E. Hansen;A. Barto;S. Zilberstein

文献摘要

被引文献

相似文献

闭环控制依赖于通常被认为是免费的感觉反馈。但如果感知产生了成本,在开环模式下采取一系列行动可能是划算的。我们描述了一种强化学习算法,当感知产生代价时,该算法学习结合开环和闭环控制。虽然我们假设传感器可靠,但开环控制的使用意味着有时必须在受控系统的当前状态不确定时采取行动。这是强化学习中隐藏状态问题的一个特例,为了解决这个问题,我们的算法依赖于短期记忆。这篇论文的主要结果是一条规则,通过修剪信息估计价值大于其成本的记忆状态,大大限制了对可能记忆状态的探索。我们证明了该规则允许收敛到最优策略。
Closed-loop control relies on sensory feedback that is usually assumed to be free. But if sensing incurs a cost, it may be cost-effective to take sequences of actions in open-loop mode. We describe a reinforcement learning algorithm that learns to combine open-loop and closed-loop control when sensing incurs a cost. Although we assume reliable sensors, use of open-loop control means that actions must sometimes be taken when the current state of the controlled system is uncertain. This is a special case of the hidden-state problem in reinforcement learning, and to cope, our algorithm relies on short-term memory. The main result of the paper is a rule that significantly limits exploration of possible memory states by pruning memory states for which the estimated value of information is greater than its cost. We prove that this rule allows convergence to an optimal policy.