Exploration bonuses and dual control

Exploration bonuses and dual control
复制标题

DOI:
10.1007/bf00115298
复制
发表时间:
1996-10-01
期刊:
影响因子:
7.5
通讯作者:
Sejnowski, TJ
Sejnowski, TJ
中科院分区:
计算机科学3区
文献类型:
--
作者:
Dayan, P;Sejnowski, TJ

文献摘要

被引文献

相似文献

在自适应最优控制中寻找勘探和开发之间的贝叶斯平衡通常是一个棘手的问题。本文展示了如何基于一种形式的对偶控制的确定性等价近似(Cozzolino,Gonzalez-Zubieta&Miller,1965)来计算次最优估计。这将探索奖金在强化学习中的现有用途系统化和扩展(Sutton,1990)。该方法有两个组成部分:一个是对世界不确定性的统计模型,另一个是将其转化为探索性行为的方法。将该方法应用于具有可移动障碍物的二维迷宫,并与Sutton的DYNA系统进行了比较。
Finding the Bayesian balance between exploration and exploitation in adaptive optimal control is in general intractable. This paper shows how to compute suboptimal estimates based on a certainty equivalence approximation (Cozzolino, Gonzalez-Zubieta & Miller, 1965) arising from a form of dual control. This systematizes and extends existing uses of exploration bonuses in reinforcement learning (Sutton, 1990). The approach has two components: a statistical model of uncertainty in the world and a way of turning this into exploratory behavior. This general approach is applied to two-dimensional mazes with moveable barriers and its performance is compared with Sutton's DYNA system.