Exploration bonuses and dual control
Exploration bonuses and dual control
复制标题
DOI:
10.1007/bf00115298
复制
发表时间:
1996-10-01
期刊:
影响因子:
7.5
通讯作者:
Sejnowski, TJ
中科院分区:
文献类型:
--
作者:
Dayan, P;Sejnowski, TJ
Finding the Bayesian balance between exploration and exploitation in adaptive optimal control is in general intractable. This paper shows how to compute suboptimal estimates based on a certainty equivalence approximation (Cozzolino, Gonzalez-Zubieta & Miller, 1965) arising from a form of dual control. This systematizes and extends existing uses of exploration bonuses in reinforcement learning (Sutton, 1990). The approach has two components: a statistical model of uncertainty in the world and a way of turning this into exploratory behavior. This general approach is applied to two-dimensional mazes with moveable barriers and its performance is compared with Sutton's DYNA system.