Policy and Value Transfer in Lifelong Reinforcement Learning
Policy and Value Transfer in Lifelong Reinforcement Learning
复制标题
DOI:
--
复制
发表时间:
2018-07
期刊:
影响因子:
--
通讯作者:
David Abel;Yuu Jinnai;Yue (Sophie) Guo;G. Konidaris;M. Littman
中科院分区:
文献类型:
--
作者:
David Abel;Yuu Jinnai;Yue (Sophie) Guo;G. Konidaris;M. Littman
We consider the problem of how best to use prior experience to bootstrap lifelong learning, where an agent faces a series of task instances drawn from some task distribution. First, we identify the initial policy that optimizes expected performance over the distribution of tasks for increasingly complex classes of policy and task distributions. We empirically demonstrate the relative performance of each policy class’ optimal element in a variety of simple task distributions. We then consider value-function initialization methods that preserve PAC guarantees while simultaneously minimizing the learning required in two learning algorithms, yielding M AX QI NIT , a practical new method for value-function-based transfer. We show that M AX QI NIT performs well in simple lifelong RL experiments.