Policy and Value Transfer in Lifelong Reinforcement Learning

Policy and Value Transfer in Lifelong Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2018-07
期刊:
--
影响因子:
--
通讯作者:
David Abel;Yuu Jinnai;Yue (Sophie) Guo;G. Konidaris;M. Littman
David Abel;Yuu Jinnai;Yue (Sophie) Guo;G. Konidaris;M. Littman
中科院分区:
其他
文献类型:
--
作者:
David Abel;Yuu Jinnai;Yue (Sophie) Guo;G. Konidaris;M. Littman

文献摘要

被引文献

相似文献

我们考虑如何最好地利用先前的经验来引导终身学习的问题,其中代理面临着从某些任务分布中提取的一系列任务实例。首先,我们确定了初始策略,该策略可以优化日益复杂的策略和任务分配类别的任务分配的预期性能。我们凭经验证明了每个策略类别的最优元素在各种简单任务分配中的相对性能。然后,我们考虑保留 PAC 保证的值函数初始化方法,同时最小化两种学习算法所需的学习,从而产生 MAX QI NIT,这是一种基于值函数的迁移的实用新方法。我们证明了 MAX QI NIT 在简单的终身强化学习实验中表现良好。
We consider the problem of how best to use prior experience to bootstrap lifelong learning, where an agent faces a series of task instances drawn from some task distribution. First, we identify the initial policy that optimizes expected performance over the distribution of tasks for increasingly complex classes of policy and task distributions. We empirically demonstrate the relative performance of each policy class’ optimal element in a variety of simple task distributions. We then consider value-function initialization methods that preserve PAC guarantees while simultaneously minimizing the learning required in two learning algorithms, yielding M AX QI NIT , a practical new method for value-function-based transfer. We show that M AX QI NIT performs well in simple lifelong RL experiments.