Policy Caches with Successor Features

Policy Caches with Successor Features
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Mark W. Nemecek;R. Parr
Mark W. Nemecek;R. Parr
中科院分区:
其他
文献类型:
--
作者:
Mark W. Nemecek;R. Parr

文献摘要

被引文献

相似文献

强化学习中的迁移是基于这样一种想法,即可以使用在一个任务中学到的东西来改善另一个任务中的学习过程。对于共享过渡动态但奖励函数不同的任务之间的转移,后继特征已被证明是一种有用的表示,可以有效地计算新任务中先前学习的策略的动作值函数。这些功能在新任务中引入策略,因此代理可能不需要为它遇到的每个新任务学习新策略,特别是如果允许在这些任务中有一定量的次优性。我们提出了新的界限,在一个新的任务中的最佳策略的性能,以及一种方法来使用这些界限来决定,当提出一个新的任务,是否使用缓存的政策或学习一个新的政策。
Transfer in reinforcement learning is based on the idea that it is possible to use what is learned in one task to improve the learning process in another task. For transfer between tasks which share transition dynamics but differ in reward function, successor features have been shown to be a useful representation which allows for efficient computation of action-value functions for previously-learned policies in new tasks. These functions induce policies in the new tasks, so an agent may not need to learn a new policy for each new task it encounters, especially if it is allowed some amount of subop-timality in those tasks. We present new bounds for the performance of optimal policies in a new task, as well as an approach to use these bounds to decide, when presented with a new task, whether to use cached policies or learn a new policy.