Importance Weighted Transfer of Samples in Reinforcement Learning

Importance Weighted Transfer of Samples in Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2018-05
期刊:
--
影响因子:
--
通讯作者:
Andrea Tirinzoni;Andrea Sessa;Matteo Pirotta;Marcello Restelli
Andrea Tirinzoni;Andrea Sessa;Matteo Pirotta;Marcello Restelli
中科院分区:
其他
文献类型:
--
作者:
Andrea Tirinzoni;Andrea Sessa;Matteo Pirotta;Marcello Restelli

文献摘要

被引文献

相似文献

我们考虑强化学习(RL)中经验样本(即元组)的转移,从一组源任务中收集经验样本,以改进给定目标任务的学习过程。大多数相关方法侧重于选择最相关的源样本来解决目标任务,但随后使用所有转移的样本而不再考虑任务模型之间的差异。在本文中,我们提出了一种基于模型的技术,可以自动估计每个源样本用于解决目标任务的相关性(重要性权重)。在所提出的方法中,所有样本都被转移并由批量强化学习算法使用来解决目标任务,但它们对学习过程的贡献与其重要性权重成正比。通过扩展监督学习文献中提供的重要性加权结果,我们对所提出的批量强化学习算法进行了有限样本分析。此外,我们根据经验将所提出的算法与最先进的方法进行比较,表明即使某些源任务与目标任务显着不同,它也能实现更好的学习性能并且对负迁移非常鲁棒。
We consider the transfer of experience samples (i.e., tuples ) in reinforcement learning (RL), collected from a set of source tasks to improve the learning process in a given target task. Most of the related approaches focus on selecting the most relevant source samples for solving the target task, but then all the transferred samples are used without considering anymore the discrepancies between the task models. In this paper, we propose a model-based technique that automatically estimates the relevance (importance weight) of each source sample for solving the target task. In the proposed approach, all the samples are transferred and used by a batch RL algorithm to solve the target task, but their contribution to the learning process is proportional to their importance weight. By extending the results for importance weighting provided in supervised learning literature, we develop a finite-sample analysis of the proposed batch RL algorithm. Furthermore, we empirically compare the proposed algorithm to state-of-the-art approaches, showing that it achieves better learning performance and is very robust to negative transfer, even when some source tasks are significantly different from the target task.