Provable Benefits of Representational Transfer in Reinforcement Learning

Provable Benefits of Representational Transfer in Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2205.14571
复制
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Alekh Agarwal;Yuda Song;Wen Sun;Kaiwen Wang;Mengdi Wang;Xuezhou Zhang
Alekh Agarwal;Yuda Song;Wen Sun;Kaiwen Wang;Mengdi Wang;Xuezhou Zhang
中科院分区:
其他
文献类型:
--
作者:
Alekh Agarwal;Yuda Song;Wen Sun;Kaiwen Wang;Mengdi Wang;Xuezhou Zhang

文献摘要

相似文献

我们研究了RL中的表征转移问题,其中代理首先在多个源任务中进行预训练以发现共享的表征,该表征随后用于在目标任务中学习好的策略。我们提出了一个新的概念,源和目标任务之间的任务相关性,并在此假设下发展了一种新的表征转移的方法。具体地说,我们表明,由于生成访问源任务,我们可以发现一个表示,使用随后的线性RL技术快速收敛到一个接近最优的政策,在目标任务。样本复杂度接近于目标任务中的地面真值特征,并且与源任务中的先前表示学习结果相当。我们补充我们的积极结果与下界没有生成访问,并验证我们的研究结果与经验评估丰富的观察MDPs,需要深入的探索。在我们的实验中,我们观察到通过预训练在目标中学习的速度加快,并且还验证了在源任务中生成访问的必要性。
We study the problem of representational transfer in RL, where an agent first pretrains in a number of source tasks to discover a shared representation, which is subsequently used to learn a good policy in a \emph{target task}. We propose a new notion of task relatedness between source and target tasks, and develop a novel approach for representational transfer under this assumption. Concretely, we show that given generative access to source tasks, we can discover a representation, using which subsequent linear RL techniques quickly converge to a near-optimal policy in the target task. The sample complexity is close to knowing the ground truth features in the target task, and comparable to prior representation learning results in the source tasks. We complement our positive results with lower bounds without generative access, and validate our findings with empirical evaluation on rich observation MDPs that require deep exploration. In our experiments, we observe a speed up in learning in the target by pre-training, and also validate the need for generative access in source tasks.