All You Need Is Supervised Learning: From Imitation Learning to Meta-RL With Upside Down RL

All You Need Is Supervised Learning: From Imitation Learning to Meta-RL With Upside Down RL
复制标题

您所需要的只是监督学习:从模仿学习到颠倒强化学习的元强化学习

DOI:
--
复制
发表时间:
2022
期刊:
arXiv.org
影响因子:
--
通讯作者:
R. Srivastava
R. Srivastava
中科院分区:
--
文献类型:
--
作者:
Kai Arulkumaran;Dylan R. Ashley;J. Schmidhuber;R. Srivastava

文献摘要

参考文献

被引文献

相似文献

倒置强化学习(UDRL)通过将回报作为输入并预测动作,颠倒了RL中目标函数中回报的传统用法。UDRL纯粹基于监督学习,并绕过了RL中的一些突出问题:引导、非政策修正和贴现因素。虽然之前对UDRL的研究在传统的在线RL设置中展示了它,但在这里我们展示了这个单一算法也可以在模仿学习和离线RL设置中工作,可以扩展到目标条件RL设置,甚至元RL设置。使用通用代理体系结构,单个UDRL代理可以跨所有范例学习。
Upside down reinforcement learning (UDRL) flips the conventional use of the return in the objective function in RL upside down, by taking returns as input and predicting actions. UDRL is based purely on supervised learning, and bypasses some prominent issues in RL: bootstrapping, off-policy corrections, and discount factors. While previous work with UDRL demonstrated it in a traditional online RL setting, here we show that this single algorithm can also work in the imitation learning and offline RL settings, be extended to the goal-conditioned RL setting, and even the meta-RL setting. With a general agent architecture, a single UDRL agent can learn across all paradigms.
DOI: --
发表时间: 2021-06
期刊: --
影响因子: --
作者:
Lili Chen;Kevin Lu;A. Rajeswaran;Kimin Lee;Aditya Grover;M. Laskin;P. Abbeel;A. Srinivas;
通讯作者: Lili Chen;Kevin Lu;A. Rajeswaran;Kimin Lee;Aditya Grover;M. Laskin;P. Abbeel;A. Srinivas;