Robustly Learning Composable Options in Deep Reinforcement Learning

Robustly Learning Composable Options in Deep Reinforcement Learning
复制标题

DOI:
10.24963/ijcai.2021/298
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
Akhil Bagaria;J. Senthil;Matthew Slivinski;G. Konidaris
Akhil Bagaria;J. Senthil;Matthew Slivinski;G. Konidaris
中科院分区:
其他
文献类型:
--
作者:
Akhil Bagaria;J. Senthil;Matthew Slivinski;G. Konidaris

文献摘要

相似文献

当高级技能能够可靠地按顺序执行时,分层强化学习(HRL)只对长期问题有效。不幸的是,学习可靠的组合技能是困难的,因为每种技能的所有组成部分在学习过程中都在不断变化。我们提出了三种方法来提高学习技能的可组合性:使用悲观和乐观分类器的组合来表示技能起始区域;学习对非平稳子目标区域具有健壮性的可重定向策略;以及使用基于模型的RL来学习健壮的选项策略。我们在四个稀疏奖励迷宫导航任务上测试了这些改进,这些任务涉及一个模拟的四足机器人。每种方法都相继提高了基线技能发现方法的健壮性,大大优于最先进的平面和分层方法。
Hierarchical reinforcement learning (HRL) is only effective for long-horizon problems when high-level skills can be reliably sequentially executed. Unfortunately, learning reliably composable skills is difficult, because all the components of every skill are constantly changing during learning. We propose three methods for improving the composability of learned skills: representing skill initiation regions using a combination of pessimistic and optimistic classifiers; learning re-targetable policies that are robust to non-stationary subgoal regions; and learning robust option policies using model-based RL. We test these improvements on four sparse-reward maze navigation tasks involving a simulated quadrupedal robot. Each method successively improves the robustness of a baseline skill discovery method, substantially outperforming state-of-the-art flat and hierarchical methods.