Hierarchical and Interpretable Skill Acquisition in Multi-task Reinforcement Learning

Hierarchical and Interpretable Skill Acquisition in Multi-task Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2017-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Tianmin Shu;Caiming Xiong;R. Socher
Tianmin Shu;Caiming Xiong;R. Socher
中科院分区:
其他
文献类型:
--
作者:
Tianmin Shu;Caiming Xiong;R. Socher

文献摘要

被引文献

相似文献

针对需要多种不同技能的复杂任务的学习策略是强化学习(RL)中的一个主要挑战。这也是其在现实世界场景中部署的要求。本文提出了一种新的高效多任务强化学习框架。我们的框架训练代理采用分层策略,决定何时使用以前学习的策略以及何时学习新技能。这使代理能够在不同的培训阶段不断获得新技能。每个学习任务对应于人类语言描述。由于代理只能通过这些描述访问以前学习的技能,因此代理始终可以为其选择提供人类可解释的描述。为了帮助代理学习复杂的时间依赖性所需的分层政策,我们提供了一个随机的时间语法,调制时依赖于以前学到的技能,何时执行新的技能。我们在Minecraft游戏上验证了我们的方法,这些游戏旨在明确测试重用以前学习的技能的能力,同时学习新技能。
Learning policies for complex tasks that require multiple different skills is a major challenge in reinforcement learning (RL). It is also a requirement for its deployment in real-world scenarios. This paper proposes a novel framework for efficient multi-task reinforcement learning. Our framework trains agents to employ hierarchical policies that decide when to use a previously learned policy and when to learn a new skill. This enables agents to continually acquire new skills during different stages of training. Each learned task corresponds to a human language description. Because agents can only access previously learned skills through these descriptions, the agent can always provide a human-interpretable description of its choices. In order to help the agent learn the complex temporal dependencies necessary for the hierarchical policy, we provide it with a stochastic temporal grammar that modulates when to rely on previously learned skills and when to execute new skills. We validate our approach on Minecraft games designed to explicitly test the ability to reuse previously learned skills while simultaneously learning new skills.