Asynchronous framework with Reptile+ algorithm to meta learn partially observable Markov decision process

Asynchronous framework with Reptile+ algorithm to meta learn partially observable Markov decision process
复制标题

DOI:
10.1007/s10489-020-01748-7
复制
发表时间:
2020-07
影响因子:
5.3
通讯作者:
Dang Quang Nguyen;Ngo Anh Vien;Viet-Hung Dang;TaeChoong Chung
Dang Quang Nguyen;Ngo Anh Vien;Viet-Hung Dang;TaeChoong Chung
中科院分区:
计算机科学2区
文献类型:
--
作者:
Dang Quang Nguyen;Ngo Anh Vien;Viet-Hung Dang;TaeChoong Chung

文献摘要

被引文献

相似文献

元学习最近在各种深度强化学习(DRL)中受到了广泛的关注。在非元学习中,我们必须训练一个深度神经网络作为控制器,使用大量数据从头开始学习特定的控制任务。这种训练方式在处理不同的相关任务时显示出许多局限性。因此,控制域元学习成为相关任务迁移学习的有力工具。然而,众所周知,元学习需要大量的计算和训练时间。本文将提出一种新的DRL框架,称为HCGF-R2-DDPG(Hybrid CPU/GPU Framework for Reptile+ and Recurrent Deep Deterministic Policy Gradient)。HCGF-R2-DDPG将元学习集成到通用异步培训架构中。拟议的框架将允许利用CPU和GPU来提高Meta网络初始化的训练速度。我们将在各种部分可观察马尔可夫决策过程(POMDP)域上评估HCGF-R2-DDPG。
Meta-learning has recently received much attention in a wide variety of deep reinforcement learning (DRL). In non-meta-learning, we have to train a deep neural network as a controller to learn a specific control task from scratch using a large amount of data. This way of training has shown many limitations in handling different related tasks. Therefore, meta-learning on control domains becomes a powerful tool for transfer learning on related tasks. However, it is widely known that meta-learning requires massive computation and training time. This paper will propose a novel DRL framework, which is called HCGF-R2-DDPG (Hybrid CPU/GPU Framework for Reptile+ and Recurrent Deep Deterministic Policy Gradient). HCGF-R2-DDPG will integrate meta-learning into a general asynchronous training architecture. The proposed framework will allow utilising both CPU and GPU to boost the training speed for the meta network initialisation. We will evaluate HCGF-R2-DDPG on various Partially Observable Markov Decision Process (POMDP) domains.