Intelligent Trainer for Dyna-Style Model-Based Deep Reinforcement Learning

Intelligent Trainer for Dyna-Style Model-Based Deep Reinforcement Learning
复制标题

DOI:
10.1109/tnnls.2020.3008249
复制
发表时间:
2020-08
影响因子:
10.4
通讯作者:
Linsen Dong;Yuanlong Li;Xiaoxia Zhou;Yonggang Wen;K. Guan
Linsen Dong;Yuanlong Li;Xiaoxia Zhou;Yonggang Wen;K. Guan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Linsen Dong;Yuanlong Li;Xiaoxia Zhou;Yonggang Wen;K. Guan

文献摘要

被引文献

相似文献

基于模型的强化学习(MBRL)是一种很有前途的替代解决方案,通过利用系统动力学模型来生成用于策略训练的合成数据,以解决规范强化学习中的高采样成本挑战。然而,MBRL框架固有地受到联合优化控制策略、学习系统动力学以及从由复杂超参数控制的两个源采样数据的复杂过程的限制。因此,训练过程涉及压倒性的手动调整,并且成本高昂。在这项研究中,我们提出了一个“强化强化”(RoR)的架构分解成两个解耦的RL层的卷积任务。内层是规范的MBRL训练过程,它被公式化为马尔可夫决策过程,称为训练过程环境(TPE)。外层充当RL代理,称为智能训练器,以学习内部TPE的最佳超参数配置。这种分解方法为实现不同的培训师设计提供了急需的灵活性,称为“培训培训师”。在我们的研究中,我们提出并优化了两种替代训练器设计:1)单头训练器和2)多头训练器。我们提出的RoR框架在OpenAI健身房中针对五个任务进行了评估。与其他三种基线方法相比,我们提出的智能训练器方法在自动调整能力方面具有竞争力的性能,在事先不知道最佳参数配置的情况下,预期采样成本节省高达56%。所提出的训练器框架可以很容易地扩展到需要昂贵的超参数调整的任务。
Model-based reinforcement learning (MBRL) has been proposed as a promising alternative solution to tackle the high sampling cost challenge in the canonical RL, by leveraging a system dynamics model to generate synthetic data for policy training purpose. The MBRL framework, nevertheless, is inherently limited by the convoluted process of jointly optimizing control policy, learning system dynamics, and sampling data from two sources controlled by complicated hyperparameters. As such, the training process involves overwhelmingly manual tuning and is prohibitively costly. In this research, we propose a “reinforcement on reinforcement” (RoR) architecture to decompose the convoluted tasks into two decoupled layers of RL. The inner layer is the canonical MBRL training process which is formulated as a Markov decision process, called training process environment (TPE). The outer layer serves as an RL agent, called intelligent trainer, to learn an optimal hyperparameter configuration for the inner TPE. This decomposition approach provides much-needed flexibility to implement different trainer designs, referred to “train the trainer.” In our research, we propose and optimize two alternative trainer designs: 1) an unihead trainer and 2) a multihead trainer. Our proposed RoR framework is evaluated for five tasks in the OpenAI gym. Compared with three other baseline methods, our proposed intelligent trainer methods have a competitive performance in autotuning capability, with up to 56% expected sampling cost saving without knowing the best parameter configurations in advance. The proposed trainer framework can be easily extended to tasks that require costly hyperparameter tuning.