Control Regularization for Reduced Variance Reinforcement Learning

Control Regularization for Reduced Variance Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
--
影响因子:
--
通讯作者:
Richard Cheng;Abhinav Verma;G. Orosz;Swarat Chaudhuri;Yisong Yue;J. Burdick
Richard Cheng;Abhinav Verma;G. Orosz;Swarat Chaudhuri;Yisong Yue;J. Burdick
中科院分区:
其他
文献类型:
--
作者:
Richard Cheng;Abhinav Verma;G. Orosz;Swarat Chaudhuri;Yisong Yue;J. Burdick

文献摘要

被引文献

相似文献

处理高方差是无模型增强学习(RL)的重大挑战。现有方法是不可靠的,从运行到使用不同的初始化/种子的运行表现出很大的差异。专注于在连续控制中产生的问题,我们提出了一种功能正则化方法来增强无模型RL。特别是,我们将深层政策的行为正规化,以便与之前的策略相似,即我们在功能空间中规范。我们表明,功能正则化会产生偏见变化的权衡,并提出一种自适应调整策略来优化这种权衡。当先验的政策具有控制理论稳定性的保证时,我们进一步表明,这种正则化大约可以保留整个学习中的稳定性保证。我们在一系列设置上从经验上验证了我们的方法,并且比单独的深度RL证明了差异,保证的动态稳定性和更有效的学习能力显着降低。
Dealing with high variance is a significant challenge in model-free reinforcement learning (RL). Existing methods are unreliable, exhibiting high variance in performance from run to run using different initializations/seeds. Focusing on problems arising in continuous control, we propose a functional regularization approach to augmenting model-free RL. In particular, we regularize the behavior of the deep policy to be similar to a policy prior, i.e., we regularize in function space. We show that functional regularization yields a bias-variance trade-off, and propose an adaptive tuning strategy to optimize this trade-off. When the policy prior has control-theoretic stability guarantees, we further show that this regularization approximately preserves those stability guarantees throughout learning. We validate our approach empirically on a range of settings, and demonstrate significantly reduced variance, guaranteed dynamic stability, and more efficient learning than deep RL alone.