Policy Search with High-Dimensional Context Variables

Policy Search with High-Dimensional Context Variables
复制标题

DOI:
10.1609/aaai.v31i1.10911
复制
发表时间:
2016-11
期刊:
--
影响因子:
--
通讯作者:
Voot Tangkaratt;H. V. Hoof;Simone Parisi;G. Neumann;Jan Peters;Masashi Sugiyama
Voot Tangkaratt;H. V. Hoof;Simone Parisi;G. Neumann;Jan Peters;Masashi Sugiyama
中科院分区:
其他
文献类型:
--
作者:
Voot Tangkaratt;H. V. Hoof;Simone Parisi;G. Neumann;Jan Peters;Masashi Sugiyama

文献摘要

相似文献

直接上下文策略搜索方法学习改进策略参数,同时将这些参数推广到不同的上下文或任务变量。然而,从高维上下文变量(如相机图像)中学习仍然是许多现实任务中的一个突出问题。一个天真的应用程序的无监督降维方法的上下文变量,如主成分分析,是不够的,任务相关的输入可能会被忽略。在基于模型的相对熵随机搜索框架下,提出了一种结合降维的上下文策略搜索方法。我们学习一个奖励模型,该模型在策略参数和上下文变量中都是局部二次的。此外,我们通过核范数正则化对上下文变量进行监督线性降维。实验结果表明,该方法优于朴素的降维,通过主成分分析和最先进的上下文策略搜索方法。
Direct contextual policy search methods learn to improve policy parameters and simultaneously generalize these parameters to different context or task variables. However, learning from high-dimensional context variables, such as camera images, is still a prominent problem in many real-world tasks. A naive application of unsupervised dimensionality reduction methods to the context variables, such as principal component analysis, is insufficient as task-relevant input may be ignored. In this paper, we propose a contextual policy search method in the model-based relative entropy stochastic search framework with integrated dimensionality reduction. We learn a model of the reward that is locally quadratic in both the policy parameters and the context variables. Furthermore, we perform supervised linear dimensionality reduction on the context variables by nuclear norm regularization. The experimental results show that the proposed method outperforms naive dimensionality reduction via principal component analysis and a state-of-the-art contextual policy search method.