Active Contextual Entropy Search

Active Contextual Entropy Search
复制标题

主动上下文熵搜索

DOI:
--
复制
发表时间:
2015
期刊:
arXiv.org
影响因子:
--
通讯作者:
J. H. Metzen
J. H. Metzen
中科院分区:
--
文献类型:
--
作者:
J. H. Metzen

文献摘要

被引文献

相似文献

上下文策略搜索允许根据不同的情况调整机器人的运动原语。例如,运动原语可能适应不同的地形倾斜度或期望的行走速度。这种适应通常可以通过修改少量超参数来实现。然而,当在真正的机器人系统上进行学习时,通常仅限于少量的试验。贝叶斯优化最近被提出作为上下文策略搜索的样本效率手段,非常适合这些条件。在这项工作中,我们扩展了熵搜索,这是贝叶斯优化的一种变体,这样它就可以用于主动上下文策略搜索,其中智能体在训练期间选择它希望学习最多的任务。模拟的经验结果表明,这可以通过较少的试验来学习成功的行为。
Contextual policy search allows adapting robotic movement primitives to different situations. For instance, a locomotion primitive might be adapted to different terrain inclinations or desired walking speeds. Such an adaptation is often achievable by modifying a small number of hyperparameters. However, learning, when performed on real robotic systems, is typically restricted to a small number of trials. Bayesian optimization has recently been proposed as a sample-efficient means for contextual policy search that is well suited under these conditions. In this work, we extend entropy search, a variant of Bayesian optimization, such that it can be used for active contextual policy search where the agent selects those tasks during training in which it expects to learn the most. Empirical results in simulation suggest that this allows learning successful behavior with less trials.