Active contextual policy search

Active contextual policy search
复制标题

主动上下文策略搜索

DOI:
--
复制
发表时间:
2014
影响因子:
6
通讯作者:
J. H. Metzen
J. H. Metzen
中科院分区:
计算机科学3区
文献类型:
--
作者:
Alexander Fabisch;J. H. Metzen

文献摘要

被引文献

相似文献

我们认为学习技能的问题是普遍适用的。学习这类技能的一种流行方法是上下文策略搜索,其中各个任务被表示为上下文向量。我们对代理能够主动选择它在学习过程中检查的任务的设置感兴趣。我们认为,有一种比同等频繁地选择每个任务更好的方法,因为一些任务可能在一开始更容易学习,而代理可以从这些任务中提取的知识可以转移到类似但更困难的任务中。我们提出的解决任务选择问题的方法将学习过程建模为一个带有客户内在奖励启发式的非平稳多臂土匪问题,从而使估计的学习进度最大化。该方法既不对底层上下文策略搜索算法做出任何假设,也不对策略表示做出任何假设。在模拟的三菱PA-10机械臂上,我们给出了一个人工基准问题和一个投球问题的实验结果,结果表明主动上下文选择可以显著提高技能的学习。
We consider the problem of learning skills that are versatilely applicable. One popular approach for learning such skills is contextual policy search in which the individual tasks are represented as context vectors. We are interested in settings in which the agent is able to actively select the tasks that it examines during the learning process. We argue that there is a better way than selecting each task equally often because some tasks might be easier to learn at the beginning and the knowledge that the agent can extract from these tasks can be transferred to similar but more difficult tasks. The methods that we propose for addressing the task-selection problem model the learning process as a nonstationary multi-armed bandit problem with custom intrinsic reward heuristics so that the estimated learning progress will be maximized. This approach does neither make any assumptions about the underlying contextual policy search algorithm nor about the policy representation. We present empirical results on an artificial benchmark problem and a ball throwing problem with a simulated Mitsubishi PA-10 robot arm which show that active context selection can improve the learning of skills considerably.