Portable Option Discovery for Automated Learning Transfer in Object-Oriented Markov Decision Processes

Portable Option Discovery for Automated Learning Transfer in Object-Oriented Markov Decision Processes
复制标题

面向对象马尔可夫决策过程中自动学习迁移的便携式选项发现

DOI:
--
复制
发表时间:
2015
期刊:
International Joint Conference on Artificial Intelligence
影响因子:
--
通讯作者:
J. MacGlashan
J. MacGlashan
中科院分区:
--
文献类型:
--
作者:
Nicholay Topin;Nicholas Haltmeyer;S. Squire;J. Winder;Marie desJardins;J. MacGlashan

文献摘要

被引文献

相似文献

我们引入了一个新颖的框架,用于在以面向对象马尔可夫决策过程(OO - MDPs)[迪乌克等人,2008]表示的复杂领域中进行选项发现和学习迁移。我们的框架,可移植选项发现(POD),扩展了现有的选项发现方法,并通过提供一种无监督方法来寻找具有不同状态空间的面向对象领域之间的映射,从而实现跨相关但不同领域的迁移。该框架还包括用于提高映射过程效率的启发式方法。我们展示了在两个应用领域中将POD应用于皮克特和巴托[2002]的策略块以及麦克格拉尚[2013]的基于选项的策略迁移的结果。我们表明,我们的方法能够有效地发现选项,在不同领域之间迁移选项,并以较低的计算开销提高学习性能。
We introduce a novel framework for option discovery and learning transfer in complex domains that are represented as object-oriented Markov decision processes (OO-MDPs) [Diuk et al., 2008]. Our framework, Portable Option Discovery (POD), extends existing option discovery methods, and enables transfer across related but different domains by providing an unsupervised method for finding a mapping between object-oriented domains with different state spaces. The framework also includes heuristic approaches for increasing the efficiency of the mapping process. We present the results of applying POD to Pickett and Barto's [2002] Policy-Blocks and MacGlashan's [2013] Option-Based Policy Transfer in two application domains. We show that our approach can discover options effectively, transfer options among different domains, and improve learning performance with low computational overhead.