TRANSFER OF LEARNING BY COMPOSING SOLUTIONS OF ELEMENTAL SEQUENTIAL TASKS

TRANSFER OF LEARNING BY COMPOSING SOLUTIONS OF ELEMENTAL SEQUENTIAL TASKS
复制标题

DOI:
10.1007/bf00992700
复制
发表时间:
1992-05-01
期刊:
影响因子:
7.5
通讯作者:
SINGH, SP
SINGH, SP
中科院分区:
计算机科学3区
文献类型:
--
作者:
SINGH, SP

文献摘要

被引文献

相似文献

虽然构建在复杂环境中运行的复杂学习代理需要学习执行多个任务,但大多数强化学习的应用都集中在单个任务上。在本文中,我考虑一类顺序决策任务(SDTs),称为复合顺序决策任务,形成时间上连接的一些元素顺序决策任务。元素SDTs不能分解成更简单的SDTs。我认为一个学习代理,必须学会解决一组元素和复合SDTs。我假设复合任务的结构是未知的学习代理。将强化学习直接应用于多个任务需要分别学习任务,这可能会浪费计算资源,包括内存和时间。我提出了一个新的学习算法和一个模块化的体系结构,学习复合SDTs的分解,并通过共享跨多个复合SDTs的元素SDTs的解决方案,实现学习的转移。复合SDT的解是通过对其组成元素SDT的解进行计算上廉价的修改来构造的。我提供了学习算法的一个方面的证明。
Although building sophisticated learning agents that operate in complex environments will require learning to perform multiple tasks, most applications of reinforcement learning have focused on single tasks. In this paper I consider a class of sequential decision tasks (SDTs), called composite sequential decision tasks, formed by temporally concatenating a number of elemental sequential decision tasks. Elemental SDTs cannot be decomposed into simpler SDTs. I consider a learning agent that has to learn to solve a set of elemental and composite SDTs. I assume that the structure of the composite tasks is unknown to the learning agent. The straightforward application of reinforcement learning to multiple tasks requires learning the tasks separately, which can waste computational resources, both memory and time. I present a new learning algorithm and a modular architecture that learns the decomposition of composite SDTs, and achieves transfer of learning by sharing the solutions of elemental SDTs across multiple composite SDTs. The solution of a composite SDT is constructed by computationally inexpensive modifications of the solutions of its constituent elemental SDTs. I provide a proof of one aspect of the learning algorithm.