Learning Multiple Goal Behavior via Task Decomposition and Dynamic Policy Merging

Learning Multiple Goal Behavior via Task Decomposition and Dynamic Policy Merging
复制标题

通过任务分解和动态策略合并学习多目标行为

DOI:
--
复制
发表时间:
1993
期刊:
影响因子:
--
通讯作者:
J. Tenenberg
J. Tenenberg
中科院分区:
--
文献类型:
--
作者:
S. Whitehead;Jonas Karlsson;J. Tenenberg

文献摘要

被引文献

相似文献

协调追求多个时变目标的能力对智能机器人来说很重要。在本章中,我们考虑了强化学习在一类简单的动态多目标任务中的应用。毫不奇怪,我们发现最直接、最单一的方法可伸缩性很差,因为状态空间的大小与目标的数量呈指数关系。作为替代方案,我们提出了一种简单的模块化体系结构,将学习和控制任务分布在一组独立的控制模块中,每个控制模块对应于代理可能遇到的每个目标。由于每个模块学习与其目标相关联的最佳策略,而不考虑其他当前目标,因此便于学习。与单片控制器相比,这极大地简化了状态表示并加快了学习时间。当机器人面对单一目标时,与该目标相关联的模块用于确定总体控制策略。当机器人面临多个目标时,来自每个相关模块的信息被合并以确定组合任务的策略。总体而言,这些合并策略产生了良好但不太理想的性能。因此,该体系结构在较差的初始性能、较慢的学习和最优的渐近策略之间进行了折衷,从而有利于较好的初始性能、快速的学习和略微次优的渐近策略。我们考虑了几种合并策略,从简单的只比较和组合关于当前状态的模块化信息的策略,到使用先行搜索构建更准确的效用估计的更复杂的策略。
An ability to coordinate the pursuit of multiple, time-varying goals is important to an intelligent robot. In this chapter we consider the application of reinforcement learning to a simple class ofdynamicmulti-goal tasks.Not surprisingly, we find that the most straightforward, monolithic approach scales poorly, since the size of the state space is exponential in the number of goals. As an alternative, we propose a simple modular architecture which distributes the learning and control task amongst a set of separate control modules, one for each goal that the agent might encounter. Learning is facilitated since each module learns the optimal policy associated with its goal without regard for other current goals. This greatly simplifies the state representation and speeds learning time compared to a single monolithic controller. When the robot is faced with a single goal, the module associated with that goal is used to determine the overall control policy. When the robot is faced with multiple goals, information from each associated module is merged to determine the policy for the combined task. In general, these merged strategies yield good but suboptimal performance. Thus, the architecture trades poor initial performance, slow learning, and an optimal asymptotic policy in favor of good initial performance, fast learning, and a slightly sub-optimal asymptotic policy. We consider several merging strategies, from simple ones that compare and combine modular information about the current state only, to more sophisticated strategies that use lookahead search to construct more accurate utility estimates.