SAN-RL: combining spreading activation networks and reinforcement learning to learn configurable behaviors

SAN-RL: combining spreading activation networks and reinforcement learning to learn configurable behaviors
复制标题

SAN-RL:结合传播激活网络和强化学习来学习可配置行为

DOI:
10.1117/12.457458
复制
发表时间:
2002
期刊:
--
影响因子:
--
通讯作者:
John H. White
John H. White
中科院分区:
--
文献类型:
--
作者:
D. Gaines;D. Wilkes;Kanok Kusumalnukool;Siripun Thongchai;K. Kawamura;John H. White

文献摘要

被引文献

相似文献

强化学习技术已经成功地允许代理学习用于实现任务的策略。智能体的整体行为可以通过适当的奖励函数来控制。然而,学习到的策略将固定于该奖励函数。如果用户希望改变他或她关于如何完成任务的偏好,则必须使用该新的奖励函数重新训练代理。我们通过将扩展激活网络和强化学习结合在一种我们称为SAN-RL的方法中来解决这一挑战。这种方法为智能体提供了一个因果结构,即扩散激活网络,将目标与可以实现这些目标的行动联系起来。这使代理能够选择与目标优先级相关的操作。我们将联合收割机与强化学习相结合,使代理能够学习策略。总之,这些方法使得能够学习可配置的行为,可以适应以满足当前偏好的策略。我们比较的方法与Q-学习机器人导航任务。我们证明了SAN-RL在学习之前表现出目标导向的行为,利用网络的因果结构在学习过程中集中搜索,并在学习后产生可配置的行为。
Reinforcement learning techniques have been successful in allowing an agent to learn a policy for achieving tasks. The overall behavior of the agent can be controlled with an appropriate reward function. However, the policy that is learned will be fixed to this reward function. If the user wishes to change his or her preference about how the task is achieved the agent must be retrained with this new reward function. We address this challenge by combining Spreading Activation Networks and Reinforcement Learning in an approach we call SAN-RL. This approach provides the agent with a causal structure, the spreading activation network, relating goals to the actions that can achieve those goals. This enables the agent to select actions relative to the goal priorities. We combine this with reinforcement learning to enable the agent to learn a policy. Together, these approaches enable the learning of a configurable behaviors, a policy that can be adapted to meet the current preferences. We compare the approach with Q-learning on a robot navigation task. We demonstrate that SAN-RL exhibits goal-directed behavior before learning, exploits the causal structure of the network to focus its search during learning and results in configurable behaviors after learning.