Biologically Inspired Reinforcement Learning: Reward-Based Decomposition for Multi-goal Environments

Biologically Inspired Reinforcement Learning: Reward-Based Decomposition for Multi-goal Environments
复制标题

DOI:
10.1007/978-3-540-27835-1_7
复制
发表时间:
2004-01
期刊:
--
影响因子:
--
通讯作者:
Weidong Zhou;R. Coggins
Weidong Zhou;R. Coggins
中科院分区:
其他
文献类型:
--
作者:
Weidong Zhou;R. Coggins

文献摘要

被引文献

相似文献

我们提出了一个基于情感的分层强化学习(HRL)算法的环境中有多个奖励来源。该系统的架构受到大脑神经生物学的启发,特别是那些负责情绪、决策和行为执行的区域,分别是杏仁核、眶额皮质和基底神经节。学习问题根据奖励来源进行分解。奖励源作为给定子任务的目标。每个子任务被分配一个人工情感指示(AEI),它预测与子任务相关联的奖励成分。AEI与顶层策略同时沿着学习,并用于在AEI发生显著变化时中断子任务执行。该算法在一个模拟的gridworld有两个来源的奖励,是部分可观察的。在相同的学习条件下,将基于情感的算法与其他HRL算法进行了实验比较。与人类设计的策略和MAXQ算法的受限形式相比,生物启发架构的使用显著加速了学习过程,并实现了更高的长期回报。
We present an emotion-based hierarchical reinforcement learning (HRL) algorithm for environments with multiple sources of reward. The architecture of the system is inspired by the neurobiology of the brain and particularly those areas responsible for emotions, decision making and behaviour execution, being the amygdala, the orbito-frontal cortex and the basal ganglia respectively. The learning problem is decomposed according to sources of reward. A reward source serves as a goal for a given subtask. Each subtask is assigned an artificial emotion indication (AEI) which predicts the reward component associated with the subtask. The AEIs are learned along with the top-level policy simultaneously and used to interrupt subtask execution when the AEIs change significantly. The algorithm is tested in a simulated gridworld which has two sources of reward and is partially observable. Experiments are performed comparing the emotion based algorithm with other HRL algorithms under the same learning conditions. The use of the biologically inspired architecture significantly accelerates the learning process and achieves higher long term reward compared to a human designed policy and a restricted form of the MAXQ algorithm.