Deep Reinforcement Learning With Explicit Context Representation

Deep Reinforcement Learning With Explicit Context Representation
复制标题

具有显式上下文表示的深度强化学习

DOI:
10.1109/tnnls.2023.3325633
复制
发表时间:
2023
影响因子:
10.4
通讯作者:
Munguia-Galeano F
Munguia-Galeano F
中科院分区:
计算机科学1区
文献类型:
--
作者:
Munguia-Galeano F

文献摘要

相似文献

虽然强化学习(RL)已经显示出解决复杂计算问题的出色能力,但大多数RL算法缺乏允许从上下文信息中学习的显式方法。另一方面,人类经常使用上下文来识别环境中元素之间的模式和关系,沿着如何避免做出错误的动作。然而,从人类的角度来看,这似乎是一个明显错误的决定,RL代理可能需要数百个步骤来学习避免。本文提出了一个离散环境的框架,称为Iota显式上下文表示(IECR)。该框架涉及使用上下文关键帧(CKF)来表示每个状态,然后可以使用该上下文关键帧来提取表示状态的示能表示的函数;此外,针对状态的示能表示引入了两个损失函数。IECR框架的新奇在于它能够从环境中提取上下文信息并从CKF的表示中学习。我们通过开发四种使用上下文学习的新算法来验证该框架:Iota深度Q网络(IDQN),Iota双深度Q网络(IDDQN),Iota决斗深度Q网络(IDuDQN)和Iota决斗双深度Q网络(IDDDQN)。此外,我们评估的框架和新的算法在五个离散的环境。我们表明,所有使用上下文信息的算法都收敛在神经网络的大约40000个训练步骤中,显著优于其最先进的同类算法。
Though reinforcement learning (RL) has shown an outstanding capability for solving complex computational problems, most RL algorithms lack an explicit method that would allow learning from contextual information. On the other hand, humans often use context to identify patterns and relations among elements in the environment, along with how to avoid making wrong actions. However, what may seem like an obviously wrong decision from a human perspective could take hundreds of steps for an RL agent to learn to avoid. This article proposes a framework for discrete environments called Iota explicit context representation (IECR). The framework involves representing each state using contextual key frames (CKFs), which can then be used to extract a function that represents the affordances of the state; in addition, two loss functions are introduced with respect to the affordances of the state. The novelty of the IECR framework lies in its capacity to extract contextual information from the environment and learn from the CKFs’ representation. We validate the framework by developing four new algorithms that learn using context: Iota deep Q-network (IDQN), Iota double deep Q-network (IDDQN), Iota dueling deep Q-network (IDuDQN), and Iota dueling double deep Q-network (IDDDQN). Furthermore, we evaluate the framework and the new algorithms in five discrete environments. We show that all the algorithms, which use contextual information, converge in around 40 000 training steps of the neural networks, significantly outperforming their state-of-the-art equivalents.