Q-learning with exploration driven by internal dynamics in chaotic neural network

Q-learning with exploration driven by internal dynamics in chaotic neural network
复制标题

DOI:
10.1109/ijcnn48605.2020.9207114
复制
发表时间:
2020-07
期刊:
2020 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Toshitaka Matsuki;Souya Inoue;K. Shibata
Toshitaka Matsuki;Souya Inoue;K. Shibata
中科院分区:
其他
文献类型:
--
作者:
Toshitaka Matsuki;Souya Inoue;K. Shibata

文献摘要

相似文献

本文提出了一种基于混沌的强化学习(RL)方法,该方法使用了混沌神经网络(NN)函数,不仅具有Actor-Critic函数,而且具有Q-学习函数。在我们提出的基于混沌的RL中,探索是基于混沌神经网络中的内部动力学进行的,并期望通过学习使动力学变得理性。Q-学习是一种非常流行的RL方法,并被广泛应用于多个研究中。重点研究了Q-学习是否适用于基于混沌的RL。在此基础上,利用Q-学习算法,证明了在网格环境下,基于混沌RL的智能体能够学习目标任务。研究还表明,随着学习的进行,由内部混沌动力学引起的网络输出的不规则性减少,智能体可以自动从探索模式切换到开发模式。此外,还证实了该智能体能够适应环境的变化,并自动恢复探测。
This paper shows chaos-based reinforcement learning (RL) using a chaotic neural network (NN) functions not only with Actor-Critic, but also with Q-learning. In chaos-based RL that we have proposed, exploration is performed based on internal dynamics in a chaotic NN and the dynamics is expected to grow rational through learning. Q-learning is a very popular RL method and widely used in several researches. We focused on whether Q-learning can be adopted to chaos-based RL. Then we demonstrated the agent can learn a goal task in a grid world environment with chaos-based RL using Q-learning. It was also shown that, as learning progresses, irregularity in the network outputs originated from the internal chaotic dynamics decreases and the agent can automatically switch from exploration mode to exploitation mode. Moreover, it was confirmed that the agent can adapt to changes in the environment and automatically resume exploration.