Self-Organization of Hierarchical Reinforcement Learning System
Self-Organization of Hierarchical Reinforcement Learning System
批准号:
13650480
负责人:
ABE Kenichi
金额:
$2.18万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2001
资助国家:
日本
项目状态:
已结题
起止时间:
2001 至 2002
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Previously, we proposed two learning algorithms, Labeling Q-learning(LQ-learning) and Switching Q-learning(SQ-learning). Although the former is the algorithm of simple structure which consists of a single agent, it can learn well in a certain kind of POMDP environments. The latter is a type of hierarchical Q-learning method (HQ-learning), which changes Q-modules by using a hierarchical learning automaton, and can work well also in a more complicated POMDP environment. In this study, we improved these two algorithms, and developed more effective HQ-learning algorithms. Further, in order to overcome more realistic environments where either or both of observations and actions take continuous values, we conducted a basic study about function approximations by neural networks. The results are following.1) We improved the SQ-learning so that it works well in noisy environments. We also demonstrated that the SQ-learning exhibits a better performance than Wiering's HQ-learning.2) We enhanced the performance of the LQ-leaning by introducing the Kohonen's self-organizing map(SOM).3) We improved the self-segmentation of sequence(SSS) algorithm by Sun and Sessions. Further, we also developed a new algorithm, called SSS(λ).4) We examined the effectiveness of SSS(λ) by applying it to the navigation task of a mobile robot. Here, the SOM was used for self-classification of continuous sonar observations.5) We proposed a statistical approximation learning(SAL) for the simultaneous recurrent neural networks, and demonstrated that it achieves the high accuracy of nonlinear function approximation. Further, we presented a novel neural network model for incremental learning.
期刊论文(42)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
M. Sakai: "Control of Chaos Dynamics in Jordan Recurrent Neural Networks"Proc. of the International Conference on Control, Automation and Systems. 292-295 (2001)
M. Sakai:“约旦循环神经网络中的混沌动力学控制”Proc。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
M. Sakai: "A Statistical Approximation Learning Method for Simultaneous Recurrent Networks"Proc. of the 15th IFAC World Congress on Automatic Control. 2491-2496 (2002)
M. Sakai:“同时循环网络的统计近似学习方法”Proc。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
H.Y.Lee: "Labeling Q-learning with SOM"Int. Conf.on Control, Automation, and Systems(ICCAS 2002). 105-109 (2002)
H.Y.Lee:“用 SOM 标记 Q 学习”Int。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
M.Sakai: "A Statistical Approximation Learning Method for Simultaneous Recurrent Networks"Proc.of the 15^<th> IFAC World Congress on Automatic Control. 2491-2496 (2002)
M.Sakai:第 15 届 IFAC 世界自动控制大会的“同时循环网络的统计近似学习方法”。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
M.Sakai: "Control of Chaos Dynamics in Jordan Recurrent Neural Networks"Proc.of the International Conference on Control, Automation and Systems. 292-295 (2001)
M.Sakai:“约旦循环神经网络中的混沌动力学控制”国际控制、自动化与系统会议论文集。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 41 条
Studies on Literary History in Bohemia
-
批准号:19K00493
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.75万
-
财政年份:2019
-
负责人:ABE Kenichi
-
依托单位:
Studies on Images of "East" in East European Literature
-
批准号:24320064
-
项目类别:Grant-in-Aid for Scientific Research (B)
-
资助金额:$6.99万
-
财政年份:2012
-
负责人:ABE Kenichi
-
依托单位:
Self-control of Memory Structure of Reinforcement Learning in Hidden Markov Environments
-
批准号:11650441
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.24万
-
财政年份:1999
-
负责人:ABE Kenichi
-
依托单位:
Study on Decentralized Learning Algorithms in Non-Markovian Environments
-
批准号:09650451
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$1.6万
-
财政年份:1997
-
负责人:ABE Kenichi
-
依托单位:
Study on Decentralized Learning Algorithms in Markovian Environments
-
批准号:06650449
-
项目类别:Grant-in-Aid for General Scientific Research (C)
-
资助金额:$1.34万
-
财政年份:1994
-
负责人:ABE Kenichi
-
依托单位: