课题基金 / 基金详情

Self-control of Memory Structure of Reinforcement Learning in Hidden Markov Environments

Self-control of Memory Structure of Reinforcement Learning in Hidden Markov Environments
隐马尔可夫环境下强化学习记忆结构的自我控制
批准号:
11650441
负责人:
ABE Kenichi
金额:
$2.24万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
1999
资助国家:
日本
项目状态:
已结题
起止时间:
1999 至 2000

项目摘要

项目成果

ABE Kenichi的其他基金

相似基金

相关文献

中文摘要
翻译
近年来,强化学习算法的研究主要集中在部分可观测马尔可夫决策问题上。POMDP的一种可能的解决方案是使用历史信息来估计状态。Q值必须以反映观察/行动对过去历史的形式进行更新。在本研究中,我们提出了两种强化学习方法,可以解决某些类型的POMDP问题。研究结果总结如下:(1)在上一期科研助学金(C)(2)的基础上,我们提出了标记Q-学习(LQ-Learning),它具有一种新的处理过去历史的记忆结构。在本研究中,我们建立了学习商学习的总体框架。设计了该框架中的各种算法,并通过仿真对这些算法进行了比较研究。然而,上述LQ学习有一个缺点,即我们必须预先定义标记机制。为了克服这一缺点,我们进一步设计了一种SOM(自组织特征映射)标记方法,在该方法中,将观察/动作对的过去历史划分为类。SOM具有一维结构,其输出节点产生标签。(2)提出了一种新的分层RL,称为开关Q-学习。SQ学习的基本思想是,非马尔可夫任务可以被自动分解成可由无记忆策略解决的子任务,而不需要任何其他信息来实现好的子目标。为了处理这种分解,SQ-学习采用了Q-模块的有序序列,其中每个模块发现一个局部控制策略。SQ-Learning采用层次化的学习自动机系统进行交换模块的学习。仿真结果表明,SQ-学习能够在不增加巨大计算负担的情况下,快速学习最优或接近最优的策略,建立一个系统地处理LQ-学习和SQ-学习的统一视图是下一步的工作。
英文摘要
Recent research on reinforcement learning (RL) algorithms has concentrated on partially observable Markov decision problems (POMDPs). A possible solution to POMDPs is to use history information to estimate state. Q values must be updated in the form reflecting past history of observation/action pairs. In this study, we developed two methods of reinforcement learning, which can solve certain types of POMDPs. The results are summarized as follows :(1) As a result of last Grant-in-Aid for Scientific Research (C)(2), we proposed Labeling Q-learning (LQ-learning), which has a new memory architecture of handling past history. In this study, we established a general framework of the LQ-learning. Various algorithms in this framework were devised, and we gave comparative study between these through simulation. The above LQ-learning, however, has the drawback that we must predefine the labeling mechanism. To overcome this drawback, we further devised a SOM (self-organizing feature map) approach of labeling, in which past history of observation/action pairs are partitioned into classes. The SOM has one-dimensional structure and the output nodes of the SOM produce labels.(2) We proposed a new type of hierarchical RL, called Switching Q-learning (SQ-learning). The basic idea of SQ-learning is that non-Markovian tasks can be automatically decomposed into subtasks solvable by memoryless policies, without any other information leading to "good" subgoals. To deal with such decomposition, SQ-learning employs ordered sequences of Q-modules in which each module discovers a local control policy. SQ-learning uses a hierarchical system of learning automata for switching module. The simulation results demonstrate that SQ-learning has the ability to quickly learn optimal or near-optimal policies without huge computational burden.It is a future work to build a unified view by which LQ-learning and SQ-learning can be dealt with systematically.
期刊论文(45)
专著(0)
科研奖励(0)
会议论文
Hae Yeon Lee: "Labeling Q-learning for Maze Problems with Partially Observable States"Proc.of 15th Korea Automatic Control Conference. Vol.2. 484-487 (2000)
Hae Yeon Lee:“Labeling Q-learning for Maze Problems with Partially Observable States”,第 15 届韩国自动控制会议论文集。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
Haeyon Lee: "Labeling Q-Learning for Partially Observable Markov Decision Process Environments"Proc.of Fifth Int.Symp.on Artificial Life and Robtics. 484-490 (2000)
Haeyon Lee:“Labeling Q-Learning for Partially Observable Markov Decision Process Environmentals”Proc.of Fifth Int.Symp.on Artificial Life and Robtics。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
HaeYeon Lee: "Labeling Q-Learning For Non-Markovian Environments"1999 IEEE International Conference on SMC. Vol.V. 487-491 (1999)
HaeYeon Lee:“为非马尔可夫环境标记 Q 学习”1999 年 IEEE 国际 SMC 会议。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
HaeYeon Lee: "Labeling Q-learning for partially observable markov decision process environments"AROB 5th '00. Vol.2. 281-284 (2000)
HaeYeon Lee:“为部分可观察的马尔可夫决策过程环境标记 Q 学习”AROB 5th 00。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
共 44 条
    Studies on Literary History in Bohemia
    • 批准号:
      19K00493
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $2.75万
    • 财政年份:
      2019
    • 负责人:
      ABE Kenichi
    • 依托单位:
    Studies on Images of "East" in East European Literature
    • 批准号:
      24320064
    • 项目类别:
      Grant-in-Aid for Scientific Research (B)
    • 资助金额:
      $6.99万
    • 财政年份:
      2012
    • 负责人:
      ABE Kenichi
    • 依托单位:
    Self-Organization of Hierarchical Reinforcement Learning System
    • 批准号:
      13650480
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $2.18万
    • 财政年份:
      2001
    • 负责人:
      ABE Kenichi
    • 依托单位:
    Study on Decentralized Learning Algorithms in Non-Markovian Environments
    • 批准号:
      09650451
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $1.6万
    • 财政年份:
      1997
    • 负责人:
      ABE Kenichi
    • 依托单位:
    海外基金