课题基金 / 基金详情

Study on Decentralized Learning Algorithms in Non-Markovian Environments

Study on Decentralized Learning Algorithms in Non-Markovian Environments
非马尔可夫环境下的分散学习算法研究
批准号:
09650451
负责人:
ABE Kenichi
金额:
$1.6万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
1997
资助国家:
日本
项目状态:
已结题
起止时间:
1997 至 1998

项目摘要

项目成果

ABE Kenichi的其他基金

相似基金

相关文献

中文摘要
翻译
本研究的结果总结如下:(1)非马尔可夫问题的正式模型是部分可观察到的马尔可夫决策问题(POMDP)。超过部分可观测性的最有用的解决方案是使用内存来估计国家。在这项研究中,我们提出了一个新的内存结构学习算法,以解决POMDPs的特定类型。代理的任务是发现一个从起始位置到目标的路径在一个部分可观察的迷宫中。agent is assumed to have life-time separable into "trials"。算法的基本框架,称为Q-学习,并将其描述为后续,让0成为一组有限的观测。在每个步骤t,当代理从环境中观察O_t epsilon OMICRON,标签,theta_t附着到观察,在那里theta_t是THETA={0, 1, 2, ·, M-1},(在每个试验开始时,所有omicron_t epsilon OMICRON的标签都初始化为0),然后对OMICRON_t =(OMICRON_t * THETA_t)定义了一个新的观测和一般强化学习算法TD( lambda),因为如果将替换轨迹应用于OMICRON=OMICRON*THETA,则为pair = (omicron_t,theta_t)具有马尔可夫属性。(2)将Q-学习标记为“Q-学习”已应用于从最近的文学中测试简单的问题。结果演示标签Q-学习的能力在附近最佳礼仪上良好地工作。(3)大多数问题将有连续性或大规模的离散观测空间。我们研究了递归神经网络(RNN)和holon网络的通用化技术,其中允许类似观测的紧凑存储。更进一步,我们已经开发出一种接近控制复杂性的方法,即:Lyapunov的表现,RNNs的表现,以及该方法是通过应用来说明确定非线性系统的某些问题。(4)我们为移动机器人进行了基于传感器的导航进行了基础实验。
英文摘要
The results of this study are summarized as follows :(1) A formal model of non-Markovian problems is the partially observable Markov decision problem (POMDP). The most useful solution to overcome partial observability is to use memory to estimate state. In this study, we proposed a new memory architecture of reinforcement learning algorithms to solve certain type of POMDPs.The agent's task is to discover a path leading from start position to goal in a partially observable maze. The agent is assumed to have life-time separable into "trials". The basic framework of the algorithm, called labeling Q-learning, is described as follows.Let 0 be the set of finite observations. At each step t, when the agent gets an observation o_t epsilon OMICRON from the environment, a label, theta_t is attached to the observation, where theta_t is an element of THETA={0, 1, 2, ・, M -1}, (in the beginning of each trial, the labels for all omicron_t epsilon OMICRON are initialized to 0).Then the pair OMICRON_t=(OMICRON_t*THETA_t) defines a new observation, and the usual reinforcementlearning algorithm TD( lambda) that uses replacing traces is applied to OMICRON=OMICRON*THETA, as if the pair = (omicron_t, theta_t) has the Markov property.(2) The labeling Q-learning was applied to test problems of simple mazes taken from the recent literature. The results demonstrated labeling Q-learning's ability to work well in near-optimal manner.(3) Most problems will have continuous or large discrete observation space. We studied generalization techniques by recurrentneural networks(RNN) and holon networks, which allow compact storage of similar observations. Further, we developed an approximate method of controlling the complexity, i.e., the Lyapunov exponent, of RNNs, and the method was demonstrated by applying it to identification problems of certain nonlinear systems.(4) We made fundamental experiments on sensor-based navigation for a mobile robot.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
本間経康: "神経回路網ダイナミクスの複雑さの制御法" 計測自動制御学会論文集. 35・1. 138-143 (1999)
本间常康:“神经网络动力学复杂性的控制方法”仪器与控制工程师学会学报 35・1(1999)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
喜多川 健: "リカレントニューラルネットワークの創発的学習手法" 計測自動制御学会論文集. 33巻11号. 1093-1098 (1997)
Ken Kitakawa:“循环神经网络的紧急学习方法”《仪器与控制工程师协会学报》,第 33 卷,第 11 期。1093-1098(1997 年)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
Fation Sevrani: "On the synthesis of brain-state-in-a-box neural models with application to associative" Neural Computation. (in press).
Fation Sevrani:“关于脑状态盒式神经模型的综合及其应用于关联”神经计算。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
喜多川 健: "リカレントニューラルネットワークの創発的学習手法" 計測自動制御学会論文集. 33・11. 1093-1098 (1997)
北川健:“循环神经网络的紧急学习方法”仪器与控制工程师协会会议记录 33・11(1997)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
18
    Studies on Literary History in Bohemia
    • 批准号:
      19K00493
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $2.75万
    • 财政年份:
      2019
    • 负责人:
      ABE Kenichi
    • 依托单位:
    Studies on Images of "East" in East European Literature
    • 批准号:
      24320064
    • 项目类别:
      Grant-in-Aid for Scientific Research (B)
    • 资助金额:
      $6.99万
    • 财政年份:
      2012
    • 负责人:
      ABE Kenichi
    • 依托单位:
    Self-Organization of Hierarchical Reinforcement Learning System
    • 批准号:
      13650480
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $2.18万
    • 财政年份:
      2001
    • 负责人:
      ABE Kenichi
    • 依托单位:
    Self-control of Memory Structure of Reinforcement Learning in Hidden Markov Environments
    • 批准号:
      11650441
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $2.24万
    • 财政年份:
      1999
    • 负责人:
      ABE Kenichi
    • 依托单位:
    海外基金