课题基金 / 基金详情

Study on Decentralized Learning Algorithms in Non-Markovian Environments

Study on Decentralized Learning Algorithms in Non-Markovian Environments
非马尔可夫环境下的分散学习算法研究
批准号:
09650451
负责人:
ABE Kenichi
金额:
$1.6万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
1997
资助国家:
日本
项目状态:
已结题
起止时间:
1997 至 1998

项目摘要

项目成果

ABE Kenichi的其他基金

相似基金

相关文献

中文摘要
翻译
The results of this study are summarized as follows:(1)A formal model of non-Markovian problems is the partially observable Markov decision problem(POMDP)。The most useful solution to overcome partial observability is to use memory to estimate state。In this study,we proposed a new memory architecture of reinforcement learning algorithms to solve certain type of POMDPs.The agent‘s task is to discover a path leading from start position to goal in a partially observable maze。The agent is assumed to have life-time separable into“trials”.The basic framework of the algorithm,called labeling Q-learning,is described as follows.Let0be the set of finite observations。At each step t,when the agent gets an observation o_t epsilon OMICRON from the environment,a label,theta_t is attached to the observation,where theta_t is an element of THETA={0,1,2,·,M-1},(in the beginning of each trial,the labels for all omicron_t epsilon OMICare initialized to0).Then the pair OMICRON_t=(OMICRON_t*THETA_t)defines a nobservation,and the usual initialized to0).Then the pair OMICRON_t=(OMICRON_t*THETA_t)defines a nobservation,and the usual reinitialized to0)。Theta_t)has the Markov property.(2)The labeling Q-learning was applied to test problems of simple mazes taken from the recent literature.The results demonstrated labeling Q-learning‘s ability to work well in near-optimal manner.(3)Most problems will have continuous or large discrete observation space。We studied generalization techniques by recurrentneural networks(RNN)and holon networks,which allow compact storage of similar observations.Further,we developed an approximate method of controlling the complexity,i.e.,the Lyapunov exponent,of RNNs,and the method was demonstrated by applying it to identification problems of certain nonlinear systems.(4)We made fundamental experiments on sensor-based navigation for a mobile robot.
英文摘要
The results of this study are summarized as follows :(1) A formal model of non-Markovian problems is the partially observable Markov decision problem (POMDP). The most useful solution to overcome partial observability is to use memory to estimate state. In this study, we proposed a new memory architecture of reinforcement learning algorithms to solve certain type of POMDPs.The agent's task is to discover a path leading from start position to goal in a partially observable maze. The agent is assumed to have life-time separable into "trials". The basic framework of the algorithm, called labeling Q-learning, is described as follows.Let 0 be the set of finite observations. At each step t, when the agent gets an observation o_t epsilon OMICRON from the environment, a label, theta_t is attached to the observation, where theta_t is an element of THETA={0, 1, 2, ・, M -1}, (in the beginning of each trial, the labels for all omicron_t epsilon OMICRON are initialized to 0).Then the pair OMICRON_t=(OMICRON_t*THETA_t) defines a new observation, and the usual reinforcementlearning algorithm TD( lambda) that uses replacing traces is applied to OMICRON=OMICRON*THETA, as if the pair = (omicron_t, theta_t) has the Markov property.(2) The labeling Q-learning was applied to test problems of simple mazes taken from the recent literature. The results demonstrated labeling Q-learning's ability to work well in near-optimal manner.(3) Most problems will have continuous or large discrete observation space. We studied generalization techniques by recurrentneural networks(RNN) and holon networks, which allow compact storage of similar observations. Further, we developed an approximate method of controlling the complexity, i.e., the Lyapunov exponent, of RNNs, and the method was demonstrated by applying it to identification problems of certain nonlinear systems.(4) We made fundamental experiments on sensor-based navigation for a mobile robot.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
本間経康: "神経回路網ダイナミクスの複雑さの制御法" 計測自動制御学会論文集. 35・1. 138-143 (1999)
本间常康:“神经网络动力学复杂性的控制方法”仪器与控制工程师学会学报 35・1(1999)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
喜多川 健: "リカレントニューラルネットワークの創発的学習手法" 計測自動制御学会論文集. 33巻11号. 1093-1098 (1997)
Ken Kitakawa:“循环神经网络的紧急学习方法”《仪器与控制工程师协会学报》,第 33 卷,第 11 期。1093-1098(1997 年)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
Fation Sevrani: "On the synthesis of brain-state-in-a-box neural models with application to associative" Neural Computation. (in press).
Fation Sevrani:“关于脑状态盒式神经模型的综合及其应用于关联”神经计算。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
喜多川 健: "リカレントニューラルネットワークの創発的学習手法" 計測自動制御学会論文集. 33・11. 1093-1098 (1997)
北川健:“循环神经网络的紧急学习方法”仪器与控制工程师协会会议记录 33・11(1997)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
18
    Studies on Literary History in Bohemia
    • 批准号:
      19K00493
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $2.75万
    • 财政年份:
      2019
    • 负责人:
      ABE Kenichi
    • 依托单位:
    Studies on Images of "East" in East European Literature
    • 批准号:
      24320064
    • 项目类别:
      Grant-in-Aid for Scientific Research (B)
    • 资助金额:
      $6.99万
    • 财政年份:
      2012
    • 负责人:
      ABE Kenichi
    • 依托单位:
    Self-Organization of Hierarchical Reinforcement Learning System
    • 批准号:
      13650480
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $2.18万
    • 财政年份:
      2001
    • 负责人:
      ABE Kenichi
    • 依托单位:
    Self-control of Memory Structure of Reinforcement Learning in Hidden Markov Environments
    • 批准号:
      11650441
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $2.24万
    • 财政年份:
      1999
    • 负责人:
      ABE Kenichi
    • 依托单位:
    海外基金