An off-policy natural gradient method for a partial observable Markov decision process

An off-policy natural gradient method for a partial observable Markov decision process
复制标题

部分可观测马尔可夫决策过程的离策略自然梯度法

DOI:
--
复制
发表时间:
2005
期刊:
Artificial Neural Networks : Formal Models and Their Applications - ICANN 2005, Lecture Notes in Computer Science 3697
影响因子:
--
通讯作者:
Y.
Y.
中科院分区:
--
文献类型:
--
作者:
Nakamura;Y.

文献摘要

参考文献

被引文献

相似文献

基于在线变分贝叶斯方法的系统辨识及其在强化学习中的应用
DOI: --
发表时间: 2003
期刊: Artificial Neural Networks and Neural Information Processing, Lecture Notes in Computer Science (Berlin : Springer-Verlag) 2714
影响因子: --
作者:
Yoshimoto;J.
通讯作者: J.
用于双足机器人 CPG 控制的自然策略梯度强化学习
DOI: --
发表时间: 2004
期刊: Parallel Problem Solving from Nature - PPSN VIII, Lecture Notes in Computer Science 3242
影响因子: --
作者:
Nakamura;Y.
通讯作者: Y.