Computational Theory of Inductive Reinforcement Learning-Bayesian Inference on Environment Search and Inductive Reconstruction
Computational Theory of Inductive Reinforcement Learning-Bayesian Inference on Environment Search and Inductive Reconstruction
批准号:
20700126
负责人:
MAKINO Takaki
金额:
$1.5万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Young Scientists (B)
财政年份:
2008
资助国家:
日本
项目状态:
已结题
起止时间:
2008 至 2010
中文摘要
本文主要研究基于贝叶斯推理技术的强化学习环境模型重构问题。在强化学习中,智能体通过反复试验学习环境模型,如果有一个合适的贝叶斯环境模型来表示环境中的不确定性,就可以实现最优探索。为此,我们提出了一种基于预测状态表示的环境描述框架TD-Network的改进方法。此外,我们对隐马尔可夫模型的非参数贝叶斯模型进行了扩展,以表示隐藏状态的层次聚类。此外,我们应用学徒式学习的框架,提出了一种基于贝叶斯推理的从他人行为中构建环境模型的方法。这些都是环境搜索和重建过程的贝叶斯重建所必需的要素。
英文摘要
This study focuses on environmental model reconstruction in reinforcement learning based on Bayesian inference techniques. In reinforcement learning, an agent learns environment model by trial-and-error; if we have a suitable Bayesian environment model that represents uncertainty in the environment, an optimal exploration can be achieved. For this purpose, we proposed new approaches that improve TD-network, an environment description framework based on predictive state representation. In addition, we extended a nonparametric Bayesian model for hidden Markov model to represent hierarchical clustering of hidden states. Moreover, we applied the framework of apprenticeship learning and proposed a method that constructs environment model from other’s actions based on Bayesian inference. These are elements that are required for Bayesian reconstruction of the process of environmental search and reconstruction.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
ベイズ確率文脈自由文法のための高速構文木サンプリング法
贝叶斯概率上下文无关文法的快速解析树采样方法
DOI:
--
发表时间:
2009
期刊:
影响因子:
--
作者:
[武井俊祐, 牧野貴樹, 高木利久]
通讯作者:
高木利久
DOI:
--
发表时间:
2009
期刊:
影响因子:
--
作者:
[Takaki Makino, Shunsuke Takei, Daichi Mochihashi, Issei Sato, Toshihisa Takagi]
通讯作者:
Toshihisa Takagi
階層状態無限隠れマルコフモデル
分层状态无限隐马尔可夫模型
DOI:
--
发表时间:
2009
期刊:
影响因子:
--
作者:
[尾崎有梨, 櫻井加奈子, 浅井健一, 戸次大介, 牧野貴樹]
通讯作者:
牧野貴樹
DOI:
--
发表时间:
2008
期刊:
影响因子:
--
作者:
[Takaki Makino, Taiki Takahashi, Hirofumi Nishinaka, and Hiroki Fukui, 戸次大介, Takaki Makino]
通讯作者:
Takaki Makino
POのP環境中でのTD-Networkの自動獲得 : 単純再帰構造による拡張
PO的P环境中TD-Network的自动获取:简单递归结构的扩展
DOI:
--
发表时间:
2008
期刊:
影响因子:
--
作者:
[芳中隆幸, 福原知宏, 増田英孝, 中川裕志, 牧野貴樹]
通讯作者:
牧野貴樹
共 23 条