Studies on Learning Algorithms for Flexibly Structured Decision Process Models
Studies on Learning Algorithms for Flexibly Structured Decision Process Models
批准号:
18540111
负责人:
KURANO Masami
金额:
$1.88万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2006
资助国家:
日本
项目状态:
已结题
起止时间:
2006 至 2007
中文摘要
本研究的目标是建立一种具有更柔性结构的不确定决策过程的自适应强化学习算法。主要研究成果如下。1.(a)研究模糊决策的可能性和可信性,并应用其扩张定理,成功地从给定的条件可信性测度出发,构造了模糊决策过程,从而使模糊环境下决策过程的公理化发展成为可能。(b)我们已经成功地推导出灵活的最优性方程的吸收半马尔可夫游戏的一般效用函数,确定最优策略。针对质量控制问题的贝叶斯分析,提出了一种新的控制图,它具有更灵活的结构,通过先验测度区间来把握未知参数。通过与常用的两种方法的比较,说明了新方法的有效性.自适应马尔可夫决策模型(MDP)的学习算法我们提出了一种自适应MDP的模式-矩阵学习算法,该算法从观测数据中学习转移矩阵的结构(模式),并利用其信息构造基于时间差分(TD)方法的自适应策略。该方法基本上适用于多链情形。3.强化学习方法的应用(a)我们研究了适用于各种神经元动态规划模型的TD或Actor Critic算法的收敛性,通过数值实验发现了其效率。(b)为了求解模糊环境下的运筹学模型,提出了一种融合模糊模拟和遗传算法的混合智能算法,并通过算例验证了该算法的有效性。
英文摘要
In this project, our objective is to establish the adaptive and reinforcement learning algorithms for uncertain decision processes with the more flexible and soft structure. The main research results are as follows. 1. Further studies on construction and analysis of flexibly structured models (a) Investigating possibility and credibility of fuzziness and applying its extension theorem, we have succeeded in constructing credibilistic process from given conditional credibility measures, by which axiomatic development of decision processes under fuzzy environment will be made to be possible. (b) We have succeeded in deriving the flexible optimality equations for an absorbing semi-Markov game with general utility functions which determine the optimal strategies. c Concerning Bayesian analysis for a quality control problem, we have proposed the new control chart which has more flexible structure, grasping the unknown parameter by a priori interval of measures. The efficiency of the new one is shown by comparing with the usual one 2. Learning algorithms for adaptive Markov decision models (MDPs) We have developed a pattern-matrix learning algorithm for adaptive MDPs which learns the structure (pattern) of transition matrices from the observed data and using its information constructs the adaptive policy based on temporal difference (TD) method. This method can be essentially applicable to the multichain case. 3. Application of reinforcement learning methods (a) We have investigated the convergence of the TD or Actor Critic algorithms applicable to various models of neuron dynamic programming, finding its efficiency by numerical experiments. (b) In order to solve several Operational Research models under fuzzy environments, we have developed the Hybrid Intelligent algorithm integrating fuzzy simulation and genetic algorithm, whose efficiency is verified by a numerical examples.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Fuzzy optimality relation for perception MDPs the average case
平均情况下感知 MDP 的模糊最优关系
DOI:
--
发表时间:
2007
期刊:
Fuzzy Sets and Systems 158
影响因子:
--
作者:
[Masaki Kameko, Mamoru Mimura, Masaki Nakagawa, Masatoshi Kokubu, 國分 雅敏, Mamoru Mimura, Masatoshi Kokubu, 蔵野正美(共著)]
通讯作者:
蔵野正美(共著)
Adaptive Markov decision processes based on temporal difference method
基于时间差分法的自适应马尔可夫决策过程
DOI:
--
发表时间:
2007
期刊:
影响因子:
--
作者:
[堀口正之, 安田正實, 堀口正之, 堀口正之]
通讯作者:
堀口正之
Fuzzy optimality relation for perceptive MDPs - the average case
感知 MDP 的模糊最优性关系 - 平均情况
DOI:
--
发表时间:
2007
期刊:
Fuzzy Sets and Systems Vol. 158
影响因子:
--
作者:
[Kurano, M., Yasuda, M., Nakagami, J., and Yoshida, Y.]
通讯作者:
Y.
A structured pattern matrix Algorithm for multichain Markov decision processes
多链马尔可夫决策过程的结构化模式矩阵算法
DOI:
--
发表时间:
2007
期刊:
Mathematical Methods of Operations Research Vol. 66
影响因子:
--
作者:
[Iki, T., Horiguchi, M., and Kurano, M.]
通讯作者:
M.
Adaptive Markov decision processes based on difference method
基于差分法的自适应马尔可夫决策过程
DOI:
--
发表时间:
2007
期刊:
影响因子:
--
作者:
[Iki, T., Horiguchi, M., Yasuda, M., and Kurano, M., 伊喜哲 一郎(共同)]
通讯作者:
伊喜哲 一郎(共同)
Studies on Flexibly Structured Models of Decision Making
-
批准号:15540105
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.18万
-
财政年份:2003
-
负责人:KURANO Masami
-
依托单位:
Studies on Flexible Structure of Dynamic Programming
-
批准号:12640104
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.24万
-
财政年份:2000
-
负责人:KURANO Masami
-
依托单位:
Studies on Mathematical Structure of Dynamic Programming
-
批准号:09640243
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$1.6万
-
财政年份:1997
-
负责人:KURANO Masami
-
依托单位:
海外基金